PR spikersoft-backend#252 (feat/533-shared-objectstore) implements this: new SpikerSoft.Storage project with IObjectStore + S3ObjectStore + shared ObjectKeys.TryMapKey shim + AddS3ObjectStore()/Use…
QA watch 2026-07-13 ~18:40Z — removal is now zero-risk. The typo'd stack's task (spiekersoft-docker-monitor, on SERVER) exited cleanly ~16:40Z and sits 0/1 Complete (on-failure policy won't…
QA watch 2026-07-13 ~18:40Z — filter fix NOT landed, failures continuing.
artpipe main venv_setup.py:263-264 is still the unfiltered call: snapshot_download(weights, local_dir_use_symlinks=False)…
QA watch 2026-07-13 ~18:40Z — status: HALF landed, keep open.
✅ ModelEnvImages half: artpipe 268f9ad (merge f7c0b15) is in main — .gitea/workflows/model-env-images.yml:50 = runs-on: [ubuntu-amd…
Verdict: bad RAM, quarantined; card/driver exonerated. Crash cluster resolved pending DIMM replacement.
- Root cause: defective physical memory region ~0xFDFC00000 (identical pfn neighborhood…
Exit-0 wedge occurrence #3 (2026-07-13 ~18:2xZ, post-GRUB-reboot recovery): metadata-extractor + security-scanner + upload-coordinator all Complete-wedged again — upload pipeline down until…
ROOT CAUSE FOUND + MITIGATION APPLIED (2026-07-13 ~18:1xZ, by joey): a bad memory sector was identified on SERVER and a GRUB memory exclusion applied for the affected 16 MB region (badram/memma…
Recurrence #8 (2026-07-13 ~18:0xZ): another flap-and-rejoin, full SERVER tier restarting (recovery in progress, expecting hands-off). **The interval is shrinking: ~45 min since #7, vs the 90-100…
Crash-#5 forensics change this ticket's thesis: prime suspect is now the iomemory-vsl4 driver, not RAM.
Evidence (journalctl, boot -1):
kernel BUG at mm/rmap.c:1102/folio_mkcleanat…
Crash #5 (~2026-07-13 16:20 UTC): SERVER dropped SSH mid-command and went unreachable ~10 minutes after a MinIO 'cluster not ready' wedge that itself followed a canceled registry push. Notable…
Recurrence #7 (2026-07-13 ~17:2xZ) — and the early-warning model is now 2-for-2: fusionio drive dropped ~17:00Z (#504), node crashed ~20 min later, same as the 02:31→02:50 sequence. This is…
⚠️ Early-warning fired (2026-07-13 ~17:0xZ): the Fusion-io device just dropped again (minio drives-online: 0, see #504) while SERVER reports Ready. Last time this exact sequence preceded the node…
Drive-drop recurrence (2026-07-13 ~17:0xZ): minio exited clean, restart:any respawned it, and the new task loops 'Read/Write quorum could not be established... drives-online: 0' — /mnt/fusionio/mi…
Crash #4: ~2026-07-13 06:2x UTC — during the sdxllightning image push (multi-GB blob commit into MinIO on /mnt/fusionio). That makes 3 of tonight's 4 crashes coincide with sustained MinIO ingest.…
Recurrence #6 (2026-07-13 ~15:1xZ, user-reported): SERVER crashed and this time recovered fully automatically within ~3 min of rejoining — all pinned services (minio/backend/mongo-router/mails…
Second disk-full within 6h (2026-07-13 ~08:03Z→~14:50Z): laptop-server filled again after the earlier ~05:15Z cleanup, killing gitea_postgres (and with it git/API/CI) for ~7h until this morning's…
Recovery timeline for last night's #5 (closing the loop): SERVER went down ~06:27Z and stayed down ~8h until the morning power-cycle (~14:35Z). Once the node rejoined, recovery was **fully…
Recurrence #5 (2026-07-13 ~06:27Z): SERVER Down/Unreachable again. Cadence is now firmly periodic: 02:50 → ~04:25 → ~06:27, roughly every 90–100 min tonight. That regularity plus load-correl…
Third failure mode confirmed — 98 GB pushes don't just fail, they HANG (2026-07-13 06:15Z). Run 10784 (with the 576b4bd cache sweep but NOT the snapshot_download filter) rebuilt sdxllightning…