Joseph Spiker spikerj
  • Joined on 2025-11-03
spikerj commented on issue spikerj/spikersoft-issues#533 2026-07-13 18:51:11 +00:00
[Backend][Refactor] Shared IObjectStore/S3 registration — dedupe per-service AmazonS3Client blocks (epic #413)

PR spikersoft-backend#252 (feat/533-shared-objectstore) implements this: new SpikerSoft.Storage project with IObjectStore + S3ObjectStore + shared ObjectKeys.TryMapKey shim + AddS3ObjectStore()/Use…

spikerj opened issue spikerj/spikersoft-issues#539 2026-07-13 18:42:24 +00:00
[Infra][Node] dreamstream6 (aarch64) down since ~16:30Z 2026-07-13 — heartbeat failure, host unreachable
spikerj commented on issue spikerj/spikersoft-issues#513 2026-07-13 18:42:10 +00:00
[Bug][Infra][Deploy] docker-monitor: new correctly-named stack exits 0 on startup (stuck 0/1) + duplicate misspelled 'spiekersoft-docker-monitor' stack running in parallel

QA watch 2026-07-13 ~18:40Z — removal is now zero-risk. The typo'd stack's task (spiekersoft-docker-monitor, on SERVER) exited cleanly ~16:40Z and sits 0/1 Complete (on-failure policy won't…

spikerj commented on issue spikerj/spikersoft-issues#537 2026-07-13 18:42:07 +00:00
[ArtPipe][CI] Phase 2 blocked: sdxllightning env image is 98 GB — unfiltered snapshot_download bakes the entire SDXL repo; push dies

QA watch 2026-07-13 ~18:40Z — filter fix NOT landed, failures continuing.

artpipe main venv_setup.py:263-264 is still the unfiltered call: snapshot_download(weights, local_dir_use_symlinks=False)…

spikerj commented on issue spikerj/spikersoft-issues#534 2026-07-13 18:41:53 +00:00
[CI][ArtPipe] Route baked model-image builds (ModelEnvImages + tier-3 finals) to the 4090 runner, off laptop-server

QA watch 2026-07-13 ~18:40Z — status: HALF landed, keep open.

ModelEnvImages half: artpipe 268f9ad (merge f7c0b15) is in main — .gitea/workflows/model-env-images.yml:50 = runs-on: [ubuntu-amd…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 17:14:26 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Verdict: bad RAM, quarantined; card/driver exonerated. Crash cluster resolved pending DIMM replacement.

  • Root cause: defective physical memory region ~0xFDFC00000 (identical pfn neighborhood…
spikerj commented on issue spikerj/spikersoft-issues#510 2026-07-13 17:10:58 +00:00
[Exploration][Infra] Resilience to SERVER node flaps — restart-policy windows, auto-reconcile sweep, de-SPOF the SERVER-pinned stateful tier

Exit-0 wedge occurrence #3 (2026-07-13 ~18:2xZ, post-GRUB-reboot recovery): metadata-extractor + security-scanner + upload-coordinator all Complete-wedged again — upload pipeline down until…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 17:04:38 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

ROOT CAUSE FOUND + MITIGATION APPLIED (2026-07-13 ~18:1xZ, by joey): a bad memory sector was identified on SERVER and a GRUB memory exclusion applied for the affected 16 MB region (badram/memma…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 17:01:28 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recurrence #8 (2026-07-13 ~18:0xZ): another flap-and-rejoin, full SERVER tier restarting (recovery in progress, expecting hands-off). **The interval is shrinking: ~45 min since #7, vs the 90-100…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 16:30:16 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Crash-#5 forensics change this ticket's thesis: prime suspect is now the iomemory-vsl4 driver, not RAM.

Evidence (journalctl, boot -1):

  • kernel BUG at mm/rmap.c:1102 / folio_mkclean at…
spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 16:25:00 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Crash #5 (~2026-07-13 16:20 UTC): SERVER dropped SSH mid-command and went unreachable ~10 minutes after a MinIO 'cluster not ready' wedge that itself followed a canceled registry push. Notable…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 16:23:11 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recurrence #7 (2026-07-13 ~17:2xZ) — and the early-warning model is now 2-for-2: fusionio drive dropped ~17:00Z (#504), node crashed ~20 min later, same as the 02:31→02:50 sequence. This is…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 16:19:09 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

⚠️ Early-warning fired (2026-07-13 ~17:0xZ): the Fusion-io device just dropped again (minio drives-online: 0, see #504) while SERVER reports Ready. Last time this exact sequence preceded the node…

spikerj commented on issue spikerj/spikersoft-issues#504 2026-07-13 16:19:08 +00:00
[Bug][Infra] MinIO doesn't recover after a transient /mnt/fusionio drive blip — stuck '0 drives provided', 'mc ready local' healthcheck misses it, needs manual restart

Drive-drop recurrence (2026-07-13 ~17:0xZ): minio exited clean, restart:any respawned it, and the new task loops 'Read/Write quorum could not be established... drives-online: 0' — /mnt/fusionio/mi…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 14:27:23 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Crash #4: ~2026-07-13 06:2x UTC — during the sdxllightning image push (multi-GB blob commit into MinIO on /mnt/fusionio). That makes 3 of tonight's 4 crashes coincide with sustained MinIO ingest.…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 14:26:33 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recurrence #6 (2026-07-13 ~15:1xZ, user-reported): SERVER crashed and this time recovered fully automatically within ~3 min of rejoining — all pinned services (minio/backend/mongo-router/mails…

spikerj commented on issue spikerj/spikersoft-issues#514 2026-07-13 14:20:02 +00:00
[Infra][CI] Gitea Actions leaves GITEA-ACTION* volumes behind → build nodes recurrently fill up (need runner auto-cleanup / scheduled prune)

Second disk-full within 6h (2026-07-13 ~08:03Z→~14:50Z): laptop-server filled again after the earlier ~05:15Z cleanup, killing gitea_postgres (and with it git/API/CI) for ~7h until this morning's…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 14:19:49 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recovery timeline for last night's #5 (closing the loop): SERVER went down ~06:27Z and stayed down ~8h until the morning power-cycle (~14:35Z). Once the node rejoined, recovery was **fully…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 06:28:42 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recurrence #5 (2026-07-13 ~06:27Z): SERVER Down/Unreachable again. Cadence is now firmly periodic: 02:50 → ~04:25 → ~06:27, roughly every 90–100 min tonight. That regularity plus load-correl…

spikerj commented on issue spikerj/spikersoft-issues#537 2026-07-13 06:15:18 +00:00
[ArtPipe][CI] Phase 2 blocked: sdxllightning env image is 98 GB — unfiltered snapshot_download bakes the entire SDXL repo; push dies

Third failure mode confirmed — 98 GB pushes don't just fail, they HANG (2026-07-13 06:15Z). Run 10784 (with the 576b4bd cache sweep but NOT the snapshot_download filter) rebuilt sdxllightning…