Joseph Spiker spikerj
  • Joined on 2025-11-03
spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 16:19:09 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

⚠️ Early-warning fired (2026-07-13 ~17:0xZ): the Fusion-io device just dropped again (minio drives-online: 0, see #504) while SERVER reports Ready. Last time this exact sequence preceded the node…

spikerj commented on issue spikerj/spikersoft-issues#504 2026-07-13 16:19:08 +00:00
[Bug][Infra] MinIO doesn't recover after a transient /mnt/fusionio drive blip — stuck '0 drives provided', 'mc ready local' healthcheck misses it, needs manual restart

Drive-drop recurrence (2026-07-13 ~17:0xZ): minio exited clean, restart:any respawned it, and the new task loops 'Read/Write quorum could not be established... drives-online: 0' — /mnt/fusionio/mi…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 14:27:23 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Crash #4: ~2026-07-13 06:2x UTC — during the sdxllightning image push (multi-GB blob commit into MinIO on /mnt/fusionio). That makes 3 of tonight's 4 crashes coincide with sustained MinIO ingest.…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 14:26:33 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recurrence #6 (2026-07-13 ~15:1xZ, user-reported): SERVER crashed and this time recovered fully automatically within ~3 min of rejoining — all pinned services (minio/backend/mongo-router/mails…

spikerj commented on issue spikerj/spikersoft-issues#514 2026-07-13 14:20:02 +00:00
[Infra][CI] Gitea Actions leaves GITEA-ACTION* volumes behind → build nodes recurrently fill up (need runner auto-cleanup / scheduled prune)

Second disk-full within 6h (2026-07-13 ~08:03Z→~14:50Z): laptop-server filled again after the earlier ~05:15Z cleanup, killing gitea_postgres (and with it git/API/CI) for ~7h until this morning's…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 14:19:49 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recovery timeline for last night's #5 (closing the loop): SERVER went down ~06:27Z and stayed down ~8h until the morning power-cycle (~14:35Z). Once the node rejoined, recovery was **fully…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 06:28:42 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recurrence #5 (2026-07-13 ~06:27Z): SERVER Down/Unreachable again. Cadence is now firmly periodic: 02:50 → ~04:25 → ~06:27, roughly every 90–100 min tonight. That regularity plus load-correl…

spikerj commented on issue spikerj/spikersoft-issues#537 2026-07-13 06:15:18 +00:00
[ArtPipe][CI] Phase 2 blocked: sdxllightning env image is 98 GB — unfiltered snapshot_download bakes the entire SDXL repo; push dies

Third failure mode confirmed — 98 GB pushes don't just fail, they HANG (2026-07-13 06:15Z). Run 10784 (with the 576b4bd cache sweep but NOT the snapshot_download filter) rebuilt sdxllightning…

spikerj opened issue spikerj/spikersoft-issues#538 2026-07-13 05:27:37 +00:00
[Infra][Network] LAN traffic to *.spikersoft.com NAT-hairpins through the UDM — local DNS records + direct Gitea→MinIO path
spikerj commented on issue spikerj/spikersoft-issues#514 2026-07-13 05:22:21 +00:00
[Infra][CI] Gitea Actions leaves GITEA-ACTION* volumes behind → build nodes recurrently fill up (need runner auto-cleanup / scheduled prune)

Escalation data point (2026-07-13 ~05:14Z): tonight this bit much harder than runner-volume leftovers — leftover ModelEnvImages layers/build-cache on laptop-server (two killed runs of the 98 GB…

spikerj opened issue spikerj/spikersoft-issues#537 2026-07-13 05:21:50 +00:00
[ArtPipe][CI] Phase 2 blocked: sdxllightning env image is 98 GB — unfiltered snapshot_download bakes the entire SDXL repo; push dies
spikerj commented on issue spikerj/spikersoft-issues#510 2026-07-13 04:45:04 +00:00
[Exploration][Infra] Resilience to SERVER node flaps — restart-policy windows, auto-reconcile sweep, de-SPOF the SERVER-pinned stateful tier

New wedge class found (2026-07-13 ~04:45Z), invisible to both item A and item B: after the 04:2x SERVER hard-down recovery, metadata-extractor, security-monitor and system-remediation each…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 04:29:59 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Crash #3: ~2026-07-13 03:5x UTC — roughly ONE HOUR after the reboot from crash #2 (03:51 UTC boot). Intervals are shrinking: ~8h → ~1h, again under memory/IO load (a ~20 GB registry push was in…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 04:28:41 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recurrence #4 (2026-07-13 ~04:2xZ): HARD DOWN againdocker node ls → SERVER Down/Unreachable (user-confirmed). Same full-outage mode as 07-12 08:29: everything hostname-pinned to SERVER…

spikerj commented on issue spikerj/spikersoft-issues#481 2026-07-13 02:54:50 +00:00
[Bug][Infra] jaeger_elasticsearch down ~24h on SERVER (exit 1) — MaxAttempts exhausted, tracing backend offline

Root cause found (2026-07-13 ~03:00Z) — diagnosable at last because the SERVER crash/reboot restored the node's log RPC endpoint:

java.lang.IllegalStateException: data path [/usr/share/el…
spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 02:53:02 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

New crash data point: SERVER went down again ~2026-07-13 03:0x UTC — second crash in ~9h (uptime was 8h21m at 02:43 UTC, so the previous boot was ~18:22 UTC 2026-07-12). Correlates with tonight's…

spikerj commented on issue spikerj/spikersoft-issues#503 2026-07-13 02:52:00 +00:00
[Bug][Infra][CRITICAL] SERVER node crash-looping — kernel 'BUG: Bad page state in process dockerd' (memory corruption / likely bad RAM); flaps all SERVER-pinned services

Recurrence 2026-07-13 ~02:50Z (user-confirmed crash; node rejoined by ~02:53Z, recovery wave in progress — minio/backend/mongo-router/docker-monitor all restarting through the usual 'swarm…

spikerj commented on issue spikerj/spikersoft-issues#504 2026-07-13 02:37:18 +00:00
[Bug][Infra] MinIO doesn't recover after a transient /mnt/fusionio drive blip — stuck '0 drives provided', 'mc ready local' healthcheck misses it, needs manual restart

MinIO down again, new failure mode (2026-07-13 ~02:31Z) — and it takes the container registry with it.

  • minio_minio task exited cleanly (state Complete, exit 0) at ~02:31Z and swarm…
spikerj opened issue spikerj/spikersoft-issues#536 2026-07-13 02:25:04 +00:00
[CI][ArtPipe] ModelEnvImages has no concurrency guard — concurrent runs race the same :latest tags (observed live)
spikerj commented on issue spikerj/spikersoft-issues#518 2026-07-13 02:02:43 +00:00
[ArtPipe] Baked images Phase 2: stood-down residents blender/sdxl/triposr, hunyuan deferred (epic #515)

Phase 2 fully staged (2026-07-13), merge in this order:

  1. spikersoft-artpipe PR #14 — extends the ModelEnvImages default set (+blender +sdxllightning +triposr; hunyuan3dpaint stays…