⚠️ Early-warning fired (2026-07-13 ~17:0xZ): the Fusion-io device just dropped again (minio drives-online: 0, see #504) while SERVER reports Ready. Last time this exact sequence preceded the node…
Drive-drop recurrence (2026-07-13 ~17:0xZ): minio exited clean, restart:any respawned it, and the new task loops 'Read/Write quorum could not be established... drives-online: 0' — /mnt/fusionio/mi…
Crash #4: ~2026-07-13 06:2x UTC — during the sdxllightning image push (multi-GB blob commit into MinIO on /mnt/fusionio). That makes 3 of tonight's 4 crashes coincide with sustained MinIO ingest.…
Recurrence #6 (2026-07-13 ~15:1xZ, user-reported): SERVER crashed and this time recovered fully automatically within ~3 min of rejoining — all pinned services (minio/backend/mongo-router/mails…
Second disk-full within 6h (2026-07-13 ~08:03Z→~14:50Z): laptop-server filled again after the earlier ~05:15Z cleanup, killing gitea_postgres (and with it git/API/CI) for ~7h until this morning's…
Recovery timeline for last night's #5 (closing the loop): SERVER went down ~06:27Z and stayed down ~8h until the morning power-cycle (~14:35Z). Once the node rejoined, recovery was **fully…
Recurrence #5 (2026-07-13 ~06:27Z): SERVER Down/Unreachable again. Cadence is now firmly periodic: 02:50 → ~04:25 → ~06:27, roughly every 90–100 min tonight. That regularity plus load-correl…
Third failure mode confirmed — 98 GB pushes don't just fail, they HANG (2026-07-13 06:15Z). Run 10784 (with the 576b4bd cache sweep but NOT the snapshot_download filter) rebuilt sdxllightning…
Escalation data point (2026-07-13 ~05:14Z): tonight this bit much harder than runner-volume leftovers — leftover ModelEnvImages layers/build-cache on laptop-server (two killed runs of the 98 GB…
New wedge class found (2026-07-13 ~04:45Z), invisible to both item A and item B: after the 04:2x SERVER hard-down recovery, metadata-extractor, security-monitor and system-remediation each…
Crash #3: ~2026-07-13 03:5x UTC — roughly ONE HOUR after the reboot from crash #2 (03:51 UTC boot). Intervals are shrinking: ~8h → ~1h, again under memory/IO load (a ~20 GB registry push was in…
Recurrence #4 (2026-07-13 ~04:2xZ): HARD DOWN again — docker node ls → SERVER Down/Unreachable (user-confirmed). Same full-outage mode as 07-12 08:29: everything hostname-pinned to SERVER…
Root cause found (2026-07-13 ~03:00Z) — diagnosable at last because the SERVER crash/reboot restored the node's log RPC endpoint:
java.lang.IllegalStateException: data path [/usr/share/el…
New crash data point: SERVER went down again ~2026-07-13 03:0x UTC — second crash in ~9h (uptime was 8h21m at 02:43 UTC, so the previous boot was ~18:22 UTC 2026-07-12). Correlates with tonight's…
Recurrence 2026-07-13 ~02:50Z (user-confirmed crash; node rejoined by ~02:53Z, recovery wave in progress — minio/backend/mongo-router/docker-monitor all restarting through the usual 'swarm…
MinIO down again, new failure mode (2026-07-13 ~02:31Z) — and it takes the container registry with it.
minio_miniotask exited cleanly (state Complete, exit 0) at ~02:31Z and swarm…
Phase 2 fully staged (2026-07-13), merge in this order:
- spikersoft-artpipe PR #14 — extends the ModelEnvImages default set (+blender +sdxllightning +triposr; hunyuan3dpaint stays…