Recurrence of #924, worse this time: gluster FUSE hang on dreamstream4 took the message bus down for ~7.5h (04:29Z → 12:16Z), and during remediation dreamstream3 crashed hard and is still down — Seq is offline until someone power-cycles ds3.
Timeline (2026-08-07, UTC)
~04:28 — gluster FUSE mount /mnt on ds4 hangs (client of staging-gfs). Root cause on the client side: glusterd on ds4 wedged — process "active" but port 24007 accepting no connections (both v4/v6 dead while gluster volume status via unix socket still worked).
04:29:45 — RabbitMQ (data bind /mnt/rabbitmq on ds4) received SIGTERM — this is swarm replacing the task after its healthcheck failed (docker exec healthchecks were failing with setns: exit status 1). Shutdown hung on the dead mount. This also retroactively explains yesterday's "unexplained SIGTERM 13:10Z" in #924.
04:29 → 12:16 — AMQP 5672 refused everywhere. All event handlers crash-looped (SpikerSoft.EventHandlers.ArtPipeProcessor "Failed to create RabbitMQ connection" every ~6s in Seq); Failed to publish system/journal event bursts at 04:34; SpikerSoft.Api RabbitMQ healthcheck red all morning.
11:40–12:14 — investigation + remediation from 4090 (no ssh; one-shot swarm jobs with docker.sock + privileged nsenter):
rabbit container was a zombie: beam.smp defunct, container PID 1 (rabbitmq-server) ignored SIGKILL (SIGKILL pending in ShdPnd, never processed — kernel-level stuck). A second rabbit task container from a previous failed shutdown had been sitting "Up 8 days (unhealthy)" the same way.
docker rm -f failed ("did not receive an exit event"); killing the two containerd-shim processes forced exit events and freed the task slots.
umount -l /mnt + remount attempts: mounts kept dying minutes later (fresh FUSE daemon died silently each time); dockerd also spammed network ymfl2rdjrzbgu6a9i9d31rhkr not found (= the rabbitmq overlay) — docker network state corrupt.
restarted glusterd (fixed volfile serving; all 4 bricks online), but node remained sick (unkillable PIDs, silent FUSE deaths).
12:14 — rebooted ds4 (systemd-run via nsenter). Swarm leader failed over 4090→ds7 cleanly.
12:16 — rabbit clean boot on ds4, 6 plugins, 49 consumers reconnected, no mnesia corruption (no #895 pattern). Event-handler fleet recovered.
~12:18 — dreamstream3 went down hard (no ping on 192.168.0.107). Not obviously caused by anything above — no workloads were changed on ds3 except a few failed image-pull attempts. ds3 also had the gluster hang yesterday; hardware/kernel suspect.
Current state
✅ RabbitMQ up on ds4, consumers healthy, ds4 rebooted clean (zombies gone, /mnt mounted via fstab).
❌dreamstream3 DOWN — needs physical power-cycle. Until then:
Seq is offline (pinned node.hostname == dreamstream3, data /mnt/seq/data). Platform-wide log ingestion is 404 — logs since ~12:18Z are being dropped. Do not trust Seq queries over the 04:29–12:16 window either (ingestion was up but half the fleet was down).
openbao-2 down (bao raft 2/3 — quorum OK, bao.spikersoft.com healthy).
redis-node-3 down (cluster 5/6), jetson-influx-3 down.
Gluster staging-gfs: brick 192.168.0.107 (ds3) offline, other 3 bricks fine — replica volume degraded but serving.
After Seq is back, check gpu-coordinator / job queues for work stuck from the outage window (#895 "stuck-Complete kick list" pattern).
This is the 2nd gluster client hang on ds4 in 2 days and ds3 crashed outright — both are aarch64 SBCs that also took the 07-29 power-outage corruption. Consider: hardware health check (SD/eMMC, PSU), kernel/glusterfs updates, or moving rabbit+seq storage/pinning to sturdier nodes.
mount localhost:/staging-gfs depends on local glusterd being healthy; fstab could list peer backup volfile servers (e.g. backupvolfile-server=192.168.0.105:192.168.0.106) so a wedged local glusterd doesn't block remounts.
Alert emails are still dead (#756 SMTP 535) — nobody got alerted for a 7.5h bus outage.
Notes for future runbook use (worked from 4090, no ssh)
Probe/fix pattern: docker service create --restart-condition none --constraint node.hostname==<node> --mount type=bind,src=/var/run/docker.sock,dst=/var/run/docker.sock docker:cli sh -c '…', then docker run --rm --privileged --pid=host alpine nsenter -t 1 -m … for host-level work; systemd-run --on-active=5 systemctl reboot to reboot safely from inside a container that dies with the node.
Unkillable container ("did not receive an exit event"): kill its containerd-shim → containerd emits the exit event and swarm reschedules.
When running mount /mnt via nsenter from an alpine container, set PATH=/usr/sbin:/usr/bin:/sbin:/bin or mount.glusterfs fails bogusly.
ds nodes cannot pull from Docker Hub (rabbit image exists only on ds4 + 4090; datalust/seq only on ds3 + 4090) — this blocked temporarily moving rabbit to ds3. Consider mirroring critical infra images into git.spikersoft.com registry.
## Summary
Recurrence of #924, worse this time: gluster FUSE hang on dreamstream4 took the message bus down for ~7.5h (04:29Z → 12:16Z), and during remediation **dreamstream3 crashed hard and is still down — Seq is offline until someone power-cycles ds3.**
## Timeline (2026-08-07, UTC)
- **~04:28** — gluster FUSE mount `/mnt` on ds4 hangs (client of `staging-gfs`). Root cause on the client side: **glusterd on ds4 wedged — process "active" but port 24007 accepting no connections** (both v4/v6 dead while `gluster volume status` via unix socket still worked).
- **04:29:45** — RabbitMQ (data bind `/mnt/rabbitmq` on ds4) received SIGTERM — this is swarm replacing the task after its healthcheck failed (docker exec healthchecks were failing with `setns: exit status 1`). Shutdown hung on the dead mount. This also retroactively explains yesterday's "unexplained SIGTERM 13:10Z" in #924.
- **04:29 → 12:16** — AMQP 5672 refused everywhere. All event handlers crash-looped (`SpikerSoft.EventHandlers.ArtPipeProcessor` "Failed to create RabbitMQ connection" every ~6s in Seq); `Failed to publish system/journal event` bursts at 04:34; SpikerSoft.Api RabbitMQ healthcheck red all morning.
- **11:40–12:14** — investigation + remediation from 4090 (no ssh; one-shot swarm jobs with docker.sock + privileged nsenter):
- rabbit container was a zombie: `beam.smp` defunct, container PID 1 (`rabbitmq-server`) **ignored SIGKILL** (SIGKILL pending in ShdPnd, never processed — kernel-level stuck). A second rabbit task container from a *previous* failed shutdown had been sitting "Up 8 days (unhealthy)" the same way.
- `docker rm -f` failed ("did not receive an exit event"); killing the two `containerd-shim` processes forced exit events and freed the task slots.
- `umount -l /mnt` + remount attempts: mounts kept dying minutes later (fresh FUSE daemon died silently each time); dockerd also spammed `network ymfl2rdjrzbgu6a9i9d31rhkr not found` (= the rabbitmq overlay) — docker network state corrupt.
- restarted glusterd (fixed volfile serving; all 4 bricks online), but node remained sick (unkillable PIDs, silent FUSE deaths).
- **12:14** — **rebooted ds4** (systemd-run via nsenter). Swarm leader failed over 4090→ds7 cleanly.
- **12:16** — rabbit clean boot on ds4, 6 plugins, **49 consumers reconnected, no mnesia corruption** (no #895 pattern). Event-handler fleet recovered.
- **~12:18** — **dreamstream3 went down hard** (no ping on 192.168.0.107). Not obviously caused by anything above — no workloads were changed on ds3 except a few failed image-pull attempts. ds3 also had the gluster hang yesterday; hardware/kernel suspect.
## Current state
- ✅ RabbitMQ up on ds4, consumers healthy, ds4 rebooted clean (zombies gone, /mnt mounted via fstab).
- ❌ **dreamstream3 DOWN — needs physical power-cycle.** Until then:
- **Seq is offline** (pinned `node.hostname == dreamstream3`, data `/mnt/seq/data`). Platform-wide log ingestion is 404 — logs since ~12:18Z are being dropped. Do not trust Seq queries over the 04:29–12:16 window either (ingestion was up but half the fleet was down).
- openbao-2 down (bao raft 2/3 — quorum OK, bao.spikersoft.com healthy).
- redis-node-3 down (cluster 5/6), jetson-influx-3 down.
- Gluster `staging-gfs`: brick 192.168.0.107 (ds3) offline, other 3 bricks fine — replica volume degraded but serving.
## Follow-ups
1. **Power-cycle ds3** (only human-actionable item). After boot: verify `/mnt`, Seq login + ingestion, openbao-2 unsealed/rejoined, redis-node-3 rejoined, brick heals (`gluster volume heal staging-gfs info`).
2. After Seq is back, check gpu-coordinator / job queues for work stuck from the outage window (#895 "stuck-Complete kick list" pattern).
3. This is the 2nd gluster client hang on ds4 in 2 days and ds3 crashed outright — both are aarch64 SBCs that also took the 07-29 power-outage corruption. Consider: hardware health check (SD/eMMC, PSU), kernel/glusterfs updates, or moving rabbit+seq storage/pinning to sturdier nodes.
4. `mount localhost:/staging-gfs` depends on local glusterd being healthy; fstab could list peer backup volfile servers (e.g. `backupvolfile-server=192.168.0.105:192.168.0.106`) so a wedged local glusterd doesn't block remounts.
5. Alert emails are still dead (#756 SMTP 535) — nobody got alerted for a 7.5h bus outage.
## Notes for future runbook use (worked from 4090, no ssh)
- Probe/fix pattern: `docker service create --restart-condition none --constraint node.hostname==<node> --mount type=bind,src=/var/run/docker.sock,dst=/var/run/docker.sock docker:cli sh -c '…'`, then `docker run --rm --privileged --pid=host alpine nsenter -t 1 -m …` for host-level work; `systemd-run --on-active=5 systemctl reboot` to reboot safely from inside a container that dies with the node.
- Unkillable container ("did not receive an exit event"): kill its `containerd-shim` → containerd emits the exit event and swarm reschedules.
- When running `mount /mnt` via nsenter from an alpine container, set `PATH=/usr/sbin:/usr/bin:/sbin:/bin` or `mount.glusterfs` fails bogusly.
- ds nodes cannot pull from Docker Hub (rabbit image exists only on ds4 + 4090; datalust/seq only on ds3 + 4090) — this blocked temporarily moving rabbit to ds3. Consider mirroring critical infra images into git.spikersoft.com registry.
Update 12:35Z — everything recovered except one manual step.
Correction: ds3 did NOT crash. Uptime is 33 days — it dropped off the network ~12:18–12:30Z (no ping) and came back on its own. So the real ds3 issue is a flaky NIC/link/switch port, not a kernel panic. This probably also explains the earlier "can't pull from Docker Hub" failures on ds3 (alpine pulled fine there after the link returned).
Seq is back: UI 200, query API 200, ingestion accepting (201). Log blackout window: ~12:18–12:31Z (plus degraded trust 04:29–12:16Z while the bus was down).
redis-node-3 and jetson-influx-3 rescheduled and running; gluster brick .107 back (heal should catch up).
openbao-2 restarted → will be sealed (raft retry-join in progress; cluster healthy on 2/3). Needs a manual bao operator unseal on openbao-2 when convenient.
RabbitMQ stable on ds4 since 12:16Z, 51+ consumers, no mnesia damage.
Remaining follow-ups from the description: check ds3's link/NIC health, gpu-coordinator stuck-job sweep, image mirroring, fstab backup volfile servers, and #756 (mail alerts still dead).
**Update 12:35Z — everything recovered except one manual step.**
- **Correction: ds3 did NOT crash.** Uptime is 33 days — it dropped off the network ~12:18–12:30Z (no ping) and came back on its own. So the real ds3 issue is a flaky NIC/link/switch port, not a kernel panic. This probably also explains the earlier "can't pull from Docker Hub" failures on ds3 (alpine pulled fine there after the link returned).
- **Seq is back**: UI 200, query API 200, ingestion accepting (201). Log blackout window: ~12:18–12:31Z (plus degraded trust 04:29–12:16Z while the bus was down).
- redis-node-3 and jetson-influx-3 rescheduled and running; gluster brick .107 back (heal should catch up).
- **openbao-2 restarted → will be sealed** (raft retry-join in progress; cluster healthy on 2/3). Needs a manual `bao operator unseal` on openbao-2 when convenient.
- RabbitMQ stable on ds4 since 12:16Z, 51+ consumers, no mnesia damage.
Remaining follow-ups from the description: check ds3's link/NIC health, gpu-coordinator stuck-job sweep, image mirroring, fstab backup volfile servers, and #756 (mail alerts still dead).
Verified 2026-08-07 — Code:spikersoft-infrastructure@86d03ff — no commit records any of this incident's follow-ups. No backupvolfile-server anywhere in the repo, no gluster/fstab runbook, and seq + rabbitmq are still hostname-pinned to the two SBCs that hung. Live:The outage is RESOLVED.docker node ls → all 10 nodes Ready/Active, dreamstream3 back at 192.168.0.107 (it was hard-down when this was filed). docker service ps seq_seq → Running on dreamstream3; openbao_openbao-2, redis-cluster_redis-node-3 and jetson-influxdb_jetson-influx-3 are all Running again. curl https://api.spikersoft.com/healthz → Healthy, RabbitMQ check reports 267/267 queues running and 78 consumers. docker node inspect dreamstream4 → Status.Addr 192.168.0.108 (the 0.0.0.0 from #925 also cleared on the reboot), and ds4 is no longer leader.
Status: the outage itself is resolved and verified live; follow-ups 2-5 (queue sweep, SBC hardware pass, fstab backup volfile servers, infra-image mirroring) are all unstarted
Closing here. Work now lives in the repo that holds the fix, so fixes #<N> in a PR will
auto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Migrated to **spikerj/spikersoft-infrastructure#177** as part of the umbrella-tracker breakup.
Verified 2026-08-07 — **Code:** `spikersoft-infrastructure@86d03ff` — no commit records any of this incident's follow-ups. No `backupvolfile-server` anywhere in the repo, no gluster/fstab runbook, and `seq` + `rabbitmq` are still hostname-pinned to the two SBCs that hung. **Live:** **The outage is RESOLVED.** `docker node ls` → all 10 nodes Ready/Active, **dreamstream3 back at 192.168.0.107** (it was hard-down when this was filed). `docker service ps seq_seq` → Running on dreamstream3; `openbao_openbao-2`, `redis-cluster_redis-node-3` and `jetson-influxdb_jetson-influx-3` are all Running again. `curl https://api.spikersoft.com/healthz` → **Healthy**, RabbitMQ check reports 267/267 queues running and 78 consumers. `docker node inspect dreamstream4` → `Status.Addr 192.168.0.108` (the 0.0.0.0 from #925 also cleared on the reboot), and ds4 is no longer leader.
Status: the outage itself is resolved and verified live; follow-ups 2-5 (queue sweep, SBC hardware pass, fstab backup volfile servers, infra-image mirroring) are all unstarted
Closing here. Work now lives in the repo that holds the fix, so `fixes #<N>` in a PR will
auto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Summary
Recurrence of #924, worse this time: gluster FUSE hang on dreamstream4 took the message bus down for ~7.5h (04:29Z → 12:16Z), and during remediation dreamstream3 crashed hard and is still down — Seq is offline until someone power-cycles ds3.
Timeline (2026-08-07, UTC)
/mnton ds4 hangs (client ofstaging-gfs). Root cause on the client side: glusterd on ds4 wedged — process "active" but port 24007 accepting no connections (both v4/v6 dead whilegluster volume statusvia unix socket still worked)./mnt/rabbitmqon ds4) received SIGTERM — this is swarm replacing the task after its healthcheck failed (docker exec healthchecks were failing withsetns: exit status 1). Shutdown hung on the dead mount. This also retroactively explains yesterday's "unexplained SIGTERM 13:10Z" in #924.SpikerSoft.EventHandlers.ArtPipeProcessor"Failed to create RabbitMQ connection" every ~6s in Seq);Failed to publish system/journal eventbursts at 04:34; SpikerSoft.Api RabbitMQ healthcheck red all morning.beam.smpdefunct, container PID 1 (rabbitmq-server) ignored SIGKILL (SIGKILL pending in ShdPnd, never processed — kernel-level stuck). A second rabbit task container from a previous failed shutdown had been sitting "Up 8 days (unhealthy)" the same way.docker rm -ffailed ("did not receive an exit event"); killing the twocontainerd-shimprocesses forced exit events and freed the task slots.umount -l /mnt+ remount attempts: mounts kept dying minutes later (fresh FUSE daemon died silently each time); dockerd also spammednetwork ymfl2rdjrzbgu6a9i9d31rhkr not found(= the rabbitmq overlay) — docker network state corrupt.Current state
node.hostname == dreamstream3, data/mnt/seq/data). Platform-wide log ingestion is 404 — logs since ~12:18Z are being dropped. Do not trust Seq queries over the 04:29–12:16 window either (ingestion was up but half the fleet was down).staging-gfs: brick 192.168.0.107 (ds3) offline, other 3 bricks fine — replica volume degraded but serving.Follow-ups
/mnt, Seq login + ingestion, openbao-2 unsealed/rejoined, redis-node-3 rejoined, brick heals (gluster volume heal staging-gfs info).mount localhost:/staging-gfsdepends on local glusterd being healthy; fstab could list peer backup volfile servers (e.g.backupvolfile-server=192.168.0.105:192.168.0.106) so a wedged local glusterd doesn't block remounts.Notes for future runbook use (worked from 4090, no ssh)
docker service create --restart-condition none --constraint node.hostname==<node> --mount type=bind,src=/var/run/docker.sock,dst=/var/run/docker.sock docker:cli sh -c '…', thendocker run --rm --privileged --pid=host alpine nsenter -t 1 -m …for host-level work;systemd-run --on-active=5 systemctl rebootto reboot safely from inside a container that dies with the node.containerd-shim→ containerd emits the exit event and swarm reschedules.mount /mntvia nsenter from an alpine container, setPATH=/usr/sbin:/usr/bin:/sbin:/binormount.glusterfsfails bogusly.Update 12:35Z — everything recovered except one manual step.
bao operator unsealon openbao-2 when convenient.Remaining follow-ups from the description: check ds3's link/NIC health, gpu-coordinator stuck-job sweep, image mirroring, fstab backup volfile servers, and #756 (mail alerts still dead).
Migrated to spikerj/spikersoft-infrastructure#177 as part of the umbrella-tracker breakup.
Verified 2026-08-07 — Code:
spikersoft-infrastructure@86d03ff— no commit records any of this incident's follow-ups. Nobackupvolfile-serveranywhere in the repo, no gluster/fstab runbook, andseq+rabbitmqare still hostname-pinned to the two SBCs that hung. Live: The outage is RESOLVED.docker node ls→ all 10 nodes Ready/Active, dreamstream3 back at 192.168.0.107 (it was hard-down when this was filed).docker service ps seq_seq→ Running on dreamstream3;openbao_openbao-2,redis-cluster_redis-node-3andjetson-influxdb_jetson-influx-3are all Running again.curl https://api.spikersoft.com/healthz→ Healthy, RabbitMQ check reports 267/267 queues running and 78 consumers.docker node inspect dreamstream4→Status.Addr 192.168.0.108(the 0.0.0.0 from #925 also cleared on the reboot), and ds4 is no longer leader.Status: the outage itself is resolved and verified live; follow-ups 2-5 (queue sweep, SBC hardware pass, fstab backup volfile servers, infra-image mirroring) are all unstarted
Closing here. Work now lives in the repo that holds the fix, so
fixes #<N>in a PR willauto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent