OUTAGE: gluster hang on ds4 → 7.5h RabbitMQ/message-bus outage (recurrence of #924); ds3 crashed and is DOWN — Seq offline, needs power-cycle #1000

Closed
opened 2026-08-07 12:23:29 +00:00 by spikerj · 2 comments
Owner

Summary

Recurrence of #924, worse this time: gluster FUSE hang on dreamstream4 took the message bus down for ~7.5h (04:29Z → 12:16Z), and during remediation dreamstream3 crashed hard and is still down — Seq is offline until someone power-cycles ds3.

Timeline (2026-08-07, UTC)

  • ~04:28 — gluster FUSE mount /mnt on ds4 hangs (client of staging-gfs). Root cause on the client side: glusterd on ds4 wedged — process "active" but port 24007 accepting no connections (both v4/v6 dead while gluster volume status via unix socket still worked).
  • 04:29:45 — RabbitMQ (data bind /mnt/rabbitmq on ds4) received SIGTERM — this is swarm replacing the task after its healthcheck failed (docker exec healthchecks were failing with setns: exit status 1). Shutdown hung on the dead mount. This also retroactively explains yesterday's "unexplained SIGTERM 13:10Z" in #924.
  • 04:29 → 12:16 — AMQP 5672 refused everywhere. All event handlers crash-looped (SpikerSoft.EventHandlers.ArtPipeProcessor "Failed to create RabbitMQ connection" every ~6s in Seq); Failed to publish system/journal event bursts at 04:34; SpikerSoft.Api RabbitMQ healthcheck red all morning.
  • 11:40–12:14 — investigation + remediation from 4090 (no ssh; one-shot swarm jobs with docker.sock + privileged nsenter):
    • rabbit container was a zombie: beam.smp defunct, container PID 1 (rabbitmq-server) ignored SIGKILL (SIGKILL pending in ShdPnd, never processed — kernel-level stuck). A second rabbit task container from a previous failed shutdown had been sitting "Up 8 days (unhealthy)" the same way.
    • docker rm -f failed ("did not receive an exit event"); killing the two containerd-shim processes forced exit events and freed the task slots.
    • umount -l /mnt + remount attempts: mounts kept dying minutes later (fresh FUSE daemon died silently each time); dockerd also spammed network ymfl2rdjrzbgu6a9i9d31rhkr not found (= the rabbitmq overlay) — docker network state corrupt.
    • restarted glusterd (fixed volfile serving; all 4 bricks online), but node remained sick (unkillable PIDs, silent FUSE deaths).
  • 12:14 — rebooted ds4 (systemd-run via nsenter). Swarm leader failed over 4090→ds7 cleanly.
  • 12:16 — rabbit clean boot on ds4, 6 plugins, 49 consumers reconnected, no mnesia corruption (no #895 pattern). Event-handler fleet recovered.
  • ~12:18 — dreamstream3 went down hard (no ping on 192.168.0.107). Not obviously caused by anything above — no workloads were changed on ds3 except a few failed image-pull attempts. ds3 also had the gluster hang yesterday; hardware/kernel suspect.

Current state

  • ✅ RabbitMQ up on ds4, consumers healthy, ds4 rebooted clean (zombies gone, /mnt mounted via fstab).
  • ❌ dreamstream3 DOWN — needs physical power-cycle. Until then:
    • Seq is offline (pinned node.hostname == dreamstream3, data /mnt/seq/data). Platform-wide log ingestion is 404 — logs since ~12:18Z are being dropped. Do not trust Seq queries over the 04:29–12:16 window either (ingestion was up but half the fleet was down).
    • openbao-2 down (bao raft 2/3 — quorum OK, bao.spikersoft.com healthy).
    • redis-node-3 down (cluster 5/6), jetson-influx-3 down.
  • Gluster staging-gfs: brick 192.168.0.107 (ds3) offline, other 3 bricks fine — replica volume degraded but serving.

Follow-ups

  1. Power-cycle ds3 (only human-actionable item). After boot: verify /mnt, Seq login + ingestion, openbao-2 unsealed/rejoined, redis-node-3 rejoined, brick heals (gluster volume heal staging-gfs info).
  2. After Seq is back, check gpu-coordinator / job queues for work stuck from the outage window (#895 "stuck-Complete kick list" pattern).
  3. This is the 2nd gluster client hang on ds4 in 2 days and ds3 crashed outright — both are aarch64 SBCs that also took the 07-29 power-outage corruption. Consider: hardware health check (SD/eMMC, PSU), kernel/glusterfs updates, or moving rabbit+seq storage/pinning to sturdier nodes.
  4. mount localhost:/staging-gfs depends on local glusterd being healthy; fstab could list peer backup volfile servers (e.g. backupvolfile-server=192.168.0.105:192.168.0.106) so a wedged local glusterd doesn't block remounts.
  5. Alert emails are still dead (#756 SMTP 535) — nobody got alerted for a 7.5h bus outage.

Notes for future runbook use (worked from 4090, no ssh)

  • Probe/fix pattern: docker service create --restart-condition none --constraint node.hostname==<node> --mount type=bind,src=/var/run/docker.sock,dst=/var/run/docker.sock docker:cli sh -c '…', then docker run --rm --privileged --pid=host alpine nsenter -t 1 -m … for host-level work; systemd-run --on-active=5 systemctl reboot to reboot safely from inside a container that dies with the node.
  • Unkillable container ("did not receive an exit event"): kill its containerd-shim → containerd emits the exit event and swarm reschedules.
  • When running mount /mnt via nsenter from an alpine container, set PATH=/usr/sbin:/usr/bin:/sbin:/bin or mount.glusterfs fails bogusly.
  • ds nodes cannot pull from Docker Hub (rabbit image exists only on ds4 + 4090; datalust/seq only on ds3 + 4090) — this blocked temporarily moving rabbit to ds3. Consider mirroring critical infra images into git.spikersoft.com registry.
## Summary Recurrence of #924, worse this time: gluster FUSE hang on dreamstream4 took the message bus down for ~7.5h (04:29Z → 12:16Z), and during remediation **dreamstream3 crashed hard and is still down — Seq is offline until someone power-cycles ds3.** ## Timeline (2026-08-07, UTC) - **~04:28** — gluster FUSE mount `/mnt` on ds4 hangs (client of `staging-gfs`). Root cause on the client side: **glusterd on ds4 wedged — process "active" but port 24007 accepting no connections** (both v4/v6 dead while `gluster volume status` via unix socket still worked). - **04:29:45** — RabbitMQ (data bind `/mnt/rabbitmq` on ds4) received SIGTERM — this is swarm replacing the task after its healthcheck failed (docker exec healthchecks were failing with `setns: exit status 1`). Shutdown hung on the dead mount. This also retroactively explains yesterday's "unexplained SIGTERM 13:10Z" in #924. - **04:29 → 12:16** — AMQP 5672 refused everywhere. All event handlers crash-looped (`SpikerSoft.EventHandlers.ArtPipeProcessor` "Failed to create RabbitMQ connection" every ~6s in Seq); `Failed to publish system/journal event` bursts at 04:34; SpikerSoft.Api RabbitMQ healthcheck red all morning. - **11:40–12:14** — investigation + remediation from 4090 (no ssh; one-shot swarm jobs with docker.sock + privileged nsenter): - rabbit container was a zombie: `beam.smp` defunct, container PID 1 (`rabbitmq-server`) **ignored SIGKILL** (SIGKILL pending in ShdPnd, never processed — kernel-level stuck). A second rabbit task container from a *previous* failed shutdown had been sitting "Up 8 days (unhealthy)" the same way. - `docker rm -f` failed ("did not receive an exit event"); killing the two `containerd-shim` processes forced exit events and freed the task slots. - `umount -l /mnt` + remount attempts: mounts kept dying minutes later (fresh FUSE daemon died silently each time); dockerd also spammed `network ymfl2rdjrzbgu6a9i9d31rhkr not found` (= the rabbitmq overlay) — docker network state corrupt. - restarted glusterd (fixed volfile serving; all 4 bricks online), but node remained sick (unkillable PIDs, silent FUSE deaths). - **12:14** — **rebooted ds4** (systemd-run via nsenter). Swarm leader failed over 4090→ds7 cleanly. - **12:16** — rabbit clean boot on ds4, 6 plugins, **49 consumers reconnected, no mnesia corruption** (no #895 pattern). Event-handler fleet recovered. - **~12:18** — **dreamstream3 went down hard** (no ping on 192.168.0.107). Not obviously caused by anything above — no workloads were changed on ds3 except a few failed image-pull attempts. ds3 also had the gluster hang yesterday; hardware/kernel suspect. ## Current state - ✅ RabbitMQ up on ds4, consumers healthy, ds4 rebooted clean (zombies gone, /mnt mounted via fstab). - ❌ **dreamstream3 DOWN — needs physical power-cycle.** Until then: - **Seq is offline** (pinned `node.hostname == dreamstream3`, data `/mnt/seq/data`). Platform-wide log ingestion is 404 — logs since ~12:18Z are being dropped. Do not trust Seq queries over the 04:29–12:16 window either (ingestion was up but half the fleet was down). - openbao-2 down (bao raft 2/3 — quorum OK, bao.spikersoft.com healthy). - redis-node-3 down (cluster 5/6), jetson-influx-3 down. - Gluster `staging-gfs`: brick 192.168.0.107 (ds3) offline, other 3 bricks fine — replica volume degraded but serving. ## Follow-ups 1. **Power-cycle ds3** (only human-actionable item). After boot: verify `/mnt`, Seq login + ingestion, openbao-2 unsealed/rejoined, redis-node-3 rejoined, brick heals (`gluster volume heal staging-gfs info`). 2. After Seq is back, check gpu-coordinator / job queues for work stuck from the outage window (#895 "stuck-Complete kick list" pattern). 3. This is the 2nd gluster client hang on ds4 in 2 days and ds3 crashed outright — both are aarch64 SBCs that also took the 07-29 power-outage corruption. Consider: hardware health check (SD/eMMC, PSU), kernel/glusterfs updates, or moving rabbit+seq storage/pinning to sturdier nodes. 4. `mount localhost:/staging-gfs` depends on local glusterd being healthy; fstab could list peer backup volfile servers (e.g. `backupvolfile-server=192.168.0.105:192.168.0.106`) so a wedged local glusterd doesn't block remounts. 5. Alert emails are still dead (#756 SMTP 535) — nobody got alerted for a 7.5h bus outage. ## Notes for future runbook use (worked from 4090, no ssh) - Probe/fix pattern: `docker service create --restart-condition none --constraint node.hostname==<node> --mount type=bind,src=/var/run/docker.sock,dst=/var/run/docker.sock docker:cli sh -c '…'`, then `docker run --rm --privileged --pid=host alpine nsenter -t 1 -m …` for host-level work; `systemd-run --on-active=5 systemctl reboot` to reboot safely from inside a container that dies with the node. - Unkillable container ("did not receive an exit event"): kill its `containerd-shim` → containerd emits the exit event and swarm reschedules. - When running `mount /mnt` via nsenter from an alpine container, set `PATH=/usr/sbin:/usr/bin:/sbin:/bin` or `mount.glusterfs` fails bogusly. - ds nodes cannot pull from Docker Hub (rabbit image exists only on ds4 + 4090; datalust/seq only on ds3 + 4090) — this blocked temporarily moving rabbit to ds3. Consider mirroring critical infra images into git.spikersoft.com registry.
Author
Owner

Update 12:35Z — everything recovered except one manual step.

  • Correction: ds3 did NOT crash. Uptime is 33 days — it dropped off the network ~12:18–12:30Z (no ping) and came back on its own. So the real ds3 issue is a flaky NIC/link/switch port, not a kernel panic. This probably also explains the earlier "can't pull from Docker Hub" failures on ds3 (alpine pulled fine there after the link returned).
  • Seq is back: UI 200, query API 200, ingestion accepting (201). Log blackout window: ~12:18–12:31Z (plus degraded trust 04:29–12:16Z while the bus was down).
  • redis-node-3 and jetson-influx-3 rescheduled and running; gluster brick .107 back (heal should catch up).
  • openbao-2 restarted → will be sealed (raft retry-join in progress; cluster healthy on 2/3). Needs a manual bao operator unseal on openbao-2 when convenient.
  • RabbitMQ stable on ds4 since 12:16Z, 51+ consumers, no mnesia damage.

Remaining follow-ups from the description: check ds3's link/NIC health, gpu-coordinator stuck-job sweep, image mirroring, fstab backup volfile servers, and #756 (mail alerts still dead).

**Update 12:35Z — everything recovered except one manual step.** - **Correction: ds3 did NOT crash.** Uptime is 33 days — it dropped off the network ~12:18–12:30Z (no ping) and came back on its own. So the real ds3 issue is a flaky NIC/link/switch port, not a kernel panic. This probably also explains the earlier "can't pull from Docker Hub" failures on ds3 (alpine pulled fine there after the link returned). - **Seq is back**: UI 200, query API 200, ingestion accepting (201). Log blackout window: ~12:18–12:31Z (plus degraded trust 04:29–12:16Z while the bus was down). - redis-node-3 and jetson-influx-3 rescheduled and running; gluster brick .107 back (heal should catch up). - **openbao-2 restarted → will be sealed** (raft retry-join in progress; cluster healthy on 2/3). Needs a manual `bao operator unseal` on openbao-2 when convenient. - RabbitMQ stable on ds4 since 12:16Z, 51+ consumers, no mnesia damage. Remaining follow-ups from the description: check ds3's link/NIC health, gpu-coordinator stuck-job sweep, image mirroring, fstab backup volfile servers, and #756 (mail alerts still dead).
Author
Owner

Migrated to spikerj/spikersoft-infrastructure#177 as part of the umbrella-tracker breakup.

Verified 2026-08-07 — Code: spikersoft-infrastructure@86d03ff — no commit records any of this incident's follow-ups. No backupvolfile-server anywhere in the repo, no gluster/fstab runbook, and seq + rabbitmq are still hostname-pinned to the two SBCs that hung. Live: The outage is RESOLVED. docker node ls → all 10 nodes Ready/Active, dreamstream3 back at 192.168.0.107 (it was hard-down when this was filed). docker service ps seq_seq → Running on dreamstream3; openbao_openbao-2, redis-cluster_redis-node-3 and jetson-influxdb_jetson-influx-3 are all Running again. curl https://api.spikersoft.com/healthz → Healthy, RabbitMQ check reports 267/267 queues running and 78 consumers. docker node inspect dreamstream4 → Status.Addr 192.168.0.108 (the 0.0.0.0 from #925 also cleared on the reboot), and ds4 is no longer leader.
Status: the outage itself is resolved and verified live; follow-ups 2-5 (queue sweep, SBC hardware pass, fstab backup volfile servers, infra-image mirroring) are all unstarted

Closing here. Work now lives in the repo that holds the fix, so fixes #<N> in a PR will
auto-close it on merge. The umbrella tracker keeps cross-repo epics only.

— Opus 5 Agent

Migrated to **spikerj/spikersoft-infrastructure#177** as part of the umbrella-tracker breakup. Verified 2026-08-07 — **Code:** `spikersoft-infrastructure@86d03ff` — no commit records any of this incident's follow-ups. No `backupvolfile-server` anywhere in the repo, no gluster/fstab runbook, and `seq` + `rabbitmq` are still hostname-pinned to the two SBCs that hung. **Live:** **The outage is RESOLVED.** `docker node ls` → all 10 nodes Ready/Active, **dreamstream3 back at 192.168.0.107** (it was hard-down when this was filed). `docker service ps seq_seq` → Running on dreamstream3; `openbao_openbao-2`, `redis-cluster_redis-node-3` and `jetson-influxdb_jetson-influx-3` are all Running again. `curl https://api.spikersoft.com/healthz` → **Healthy**, RabbitMQ check reports 267/267 queues running and 78 consumers. `docker node inspect dreamstream4` → `Status.Addr 192.168.0.108` (the 0.0.0.0 from #925 also cleared on the reboot), and ds4 is no longer leader. Status: the outage itself is resolved and verified live; follow-ups 2-5 (queue sweep, SBC hardware pass, fstab backup volfile servers, infra-image mirroring) are all unstarted Closing here. Work now lives in the repo that holds the fix, so `fixes #<N>` in a PR will auto-close it on merge. The umbrella tracker keeps cross-repo epics only. — Opus 5 Agent
Sign in to join this conversation.