[Critical][Infra][Outage] GlusterFS FUSE mounts hung on ds3+ds4 since 08-05 20:16Z — Seq ingestion/login dead, RabbitMQ (message bus) down since 13:10Z #924

Closed
opened 2026-08-06 13:44:11 +00:00 by spikerj · 2 comments
Owner

Live incident (2026-08-06): GlusterFS FUSE mounts hung on dreamstream3 + dreamstream4

Discovered while investigating "can't log into Seq". Root cause is NOT Seq — the localhost:/staging-gfs FUSE mount at /mnt is hard-hung on ds3 and ds4 (healthy on ds2; SERVER doesn't mount gluster). Everything with data under /mnt on those two nodes is stalled.

Impact (verified, not speculative)

  1. Seq (ds3, data bind /mnt/seq/data)
    • UI login hangs forever: any HTTP request with a body never completes (login POST = write path). GETs work from cache, so the UI loads and looks fine.
    • Event ingestion returns 503 Service Unavailable on :5341 — verified with a direct in-overlay probe on ds3. Platform-wide log ingestion is DEAD; nothing is being recorded (ironically including the errors this outage is causing).
  2. RabbitMQ (ds4, data bind /mnt/rabbitmq)
    • Got a SIGTERM at 2026-08-06 13:10:18Z (origin unknown) and is hung mid-shutdown — beam can't flush to the hung mount. Nothing logged since.
    • Verified from an in-overlay probe: AMQP 5672 refuses connections (mgmt 15672 still open). Message bus is down now; all consumers/publishers are failing. Swarm still shows the task "Running" so no restart is being attempted.
  3. Anything else touching /mnt from ds3/ds4 will block in D-state. (OpenBao raft, redis-cluster, jetson-influx on those nodes use local volumes — they're fine.)

Timeline / evidence

  • 2026-08-05 20:16:38-46Z — brick logs on ds3: ds2's FUSE clients disconnect; glusterd on ds3 logs peer 192.168.0.106 disconnect plus Lock for vol staging-gfs not held / Lock not released for staging-gfs.
  • 20:17:50Z — ds3's mount log: client-1 reconnects, "fds open - Delaying child_up until they are re-opened [{count=845}]", then CHILD_UP … and silence since. The FUSE clients on ds3/ds4 never actually recovered — stat /mnt hangs indefinitely on both (verified with in-node probes; df wedges the moment it touches /mnt).
  • Prior days' mount logs show repeated client_pre_lk_v2 … remote_fd is -1 … EBADFD lock-recovery failures — classic post-brick-bounce fd/lock recovery deadlock.
  • Volume layout: staging-gfs, bricks /gluster/volume1 on 192.168.0.105/106/107/108 (ds1-ds4). All peers state=3 (in cluster). Brick processes are up — this is a client-side (FUSE) hang, not brick loss.
  • Node root disks are fine (ds3 nvme 28% used) — this is not a disk-space 503.

Remediation runbook (needs sudo on ds3 + ds4)

Per node (ds4 first — message bus):

  1. From a manager: scale the affected service to 0 so nothing new touches the mount: docker service scale rabbitmq_rabbitmq=0 (ds4) / docker service scale seq_seq=0 (ds3). Stops may hang — that's expected (D-state).
  2. On the node: sudo umount -l /mnt then remount: sudo mount /mnt (or sudo mount -t glusterfs localhost:/staging-gfs /mnt). Verify: timeout 3 stat /mnt && timeout 3 ls /mnt.
  3. If umount/remount can't clear it (D-state pins), reboot the node (drain first: docker node update --availability drain <node>; undrain after).
  4. Scale services back to 1. Verify:
    • Seq: login works; curl -X POST http://<seq>:5341/api/events/raw … returns 201; new events flowing.
    • RabbitMQ: watch boot logs for mnesia corruption (#895 pattern — if consumers get refused / user table empty, use the #895 wipe-and-fresh-boot runbook); confirm consumers reconnect and kick any stuck-Complete jobs per #895 list.
  5. Check ds1/ds2 FUSE clients afterwards (timeout 3 stat /mnt) — only ds3/ds4 were hung as of 13:45Z.

Notes

  • The SIGTERM to rabbit at 13:10Z is unexplained — docker desired-state is still Running, so it wasn't a swarm-initiated update. If nobody did it manually, worth finding the sender after recovery.
  • Related: #895 (power-outage rabbit corruption — same data dir at risk again), #413 (epic to get off GlusterFS — this incident is more ammunition).
  • Separate minor finding while probing: the swarm ingress mesh is entirely dead from the 4090 node (every published port, e.g. 30002/8086/16686, times out from 4090; fine from other nodes), and ds4's swarm-registered addr shows 0.0.0.0. Not the cause of this incident, but worth a look.

Filed by Claude (Fable 5 agent) from the 4090 session; diagnosis via read-only one-shot swarm jobs.

# Live incident (2026-08-06): GlusterFS FUSE mounts hung on dreamstream3 + dreamstream4 Discovered while investigating "can't log into Seq". Root cause is NOT Seq — the `localhost:/staging-gfs` FUSE mount at `/mnt` is **hard-hung on ds3 and ds4** (healthy on ds2; SERVER doesn't mount gluster). Everything with data under `/mnt` on those two nodes is stalled. ## Impact (verified, not speculative) 1. **Seq (ds3, data bind `/mnt/seq/data`)** - UI login hangs forever: any HTTP request **with a body** never completes (login POST = write path). GETs work from cache, so the UI loads and *looks* fine. - **Event ingestion returns `503 Service Unavailable`** on `:5341` — verified with a direct in-overlay probe on ds3. Platform-wide log ingestion is DEAD; nothing is being recorded (ironically including the errors this outage is causing). 2. **RabbitMQ (ds4, data bind `/mnt/rabbitmq`)** - Got a SIGTERM at **2026-08-06 13:10:18Z** (origin unknown) and is hung mid-shutdown — beam can't flush to the hung mount. Nothing logged since. - Verified from an in-overlay probe: **AMQP 5672 refuses connections** (mgmt 15672 still open). **Message bus is down now**; all consumers/publishers are failing. Swarm still shows the task "Running" so no restart is being attempted. 3. Anything else touching `/mnt` from ds3/ds4 will block in D-state. (OpenBao raft, redis-cluster, jetson-influx on those nodes use local volumes — they're fine.) ## Timeline / evidence - `2026-08-05 20:16:38-46Z` — brick logs on ds3: ds2's FUSE clients disconnect; glusterd on ds3 logs peer `192.168.0.106` disconnect **plus `Lock for vol staging-gfs not held` / `Lock not released for staging-gfs`**. - `20:17:50Z` — ds3's mount log: client-1 reconnects, "fds open - Delaying child_up until they are re-opened [{count=845}]", then CHILD_UP … and **silence since**. The FUSE clients on ds3/ds4 never actually recovered — `stat /mnt` hangs indefinitely on both (verified with in-node probes; `df` wedges the moment it touches `/mnt`). - Prior days' mount logs show repeated `client_pre_lk_v2 … remote_fd is -1 … EBADFD` lock-recovery failures — classic post-brick-bounce fd/lock recovery deadlock. - Volume layout: `staging-gfs`, bricks `/gluster/volume1` on 192.168.0.105/106/107/108 (ds1-ds4). All peers state=3 (in cluster). Brick processes are up — this is a **client-side (FUSE) hang**, not brick loss. - Node root disks are fine (ds3 nvme 28% used) — this is not a disk-space 503. ## Remediation runbook (needs sudo on ds3 + ds4) Per node (ds4 first — message bus): 1. From a manager: scale the affected service to 0 so nothing new touches the mount: `docker service scale rabbitmq_rabbitmq=0` (ds4) / `docker service scale seq_seq=0` (ds3). Stops may hang — that's expected (D-state). 2. On the node: `sudo umount -l /mnt` then remount: `sudo mount /mnt` (or `sudo mount -t glusterfs localhost:/staging-gfs /mnt`). Verify: `timeout 3 stat /mnt && timeout 3 ls /mnt`. 3. If umount/remount can't clear it (D-state pins), **reboot the node** (drain first: `docker node update --availability drain <node>`; undrain after). 4. Scale services back to 1. Verify: - Seq: login works; `curl -X POST http://<seq>:5341/api/events/raw …` returns 201; new events flowing. - RabbitMQ: watch boot logs for mnesia corruption (#895 pattern — if consumers get refused / user table empty, use the #895 wipe-and-fresh-boot runbook); confirm consumers reconnect and kick any stuck-Complete jobs per #895 list. 5. Check ds1/ds2 FUSE clients afterwards (`timeout 3 stat /mnt`) — only ds3/ds4 were hung as of 13:45Z. ## Notes - The SIGTERM to rabbit at 13:10Z is unexplained — docker desired-state is still Running, so it wasn't a swarm-initiated update. If nobody did it manually, worth finding the sender after recovery. - Related: #895 (power-outage rabbit corruption — same data dir at risk again), #413 (epic to get off GlusterFS — this incident is more ammunition). - Separate minor finding while probing: the swarm **ingress mesh is entirely dead from the 4090 node** (every published port, e.g. 30002/8086/16686, times out from 4090; fine from other nodes), and ds4's swarm-registered addr shows `0.0.0.0`. Not the cause of this incident, but worth a look. *Filed by Claude (Fable 5 agent) from the 4090 session; diagnosis via read-only one-shot swarm jobs.*
Author
Owner

Resolved 2026-08-06 ~13:56Z — user ran the runbook (remount on ds3+ds4, services scaled back up). Post-recovery verification:

  • stat/ls/df on /mnt instant on both ds3 and ds4; both show the shared volume (264.8G used) again.
  • Seq: login POST responds in <1s (was infinite hang); ingestion POST :5341/api/events/raw accepted ({"MinimumLevelAccepted":null}) — log ingestion restored.
  • RabbitMQ: clean boot at 13:56:00Z, mnesia tables loaded, unclean-shutdown queue scans dropped 0 messages, 2 invalid .qi segment files self-healed, "Server startup complete; 6 plugins started", 72 AMQP connections accepted in the first 15 min — consumers reconnected. No #895-style corruption.

Remaining loose ends moved out of this ticket: unexplained SIGTERM sender (13:10:18Z), 4090 ingress-mesh dead, ds4 swarm addr 0.0.0.0. Log-ingestion blackout window: ~2026-08-05 20:16Z → 08-06 13:56Z.

**Resolved 2026-08-06 ~13:56Z** — user ran the runbook (remount on ds3+ds4, services scaled back up). Post-recovery verification: - `stat`/`ls`/`df` on `/mnt` instant on both ds3 and ds4; both show the shared volume (264.8G used) again. - **Seq**: login POST responds in <1s (was infinite hang); ingestion `POST :5341/api/events/raw` accepted (`{"MinimumLevelAccepted":null}`) — log ingestion restored. - **RabbitMQ**: clean boot at 13:56:00Z, mnesia tables loaded, unclean-shutdown queue scans dropped **0** messages, 2 invalid `.qi` segment files self-healed, "Server startup complete; 6 plugins started", **72 AMQP connections accepted in the first 15 min** — consumers reconnected. No #895-style corruption. Remaining loose ends moved out of this ticket: unexplained SIGTERM sender (13:10:18Z), 4090 ingress-mesh dead, ds4 swarm addr `0.0.0.0`. Log-ingestion blackout window: ~2026-08-05 20:16Z → 08-06 13:56Z.
Author
Owner

Correction to the "4090 ingress mesh dead" side-note — investigated today, and the ingress mesh on 4090 was never broken. IPv4 ingress works in all directions (inbound from other nodes, outbound vxlan, local via 127.0.0.1 and LAN IP). What actually happens, on EVERY node (verified identical on ds2): dockerd binds IPv6 wildcard listeners on all published ports but never accepts from them (swarm v6 ingress is unsupported in docker 24; ss shows Recv-Q backlog rotting in the accept queue). Since localhost resolves to ::1 first, any curl localhost:<published-port> hangs — which is exactly what my probes did. Workaround: use 127.0.0.1, or prefer v4 in /etc/gai.conf (precedence ::ffff:0:0/96 100). Filed the related hygiene items separately.

**Correction to the "4090 ingress mesh dead" side-note** — investigated today, and the ingress mesh on 4090 was never broken. IPv4 ingress works in all directions (inbound from other nodes, outbound vxlan, local via 127.0.0.1 and LAN IP). What actually happens, on EVERY node (verified identical on ds2): dockerd binds IPv6 wildcard listeners on all published ports but never accepts from them (swarm v6 ingress is unsupported in docker 24; `ss` shows Recv-Q backlog rotting in the accept queue). Since `localhost` resolves to `::1` first, any `curl localhost:<published-port>` hangs — which is exactly what my probes did. Workaround: use `127.0.0.1`, or prefer v4 in `/etc/gai.conf` (`precedence ::ffff:0:0/96 100`). Filed the related hygiene items separately.
Sign in to join this conversation.