[Infra][Hygiene] Swarm traps from #924 triage: localhost-v6 published-port hang (all nodes), unwired runner cache configs/dead ports, ds4 advertise-addr 0.0.0.0 #925

Closed
opened 2026-08-06 14:12:15 +00:00 by spikerj · 2 comments
Owner

Follow-ups from the #924 investigation. None are outages; all are traps that cost diagnostic time or will bite later.

1. localhost:<published-port> hangs on every node (IPv6 accept-queue trap)

dockerd (24.0.2, swarm mode) binds [::]:<port> for every published port but never accepts connections on them (v6 ingress unsupported). localhost resolves to ::1 first, so curl localhost:30002 etc. hangs forever on any node, while 127.0.0.1 works. This single quirk masqueraded as "4090 ingress mesh dead" during the #924 incident.
Fix: per node, enable v4 preference: uncomment/add precedence ::ffff:0:0/96 100 in /etc/gai.conf. One line, no restart needed.

2. Runner cache config is dead weight / never wired in

  • gitea-act-runner/*-config.yaml files (incl. 4090-config.yaml with cache.host: "<4090-LAN-IP>" placeholder) are NOT mounted into any runner service — runners run act_runner defaults (cache on random port + container IP).
  • The published cache ports (10089/10090/10091/10092/10094-10096 → 8088) therefore point at nothing and refuse/timeout from everywhere.
  • No workflow uses actions/cache (CI uses the nx-cache-server instead), so nothing is actually broken — but the dead ports + unmounted configs are misleading during incident triage.
    Fix (either): (a) drop the published :8088 ports and delete/mark the config files, or (b) actually mount the configs (fix cache.port semantics: act_runner listens ON cache.port — the 10092:8088 mapping assumption in the config comments is wrong) if runner cache is ever wanted.

3. dreamstream4 swarm addr registered as 0.0.0.0

docker node inspect dreamstream4 shows Status.Addr: 0.0.0.0 (every other node shows its LAN IP) — and ds4 is currently the raft leader. Everything works today (overlay, ingress, rabbit), but a 0.0.0.0 advertise address is the classic cause of weird overlay/gossip behavior after restarts. Likely from a join/daemon start without --advertise-addr (possibly the Jul 29 power-outage boot).
Fix: on ds4, set the advertise address explicitly (daemon.json or re-init of the swarm membership) during a quiet window; demote from leader first (docker node demote after another manager takes over) to make it boring.

Diagnosis details in #924 comments. Filed by Claude (Fable 5 agent).

Follow-ups from the #924 investigation. None are outages; all are traps that cost diagnostic time or will bite later. ### 1. `localhost:<published-port>` hangs on every node (IPv6 accept-queue trap) dockerd (24.0.2, swarm mode) binds `[::]:<port>` for every published port but never accepts connections on them (v6 ingress unsupported). `localhost` resolves to `::1` first, so `curl localhost:30002` etc. hangs forever on any node, while `127.0.0.1` works. This single quirk masqueraded as "4090 ingress mesh dead" during the #924 incident. **Fix:** per node, enable v4 preference: uncomment/add `precedence ::ffff:0:0/96 100` in `/etc/gai.conf`. One line, no restart needed. ### 2. Runner cache config is dead weight / never wired in - `gitea-act-runner/*-config.yaml` files (incl. `4090-config.yaml` with `cache.host: "<4090-LAN-IP>"` placeholder) are NOT mounted into any runner service — runners run act_runner defaults (cache on random port + container IP). - The published cache ports (10089/10090/10091/10092/10094-10096 → 8088) therefore point at nothing and refuse/timeout from everywhere. - No workflow uses `actions/cache` (CI uses the nx-cache-server instead), so nothing is actually broken — but the dead ports + unmounted configs are misleading during incident triage. **Fix (either):** (a) drop the published `:8088` ports and delete/mark the config files, or (b) actually mount the configs (fix `cache.port` semantics: act_runner listens ON `cache.port` — the `10092:8088` mapping assumption in the config comments is wrong) if runner cache is ever wanted. ### 3. dreamstream4 swarm addr registered as `0.0.0.0` `docker node inspect dreamstream4` shows `Status.Addr: 0.0.0.0` (every other node shows its LAN IP) — and ds4 is currently the raft **leader**. Everything works today (overlay, ingress, rabbit), but a 0.0.0.0 advertise address is the classic cause of weird overlay/gossip behavior after restarts. Likely from a join/daemon start without `--advertise-addr` (possibly the Jul 29 power-outage boot). **Fix:** on ds4, set the advertise address explicitly (daemon.json or re-init of the swarm membership) during a quiet window; demote from leader first (`docker node demote` after another manager takes over) to make it boring. *Diagnosis details in #924 comments. Filed by Claude (Fable 5 agent).*
Author
Owner

Progress 2026-08-06 (all three items advanced) + corrections from deeper diagnosis:

1. localhost hang — root cause CORRECTED, fixed on 4090

It is not glibc/gai.conf at all: getent/python/wget resolve localhost → 127.0.0.1 and always worked. curl (≥ ~7.86) hardcodes localhost → ::1 + 127.0.0.1 per RFC 6761, ignores the resolver, tries ::1 first, and the connect succeeds into dockerd's never-accepted v6 socket (ss shows Recv-Q piling on the LISTEN socket), so happy-eyeballs commits and hangs. Ergo gai.conf can't help curl — the ds2 gai.conf line was reverted (the one Joey added on 4090 is harmless, keep or drop).
Fix applied: ipv4 in ~/.curlrc on the 4090 — curl localhost:<port> now returns in ms. Same one-liner recommended on any node where humans curl (SERVER): echo ipv4 >> ~/.curlrc. Rule of thumb regardless: probe node ports with 127.0.0.1.

2. Runner cache — DONE live + PR

  • Correction to this ticket's original text: configs ARE wired for the flagship 4090 / laptop / arm runners (node-local copies bind-mounted); only 4090 slots 2-4 are config-less by design. The real defects: act_runner listens on cache.port (verified: flagship bound :10092 in-container) so the 100xx→8088 publishes never matched anything; the 4090's live copy still had the literal <4090-LAN-IP> placeholder; arm-config pointed at a wrong subnet.
  • Since no workflow uses actions/cache (nx-cache-server does CI caching): dead published ports removed from all 7 live services via docker service update --publish-rm (lesson: removing a service's last published port detaches ingress and DOES restart the task — all 7 restarted gracefully and re-declared, no jobs were in flight); 4090's live config copy fixed in place (cache off, ignored 3.0.2 fields dropped) and the runner bounced clean — unknown-field warnings and the stray :10092 listener are gone.
  • Repo alignment: infrastructure PR #159 (stack ports gone + all templates cache-off + runbook fill-in step retired). Post-merge nicety: refresh the stale node-local copies on laptop/ds5/ds6 whenever convenient.

3. ds4 addr 0.0.0.0 — downgraded to trivial

ManagerStatus.Addr is CORRECT (192.168.0.108:2377) — raft is fine, and the ingress path to ds4-hosted containers verified working end-to-end (rabbit mgmt UI 200 via mesh from another node). Only the agent's Status.Addr is 0.0.0.0 — the July 29 boot race (dockerd auto-detected before the NIC was up). No leave/rejoin needed (retracting this ticket's original suggestion): a plain sudo systemctl restart docker on ds4, now that networking is up, should re-detect. Containers survive a daemon restart; leadership bounces harmlessly (9 other managers). Pending: Joey to run it; verify with docker node inspect dreamstream4 --format '{{.Status.Addr}}'.

**Progress 2026-08-06 (all three items advanced) + corrections from deeper diagnosis:** ### 1. localhost hang — root cause CORRECTED, fixed on 4090 It is not glibc/gai.conf at all: `getent`/python/wget resolve localhost → 127.0.0.1 and always worked. **curl** (≥ ~7.86) hardcodes `localhost` → `::1` + `127.0.0.1` per RFC 6761, ignores the resolver, tries `::1` first, and the connect **succeeds** into dockerd's never-accepted v6 socket (`ss` shows Recv-Q piling on the LISTEN socket), so happy-eyeballs commits and hangs. Ergo gai.conf can't help curl — the ds2 gai.conf line was reverted (the one Joey added on 4090 is harmless, keep or drop). **Fix applied:** `ipv4` in `~/.curlrc` on the 4090 — `curl localhost:<port>` now returns in ms. Same one-liner recommended on any node where humans curl (SERVER): `echo ipv4 >> ~/.curlrc`. Rule of thumb regardless: probe node ports with `127.0.0.1`. ### 2. Runner cache — DONE live + PR - Correction to this ticket's original text: configs ARE wired for the flagship 4090 / laptop / arm runners (node-local copies bind-mounted); only 4090 slots 2-4 are config-less by design. The real defects: act_runner **listens on** `cache.port` (verified: flagship bound :10092 in-container) so the `100xx→8088` publishes never matched anything; the 4090's live copy still had the literal `<4090-LAN-IP>` placeholder; arm-config pointed at a wrong subnet. - Since no workflow uses actions/cache (nx-cache-server does CI caching): dead published ports removed from all 7 live services via `docker service update --publish-rm` (lesson: removing a service's last published port detaches ingress and DOES restart the task — all 7 restarted gracefully and re-declared, no jobs were in flight); 4090's live config copy fixed in place (cache off, ignored 3.0.2 fields dropped) and the runner bounced clean — unknown-field warnings and the stray :10092 listener are gone. - Repo alignment: **infrastructure PR #159** (stack ports gone + all templates cache-off + runbook fill-in step retired). Post-merge nicety: refresh the stale node-local copies on laptop/ds5/ds6 whenever convenient. ### 3. ds4 addr 0.0.0.0 — downgraded to trivial `ManagerStatus.Addr` is CORRECT (`192.168.0.108:2377`) — raft is fine, and the ingress path to ds4-hosted containers verified working end-to-end (rabbit mgmt UI 200 via mesh from another node). Only the agent's `Status.Addr` is 0.0.0.0 — the July 29 boot race (dockerd auto-detected before the NIC was up). **No leave/rejoin needed** (retracting this ticket's original suggestion): a plain `sudo systemctl restart docker` on ds4, now that networking is up, should re-detect. Containers survive a daemon restart; leadership bounces harmlessly (9 other managers). Pending: Joey to run it; verify with `docker node inspect dreamstream4 --format '{{.Status.Addr}}'`.
Author
Owner

Migrated to spikerj/spikersoft-infrastructure#190 as part of the umbrella-tracker breakup.

Verified 2026-08-07 — Code: spikersoft-infrastructure@86d03ff — item 2 is DONE: gitea-act-runner/docker-stack.yml:58-70 documents the 2026-08-06 removal of the dead host:100xx -> 8088 cache-port mappings (live services already updated with docker service update --publish-rm), and the per-node cache is disabled in the configs. Items 1 and 3 have no repo artifact — git grep "gai.conf\|precedence ::ffff" and git grep advertise-addr both return nothing. Live: Item 1: fixed on the 4090 only. /etc/gai.conf:66 has an active precedence ::ffff:0:0/96 100, and curl http://localhost:30002/ now returns 200 instead of hanging — but the ticket asks for this per node, and there is nothing to show the other nine got it. Item 3: fixed live but not durably. docker node inspect dreamstream4 now reports Status.Addr 192.168.0.108 (was 0.0.0.0) and ds4 is no longer the raft leader — cleared by the 2026-08-07 reboot rather than by configuration, so a future join/daemon start without --advertise-addr reproduces it.
Status: partially done — runner cache ports removed; gai.conf fixed on the 4090 only; ds4's advertise-addr corrected by a reboot but not made durable

Closing here. Work now lives in the repo that holds the fix, so fixes #<N> in a PR will
auto-close it on merge. The umbrella tracker keeps cross-repo epics only.

— Opus 5 Agent

Migrated to **spikerj/spikersoft-infrastructure#190** as part of the umbrella-tracker breakup. Verified 2026-08-07 — **Code:** `spikersoft-infrastructure@86d03ff` — **item 2 is DONE**: `gitea-act-runner/docker-stack.yml:58-70` documents the 2026-08-06 removal of the dead `host:100xx -> 8088` cache-port mappings (live services already updated with `docker service update --publish-rm`), and the per-node cache is disabled in the configs. Items 1 and 3 have **no repo artifact** — `git grep "gai.conf\|precedence ::ffff"` and `git grep advertise-addr` both return nothing. **Live:** **Item 1: fixed on the 4090 only.** `/etc/gai.conf:66` has an active `precedence ::ffff:0:0/96 100`, and `curl http://localhost:30002/` now returns 200 instead of hanging — but the ticket asks for this per node, and there is nothing to show the other nine got it. **Item 3: fixed live but not durably.** `docker node inspect dreamstream4` now reports `Status.Addr 192.168.0.108` (was `0.0.0.0`) and ds4 is no longer the raft leader — cleared by the 2026-08-07 reboot rather than by configuration, so a future join/daemon start without `--advertise-addr` reproduces it. Status: partially done — runner cache ports removed; gai.conf fixed on the 4090 only; ds4's advertise-addr corrected by a reboot but not made durable Closing here. Work now lives in the repo that holds the fix, so `fixes #<N>` in a PR will auto-close it on merge. The umbrella tracker keeps cross-repo epics only. — Opus 5 Agent
Sign in to join this conversation.