Follow-ups from the #924 investigation. None are outages; all are traps that cost diagnostic time or will bite later.
1. localhost:<published-port> hangs on every node (IPv6 accept-queue trap)
dockerd (24.0.2, swarm mode) binds [::]:<port> for every published port but never accepts connections on them (v6 ingress unsupported). localhost resolves to ::1 first, so curl localhost:30002 etc. hangs forever on any node, while 127.0.0.1 works. This single quirk masqueraded as "4090 ingress mesh dead" during the #924 incident. Fix: per node, enable v4 preference: uncomment/add precedence ::ffff:0:0/96 100 in /etc/gai.conf. One line, no restart needed.
2. Runner cache config is dead weight / never wired in
gitea-act-runner/*-config.yaml files (incl. 4090-config.yaml with cache.host: "<4090-LAN-IP>" placeholder) are NOT mounted into any runner service — runners run act_runner defaults (cache on random port + container IP).
The published cache ports (10089/10090/10091/10092/10094-10096 → 8088) therefore point at nothing and refuse/timeout from everywhere.
No workflow uses actions/cache (CI uses the nx-cache-server instead), so nothing is actually broken — but the dead ports + unmounted configs are misleading during incident triage. Fix (either): (a) drop the published :8088 ports and delete/mark the config files, or (b) actually mount the configs (fix cache.port semantics: act_runner listens ON cache.port — the 10092:8088 mapping assumption in the config comments is wrong) if runner cache is ever wanted.
3. dreamstream4 swarm addr registered as 0.0.0.0
docker node inspect dreamstream4 shows Status.Addr: 0.0.0.0 (every other node shows its LAN IP) — and ds4 is currently the raft leader. Everything works today (overlay, ingress, rabbit), but a 0.0.0.0 advertise address is the classic cause of weird overlay/gossip behavior after restarts. Likely from a join/daemon start without --advertise-addr (possibly the Jul 29 power-outage boot). Fix: on ds4, set the advertise address explicitly (daemon.json or re-init of the swarm membership) during a quiet window; demote from leader first (docker node demote after another manager takes over) to make it boring.
Diagnosis details in #924 comments. Filed by Claude (Fable 5 agent).
Follow-ups from the #924 investigation. None are outages; all are traps that cost diagnostic time or will bite later.
### 1. `localhost:<published-port>` hangs on every node (IPv6 accept-queue trap)
dockerd (24.0.2, swarm mode) binds `[::]:<port>` for every published port but never accepts connections on them (v6 ingress unsupported). `localhost` resolves to `::1` first, so `curl localhost:30002` etc. hangs forever on any node, while `127.0.0.1` works. This single quirk masqueraded as "4090 ingress mesh dead" during the #924 incident.
**Fix:** per node, enable v4 preference: uncomment/add `precedence ::ffff:0:0/96 100` in `/etc/gai.conf`. One line, no restart needed.
### 2. Runner cache config is dead weight / never wired in
- `gitea-act-runner/*-config.yaml` files (incl. `4090-config.yaml` with `cache.host: "<4090-LAN-IP>"` placeholder) are NOT mounted into any runner service — runners run act_runner defaults (cache on random port + container IP).
- The published cache ports (10089/10090/10091/10092/10094-10096 → 8088) therefore point at nothing and refuse/timeout from everywhere.
- No workflow uses `actions/cache` (CI uses the nx-cache-server instead), so nothing is actually broken — but the dead ports + unmounted configs are misleading during incident triage.
**Fix (either):** (a) drop the published `:8088` ports and delete/mark the config files, or (b) actually mount the configs (fix `cache.port` semantics: act_runner listens ON `cache.port` — the `10092:8088` mapping assumption in the config comments is wrong) if runner cache is ever wanted.
### 3. dreamstream4 swarm addr registered as `0.0.0.0`
`docker node inspect dreamstream4` shows `Status.Addr: 0.0.0.0` (every other node shows its LAN IP) — and ds4 is currently the raft **leader**. Everything works today (overlay, ingress, rabbit), but a 0.0.0.0 advertise address is the classic cause of weird overlay/gossip behavior after restarts. Likely from a join/daemon start without `--advertise-addr` (possibly the Jul 29 power-outage boot).
**Fix:** on ds4, set the advertise address explicitly (daemon.json or re-init of the swarm membership) during a quiet window; demote from leader first (`docker node demote` after another manager takes over) to make it boring.
*Diagnosis details in #924 comments. Filed by Claude (Fable 5 agent).*
Progress 2026-08-06 (all three items advanced) + corrections from deeper diagnosis:
1. localhost hang — root cause CORRECTED, fixed on 4090
It is not glibc/gai.conf at all: getent/python/wget resolve localhost → 127.0.0.1 and always worked. curl (≥ ~7.86) hardcodes localhost → ::1 + 127.0.0.1 per RFC 6761, ignores the resolver, tries ::1 first, and the connect succeeds into dockerd's never-accepted v6 socket (ss shows Recv-Q piling on the LISTEN socket), so happy-eyeballs commits and hangs. Ergo gai.conf can't help curl — the ds2 gai.conf line was reverted (the one Joey added on 4090 is harmless, keep or drop). Fix applied:ipv4 in ~/.curlrc on the 4090 — curl localhost:<port> now returns in ms. Same one-liner recommended on any node where humans curl (SERVER): echo ipv4 >> ~/.curlrc. Rule of thumb regardless: probe node ports with 127.0.0.1.
2. Runner cache — DONE live + PR
Correction to this ticket's original text: configs ARE wired for the flagship 4090 / laptop / arm runners (node-local copies bind-mounted); only 4090 slots 2-4 are config-less by design. The real defects: act_runner listens oncache.port (verified: flagship bound :10092 in-container) so the 100xx→8088 publishes never matched anything; the 4090's live copy still had the literal <4090-LAN-IP> placeholder; arm-config pointed at a wrong subnet.
Since no workflow uses actions/cache (nx-cache-server does CI caching): dead published ports removed from all 7 live services via docker service update --publish-rm (lesson: removing a service's last published port detaches ingress and DOES restart the task — all 7 restarted gracefully and re-declared, no jobs were in flight); 4090's live config copy fixed in place (cache off, ignored 3.0.2 fields dropped) and the runner bounced clean — unknown-field warnings and the stray :10092 listener are gone.
Repo alignment: infrastructure PR #159 (stack ports gone + all templates cache-off + runbook fill-in step retired). Post-merge nicety: refresh the stale node-local copies on laptop/ds5/ds6 whenever convenient.
3. ds4 addr 0.0.0.0 — downgraded to trivial
ManagerStatus.Addr is CORRECT (192.168.0.108:2377) — raft is fine, and the ingress path to ds4-hosted containers verified working end-to-end (rabbit mgmt UI 200 via mesh from another node). Only the agent's Status.Addr is 0.0.0.0 — the July 29 boot race (dockerd auto-detected before the NIC was up). No leave/rejoin needed (retracting this ticket's original suggestion): a plain sudo systemctl restart docker on ds4, now that networking is up, should re-detect. Containers survive a daemon restart; leadership bounces harmlessly (9 other managers). Pending: Joey to run it; verify with docker node inspect dreamstream4 --format '{{.Status.Addr}}'.
**Progress 2026-08-06 (all three items advanced) + corrections from deeper diagnosis:**
### 1. localhost hang — root cause CORRECTED, fixed on 4090
It is not glibc/gai.conf at all: `getent`/python/wget resolve localhost → 127.0.0.1 and always worked. **curl** (≥ ~7.86) hardcodes `localhost` → `::1` + `127.0.0.1` per RFC 6761, ignores the resolver, tries `::1` first, and the connect **succeeds** into dockerd's never-accepted v6 socket (`ss` shows Recv-Q piling on the LISTEN socket), so happy-eyeballs commits and hangs. Ergo gai.conf can't help curl — the ds2 gai.conf line was reverted (the one Joey added on 4090 is harmless, keep or drop).
**Fix applied:** `ipv4` in `~/.curlrc` on the 4090 — `curl localhost:<port>` now returns in ms. Same one-liner recommended on any node where humans curl (SERVER): `echo ipv4 >> ~/.curlrc`. Rule of thumb regardless: probe node ports with `127.0.0.1`.
### 2. Runner cache — DONE live + PR
- Correction to this ticket's original text: configs ARE wired for the flagship 4090 / laptop / arm runners (node-local copies bind-mounted); only 4090 slots 2-4 are config-less by design. The real defects: act_runner **listens on** `cache.port` (verified: flagship bound :10092 in-container) so the `100xx→8088` publishes never matched anything; the 4090's live copy still had the literal `<4090-LAN-IP>` placeholder; arm-config pointed at a wrong subnet.
- Since no workflow uses actions/cache (nx-cache-server does CI caching): dead published ports removed from all 7 live services via `docker service update --publish-rm` (lesson: removing a service's last published port detaches ingress and DOES restart the task — all 7 restarted gracefully and re-declared, no jobs were in flight); 4090's live config copy fixed in place (cache off, ignored 3.0.2 fields dropped) and the runner bounced clean — unknown-field warnings and the stray :10092 listener are gone.
- Repo alignment: **infrastructure PR #159** (stack ports gone + all templates cache-off + runbook fill-in step retired). Post-merge nicety: refresh the stale node-local copies on laptop/ds5/ds6 whenever convenient.
### 3. ds4 addr 0.0.0.0 — downgraded to trivial
`ManagerStatus.Addr` is CORRECT (`192.168.0.108:2377`) — raft is fine, and the ingress path to ds4-hosted containers verified working end-to-end (rabbit mgmt UI 200 via mesh from another node). Only the agent's `Status.Addr` is 0.0.0.0 — the July 29 boot race (dockerd auto-detected before the NIC was up). **No leave/rejoin needed** (retracting this ticket's original suggestion): a plain `sudo systemctl restart docker` on ds4, now that networking is up, should re-detect. Containers survive a daemon restart; leadership bounces harmlessly (9 other managers). Pending: Joey to run it; verify with `docker node inspect dreamstream4 --format '{{.Status.Addr}}'`.
Verified 2026-08-07 — Code:spikersoft-infrastructure@86d03ff — item 2 is DONE: gitea-act-runner/docker-stack.yml:58-70 documents the 2026-08-06 removal of the dead host:100xx -> 8088 cache-port mappings (live services already updated with docker service update --publish-rm), and the per-node cache is disabled in the configs. Items 1 and 3 have no repo artifact — git grep "gai.conf\|precedence ::ffff" and git grep advertise-addr both return nothing. Live:Item 1: fixed on the 4090 only./etc/gai.conf:66 has an active precedence ::ffff:0:0/96 100, and curl http://localhost:30002/ now returns 200 instead of hanging — but the ticket asks for this per node, and there is nothing to show the other nine got it. Item 3: fixed live but not durably.docker node inspect dreamstream4 now reports Status.Addr 192.168.0.108 (was 0.0.0.0) and ds4 is no longer the raft leader — cleared by the 2026-08-07 reboot rather than by configuration, so a future join/daemon start without --advertise-addr reproduces it.
Status: partially done — runner cache ports removed; gai.conf fixed on the 4090 only; ds4's advertise-addr corrected by a reboot but not made durable
Closing here. Work now lives in the repo that holds the fix, so fixes #<N> in a PR will
auto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Migrated to **spikerj/spikersoft-infrastructure#190** as part of the umbrella-tracker breakup.
Verified 2026-08-07 — **Code:** `spikersoft-infrastructure@86d03ff` — **item 2 is DONE**: `gitea-act-runner/docker-stack.yml:58-70` documents the 2026-08-06 removal of the dead `host:100xx -> 8088` cache-port mappings (live services already updated with `docker service update --publish-rm`), and the per-node cache is disabled in the configs. Items 1 and 3 have **no repo artifact** — `git grep "gai.conf\|precedence ::ffff"` and `git grep advertise-addr` both return nothing. **Live:** **Item 1: fixed on the 4090 only.** `/etc/gai.conf:66` has an active `precedence ::ffff:0:0/96 100`, and `curl http://localhost:30002/` now returns 200 instead of hanging — but the ticket asks for this per node, and there is nothing to show the other nine got it. **Item 3: fixed live but not durably.** `docker node inspect dreamstream4` now reports `Status.Addr 192.168.0.108` (was `0.0.0.0`) and ds4 is no longer the raft leader — cleared by the 2026-08-07 reboot rather than by configuration, so a future join/daemon start without `--advertise-addr` reproduces it.
Status: partially done — runner cache ports removed; gai.conf fixed on the 4090 only; ds4's advertise-addr corrected by a reboot but not made durable
Closing here. Work now lives in the repo that holds the fix, so `fixes #<N>` in a PR will
auto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Follow-ups from the #924 investigation. None are outages; all are traps that cost diagnostic time or will bite later.
1.
localhost:<published-port>hangs on every node (IPv6 accept-queue trap)dockerd (24.0.2, swarm mode) binds
[::]:<port>for every published port but never accepts connections on them (v6 ingress unsupported).localhostresolves to::1first, socurl localhost:30002etc. hangs forever on any node, while127.0.0.1works. This single quirk masqueraded as "4090 ingress mesh dead" during the #924 incident.Fix: per node, enable v4 preference: uncomment/add
precedence ::ffff:0:0/96 100in/etc/gai.conf. One line, no restart needed.2. Runner cache config is dead weight / never wired in
gitea-act-runner/*-config.yamlfiles (incl.4090-config.yamlwithcache.host: "<4090-LAN-IP>"placeholder) are NOT mounted into any runner service — runners run act_runner defaults (cache on random port + container IP).actions/cache(CI uses the nx-cache-server instead), so nothing is actually broken — but the dead ports + unmounted configs are misleading during incident triage.Fix (either): (a) drop the published
:8088ports and delete/mark the config files, or (b) actually mount the configs (fixcache.portsemantics: act_runner listens ONcache.port— the10092:8088mapping assumption in the config comments is wrong) if runner cache is ever wanted.3. dreamstream4 swarm addr registered as
0.0.0.0docker node inspect dreamstream4showsStatus.Addr: 0.0.0.0(every other node shows its LAN IP) — and ds4 is currently the raft leader. Everything works today (overlay, ingress, rabbit), but a 0.0.0.0 advertise address is the classic cause of weird overlay/gossip behavior after restarts. Likely from a join/daemon start without--advertise-addr(possibly the Jul 29 power-outage boot).Fix: on ds4, set the advertise address explicitly (daemon.json or re-init of the swarm membership) during a quiet window; demote from leader first (
docker node demoteafter another manager takes over) to make it boring.Diagnosis details in #924 comments. Filed by Claude (Fable 5 agent).
Progress 2026-08-06 (all three items advanced) + corrections from deeper diagnosis:
1. localhost hang — root cause CORRECTED, fixed on 4090
It is not glibc/gai.conf at all:
getent/python/wget resolve localhost → 127.0.0.1 and always worked. curl (≥ ~7.86) hardcodeslocalhost→::1+127.0.0.1per RFC 6761, ignores the resolver, tries::1first, and the connect succeeds into dockerd's never-accepted v6 socket (ssshows Recv-Q piling on the LISTEN socket), so happy-eyeballs commits and hangs. Ergo gai.conf can't help curl — the ds2 gai.conf line was reverted (the one Joey added on 4090 is harmless, keep or drop).Fix applied:
ipv4in~/.curlrcon the 4090 —curl localhost:<port>now returns in ms. Same one-liner recommended on any node where humans curl (SERVER):echo ipv4 >> ~/.curlrc. Rule of thumb regardless: probe node ports with127.0.0.1.2. Runner cache — DONE live + PR
cache.port(verified: flagship bound :10092 in-container) so the100xx→8088publishes never matched anything; the 4090's live copy still had the literal<4090-LAN-IP>placeholder; arm-config pointed at a wrong subnet.docker service update --publish-rm(lesson: removing a service's last published port detaches ingress and DOES restart the task — all 7 restarted gracefully and re-declared, no jobs were in flight); 4090's live config copy fixed in place (cache off, ignored 3.0.2 fields dropped) and the runner bounced clean — unknown-field warnings and the stray :10092 listener are gone.3. ds4 addr 0.0.0.0 — downgraded to trivial
ManagerStatus.Addris CORRECT (192.168.0.108:2377) — raft is fine, and the ingress path to ds4-hosted containers verified working end-to-end (rabbit mgmt UI 200 via mesh from another node). Only the agent'sStatus.Addris 0.0.0.0 — the July 29 boot race (dockerd auto-detected before the NIC was up). No leave/rejoin needed (retracting this ticket's original suggestion): a plainsudo systemctl restart dockeron ds4, now that networking is up, should re-detect. Containers survive a daemon restart; leadership bounces harmlessly (9 other managers). Pending: Joey to run it; verify withdocker node inspect dreamstream4 --format '{{.Status.Addr}}'.Migrated to spikerj/spikersoft-infrastructure#190 as part of the umbrella-tracker breakup.
Verified 2026-08-07 — Code:
spikersoft-infrastructure@86d03ff— item 2 is DONE:gitea-act-runner/docker-stack.yml:58-70documents the 2026-08-06 removal of the deadhost:100xx -> 8088cache-port mappings (live services already updated withdocker service update --publish-rm), and the per-node cache is disabled in the configs. Items 1 and 3 have no repo artifact —git grep "gai.conf\|precedence ::ffff"andgit grep advertise-addrboth return nothing. Live: Item 1: fixed on the 4090 only./etc/gai.conf:66has an activeprecedence ::ffff:0:0/96 100, andcurl http://localhost:30002/now returns 200 instead of hanging — but the ticket asks for this per node, and there is nothing to show the other nine got it. Item 3: fixed live but not durably.docker node inspect dreamstream4now reportsStatus.Addr 192.168.0.108(was0.0.0.0) and ds4 is no longer the raft leader — cleared by the 2026-08-07 reboot rather than by configuration, so a future join/daemon start without--advertise-addrreproduces it.Status: partially done — runner cache ports removed; gai.conf fixed on the 4090 only; ds4's advertise-addr corrected by a reboot but not made durable
Closing here. Work now lives in the repo that holds the fix, so
fixes #<N>in a PR willauto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent