Problem 1 (exit-0) is fixed on master — Program.cs now registers IDockerClient with the annotation explaining exactly this failure: ServiceAutoReconciler (#510-B) resolved `IDockerClient…
Code-side this is fixed on master (198ad66 — backups pinned node.hostname == SERVER like its siblings, plus the $$ escaping so pg_dump's vars survive the compose parser). What remains…
QA Team — corruption verdict 2026-07-14 ~03:00Z: durable, not transient. The missing chunk number 0 for toast value 131293 in pg_toast_2619 error re-fired at 02:44:50Z — well after…
Root cause (confirmed in code + config)
Three numbers that don't add up:
- Declared: image-description leases
GPU_VRAM_REQUIREMENT_MB=16000(its stack file) - Real: Qwen3-VL-8B fp16…
QA Team — predictive heads-up 2026-07-14 ~02:55Z: the kernel-error class that preceded dreamstream4's panic is actively accumulating on other nodes. system-remediation in the last 4h:
-…
QA Team — evidence update + topology correction 2026-07-14 ~02:55Z:
Correction: the contended GPU is the 4090 node's card (the #500 "4090 lane" is in service), not a GPU on SERVER —…
Registry/source findings from repo-side investigation (no host access from this session):
1. The image is unpullable — its namespace no longer exists. The stack pins `git.spikersoft.com/jspik…
Cross-linking evidence from tonight: dreamstream4 being Down is very likely the cause of the GlusterFS /mnt/infrastructure distress Joey observed ~02:14-02:25Z (and possibly of the brief…
QA Team — de-escalation update 2026-07-14 ~02:50Z (follow-up to the URGENT note above):
- The postgres transport errors stopped at 02:29:51Z and have not recurred; timing tracks…
QA Team — watch 2026-07-14 ~02:45Z: gitea-runners_arm_v8_2_act_runner is still crash-looping on dreamstream6 (task: non-zero exit (1) every ~20s, currently 0/1) — roughly 6 hours…
QA Team — root cause CONFIRMED 2026-07-14 ~02:45Z: the 'I/O error' was the reporters failing to reach InfluxDB itself. The main influxDB_influxdb service is pinned to dreamstream4 and…
QA Team — URGENT escalation 2026-07-14 ~02:45Z. The data-integrity question from this ticket is now CONFIRMED as an active failure.
1. Node status: dreamstream4 rejoined ~02:40Z (all…
QA Team — supporting datapoint 2026-07-14 ~02:32Z: from the swarm manager (LAN), https://git.spikersoft.com hard-refused connections (`dial tcp 204.197.150.99:443: connect: connection…
QA Team — watch 2026-07-14 ~02:35Z: the junk docker.node.unknown stream now has a measurable downstream cost beyond RabbitMQ spam. The API is throwing on it:
QA Team — post-close verification 2026-07-14 ~02:20Z: the fix holds in prod.
https://learn.spikersoft.com/spikersoft/quarantine/→ 404- `https://learn.spikersoft.com/spikersoft/uplo…
QA Team — state update 2026-07-14 ~02:20Z:
- The correctly-named stack is now healthy:
spikersoft-docker-monitor_spikersoft-docker-monitorshows 1/1 Running (no more exit-0-on-start…
QA Team — watch 2026-07-14 ~02:20Z:
- dreamstream6 is back:
docker node lsshows Ready/Active. - More evidence for the ds6 DNS/pull problem discussed above: at ~01:20Z swarm tried…