QA Team — alert-tuning addendum 2026-07-14 ~03:15Z: the fleet kernel errors turn out to be PCIe AER corrected errors (see #552 for the full diagnosis). AER corrected errors can burst by…
Fixed and CI-verified. Angular PR #191 (merged): the three i18n asset imports in testing/transloco-test-setup.ts are now the workspace's second documented boundary exemption (inline…
QA Team — kernel-error signature identified 2026-07-14 ~03:15Z (kernel log pulled from a dreamstream node by ops):
The fleet 'kernel errors' are **PCIe AER corrected-error bursts on the NVMe…
QA Team — refinement/correction 2026-07-14 ~03:10Z on the SignalR errors noted earlier: Jaeger traces show the 02:27:48–49Z ForwardClusterEventToSignalR failures were **Redis PUBLISH…
Problem 1 (exit-0) is fixed on master — Program.cs now registers IDockerClient with the annotation explaining exactly this failure: ServiceAutoReconciler (#510-B) resolved `IDockerClient…
Code-side this is fixed on master (198ad66 — backups pinned node.hostname == SERVER like its siblings, plus the $$ escaping so pg_dump's vars survive the compose parser). What remains…
QA Team — corruption verdict 2026-07-14 ~03:00Z: durable, not transient. The missing chunk number 0 for toast value 131293 in pg_toast_2619 error re-fired at 02:44:50Z — well after…
Root cause (confirmed in code + config)
Three numbers that don't add up:
- Declared: image-description leases
GPU_VRAM_REQUIREMENT_MB=16000(its stack file) - Real: Qwen3-VL-8B fp16…
QA Team — predictive heads-up 2026-07-14 ~02:55Z: the kernel-error class that preceded dreamstream4's panic is actively accumulating on other nodes. system-remediation in the last 4h:
-…
QA Team — evidence update + topology correction 2026-07-14 ~02:55Z:
Correction: the contended GPU is the 4090 node's card (the #500 "4090 lane" is in service), not a GPU on SERVER —…
Registry/source findings from repo-side investigation (no host access from this session):
1. The image is unpullable — its namespace no longer exists. The stack pins `git.spikersoft.com/jspik…
Cross-linking evidence from tonight: dreamstream4 being Down is very likely the cause of the GlusterFS /mnt/infrastructure distress Joey observed ~02:14-02:25Z (and possibly of the brief…
QA Team — de-escalation update 2026-07-14 ~02:50Z (follow-up to the URGENT note above):
- The postgres transport errors stopped at 02:29:51Z and have not recurred; timing tracks…
QA Team — watch 2026-07-14 ~02:45Z: gitea-runners_arm_v8_2_act_runner is still crash-looping on dreamstream6 (task: non-zero exit (1) every ~20s, currently 0/1) — roughly 6 hours…
QA Team — root cause CONFIRMED 2026-07-14 ~02:45Z: the 'I/O error' was the reporters failing to reach InfluxDB itself. The main influxDB_influxdb service is pinned to dreamstream4 and…
QA Team — URGENT escalation 2026-07-14 ~02:45Z. The data-integrity question from this ticket is now CONFIRMED as an active failure.
1. Node status: dreamstream4 rejoined ~02:40Z (all…
QA Team — supporting datapoint 2026-07-14 ~02:32Z: from the swarm manager (LAN), https://git.spikersoft.com hard-refused connections (`dial tcp 204.197.150.99:443: connect: connection…
QA Team — watch 2026-07-14 ~02:35Z: the junk docker.node.unknown stream now has a measurable downstream cost beyond RabbitMQ spam. The API is throwing on it:
[02:27:59 ERR] [Production/…