Fix in backend PR #279 — took your first fix shape (score all, then top K), not the fail-loudly one: the fallback exists to keep RAG alive through an index outage, and it can genuinely do…
Fix in backend PR #278 — both actionable fix shapes, coordinator-side.
Fix shape 2 (the arbitration defect) — churn-as-pressure, works today with zero new telemetry. The structural hole…
QA Team — triage update 2026-07-14: per spikerj, the committed appsettings credentials inventoried above are known development values — no emergency rotation needed. The inventory stands…
QA Team — a perfect concrete specimen for this ticket's severity-classification ask, 2026-07-14 ~13:35Z: system-remediation just dispatched a HIGH 'Kernel error on 4090' ops email…
QA Team — diagnosis sharpened 2026-07-14 ~09:45Z, and severity is worse than filed:
- The coderunner occurrence is not a 1-3 min transient: replica .4 has now been bouncing between…
QA Team — inventory input for this epic, 2026-07-14: a full fact-checking audit of every spikersoft-backend project (README rebuild, 100+ agents reading every csproj/appsettings/Program.cs)…
Recovered. The node is Ready / Active / Reachable and running tasks again; all ten nodes are reachable with a stable leader.
Closing on observed state, with the same caveat I left on #552:…
Recovered. The node is Ready / Active / Reachable, it is carrying tasks again, and the whole manager set is reachable with a stable leader.
The concrete casualty in this ticket is also…
Found it: it's Traefik's InfluxDB metrics push
Traefik's metrics.influxdb2 exporter pushes on a 10s interval by default — an exact match for the reported cadence (90 rejections per…