Observed during the #885 broker outage: SpikerSoft.EventHandlers.ArtPipeProcessor (resident artpipe-model-safety service) attempts its RabbitMQ connection ~0s after boot, and a failure is fatal: Failed to create RabbitMQ connection → Hosting failed to start → exit 1. Swarm's on-failure restart (5s delay) turns that into a ~7s crash-loop — 3 error events per cycle, ~1,500+ Seq errors over the 2h outage, plus constant container churn on the GPU node.
Contrast: MassTransit consumers auto-recover and re-declare topology when the broker returns (per rabbitmq/docker-stack.yml deploy notes); this host's bespoke RabbitMQ.Connection bootstrap does not.
Fix: wrap broker connection creation at startup in retry-with-backoff (e.g. Polly exponential, cap ~30-60s, log Warning per attempt and Error only after N failures), or let the host start degraded and connect lazily with the client's built-in automatic recovery. Applies to whatever shares this bootstrap in SpikerSoft.EventHandlers.Infrastructure — audit the other EventHandler hosts for the same instant-fatal pattern (they were quiet during #885 only because none of them restarted).
Repro/observe: stop the broker task; watch spikersoft-artpipe-model-safety restart-loop and Seq fill with the triplet every ~7s (first occurrence today 12:17:22Z, 2026-07-28).
Observed during the #885 broker outage: `SpikerSoft.EventHandlers.ArtPipeProcessor` (resident `artpipe-model-safety` service) attempts its RabbitMQ connection ~0s after boot, and a failure is fatal: `Failed to create RabbitMQ connection` → `Hosting failed to start` → exit 1. Swarm's `on-failure` restart (5s delay) turns that into a ~7s crash-loop — 3 error events per cycle, ~1,500+ Seq errors over the 2h outage, plus constant container churn on the GPU node.
Contrast: MassTransit consumers auto-recover and re-declare topology when the broker returns (per rabbitmq/docker-stack.yml deploy notes); this host's bespoke `RabbitMQ.Connection` bootstrap does not.
**Fix:** wrap broker connection creation at startup in retry-with-backoff (e.g. Polly exponential, cap ~30-60s, log Warning per attempt and Error only after N failures), or let the host start degraded and connect lazily with the client's built-in automatic recovery. Applies to whatever shares this bootstrap in `SpikerSoft.EventHandlers.Infrastructure` — audit the other EventHandler hosts for the same instant-fatal pattern (they were quiet during #885 only because none of them restarted).
Repro/observe: stop the broker task; watch `spikersoft-artpipe-model-safety` restart-loop and Seq fill with the triplet every ~7s (first occurrence today 12:17:22Z, 2026-07-28).
Re-verified against spikersoft-backendorigin/master @ 766cb8e4 (2026-07-29) — still unfixed, and the requested audit of the other EventHandler hosts is done. Notes updated so nobody has to re-derive this:
ArtPipeProcessor, confirmed:SpikerSoft.EventHandlers.ArtPipeProcessor/Services/ArtPipeStageConsumer.cs:133 calls factory.CreateConnectionAsync(...) with no retry wrapper. ExecuteAsync awaits it at :105 inside a try whose catch (Exception ex) logs "Fatal error in ArtPipeStageConsumer" and rethrows (:112) — so a broker-down boot is fatal exactly as described.
The AutomaticRecoveryEnabled = true / NetworkRecoveryInterval / TopologyRecoveryEnabled settings at :128 are a red herring for this failure mode: the RabbitMQ client only recovers a connection that was established at least once. It does nothing for a failure on the initialCreateConnectionAsync.
Audit result — the pattern is fleet-wide, not ArtPipeProcessor-specific. All 24SpikerSoft.EventHandlers.*/Services/*Consumer.cs hosts share it: every one creates its connection with zero retry/backoff, and every one carries the same "Fatal error in …" + throw; in ExecuteAsync. Full list:
The original note guessed the others "were quiet during #885 only because none of them restarted" — that is confirmed: they are equally fatal, they just happened not to be rescheduled. So the fix should land as a shared bootstrap helper in SpikerSoft.EventHandlers.Infrastructure applied to all 24, not a one-off patch to ArtPipeProcessor. The 2026-07-29 outage (#895) is the second occurrence of the same class.
Keeping this open — no implementing code exists yet.
Re-verified against `spikersoft-backend` `origin/master` @ 766cb8e4 (2026-07-29) — **still unfixed**, and the requested audit of the other EventHandler hosts is done. Notes updated so nobody has to re-derive this:
**ArtPipeProcessor, confirmed:** `SpikerSoft.EventHandlers.ArtPipeProcessor/Services/ArtPipeStageConsumer.cs:133` calls `factory.CreateConnectionAsync(...)` with no retry wrapper. `ExecuteAsync` awaits it at :105 inside a `try` whose `catch (Exception ex)` logs "Fatal error in ArtPipeStageConsumer" and rethrows (:112) — so a broker-down boot is fatal exactly as described.
The `AutomaticRecoveryEnabled = true` / `NetworkRecoveryInterval` / `TopologyRecoveryEnabled` settings at :128 are a red herring for this failure mode: the RabbitMQ client only recovers a connection that was established at least once. It does nothing for a failure on the *initial* `CreateConnectionAsync`.
**Audit result — the pattern is fleet-wide, not ArtPipeProcessor-specific.** All **24** `SpikerSoft.EventHandlers.*/Services/*Consumer.cs` hosts share it: every one creates its connection with zero retry/backoff, and every one carries the same `"Fatal error in …"` + `throw;` in `ExecuteAsync`. Full list:
ArtPipeStageConsumer, ArtStudioMetricsConsumer, BlogMediaMetadataConsumer, BlogMediaMoveConsumer, BlogMediaReceivedConsumer, BlogMediaScanConsumer, BlogMediaStripConsumer, BookManagementConsumer, FileMovementConsumer, GameEventsConsumer, LessonVideoMoveConsumer, LessonVideoReceivedConsumer, LessonVideoScanConsumer, LessonVideoTranscodeConsumer, CoverVariantsConsumer, MetadataExtractionConsumer, PhotographPipelineConsumer, ScheduledTaskConsumer, SecurityScanConsumer, BookCreatedConsumer, FileMovedConsumer, MetadataExtractedConsumer, ScanEventsConsumer, UploadReceivedConsumer.
The original note guessed the others "were quiet during #885 only because none of them restarted" — that is confirmed: they are equally fatal, they just happened not to be rescheduled. So the fix should land as a **shared bootstrap helper** in `SpikerSoft.EventHandlers.Infrastructure` applied to all 24, not a one-off patch to ArtPipeProcessor. The 2026-07-29 outage (#895) is the second occurrence of the same class.
Keeping this open — no implementing code exists yet.
Verified 2026-08-07 — Code:spikersoft-backend@9810202 — not fixed.SpikerSoft.EventHandlers.Infrastructure/Extensions/ServiceCollectionExtensions.cs, AddEventHandlerRabbitMQ (from :307): the connection is still built as a bare factory.CreateConnectionAsync().GetAwaiter().GetResult() at :340, with :346 logging Failed to create RabbitMQ connection and rethrowing — no Polly, no retry, no backoff anywhere in the method. SpikerSoft.Api/Extensions/ServiceCollectionExtensions.cs carries the same string. The health checks added at :438/:442 detect the condition but do not prevent the fatal bootstrap. Live: Reproduced again in today's umbrella #1000 outage: SpikerSoft.EventHandlers.ArtPipeProcessor logged Failed to create RabbitMQ connection every ~6s for the whole 7.5h broker outage (04:29-12:16Z). Broker is healthy now — curl https://api.spikersoft.com/healthz reports RabbitMQ Healthy, 267/267 queues, 78 consumers — so the crash-loop is quiescent but the defect is unchanged.
Status: not started — the bare CreateConnectionAsync bootstrap is unchanged, and it re-fired for 7.5h during today's outage
Closing here. Work now lives in the repo that holds the fix, so fixes #<N> in a PR will
auto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Migrated to **spikerj/spikersoft-backend#559** as part of the umbrella-tracker breakup.
Verified 2026-08-07 — **Code:** `spikersoft-backend@9810202` — **not fixed.** `SpikerSoft.EventHandlers.Infrastructure/Extensions/ServiceCollectionExtensions.cs`, `AddEventHandlerRabbitMQ` (from `:307`): the connection is still built as a bare `factory.CreateConnectionAsync().GetAwaiter().GetResult()` at `:340`, with `:346` logging `Failed to create RabbitMQ connection` and rethrowing — no Polly, no retry, no backoff anywhere in the method. `SpikerSoft.Api/Extensions/ServiceCollectionExtensions.cs` carries the same string. The health checks added at `:438`/`:442` detect the condition but do not prevent the fatal bootstrap. **Live:** Reproduced again in today's umbrella #1000 outage: `SpikerSoft.EventHandlers.ArtPipeProcessor` logged `Failed to create RabbitMQ connection` every ~6s for the whole 7.5h broker outage (04:29-12:16Z). Broker is healthy now — `curl https://api.spikersoft.com/healthz` reports RabbitMQ Healthy, 267/267 queues, 78 consumers — so the crash-loop is quiescent but the defect is unchanged.
Status: not started — the bare CreateConnectionAsync bootstrap is unchanged, and it re-fired for 7.5h during today's outage
Closing here. Work now lives in the repo that holds the fix, so `fixes #<N>` in a PR will
auto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Observed during the #885 broker outage:
SpikerSoft.EventHandlers.ArtPipeProcessor(residentartpipe-model-safetyservice) attempts its RabbitMQ connection ~0s after boot, and a failure is fatal:Failed to create RabbitMQ connection→Hosting failed to start→ exit 1. Swarm'son-failurerestart (5s delay) turns that into a ~7s crash-loop — 3 error events per cycle, ~1,500+ Seq errors over the 2h outage, plus constant container churn on the GPU node.Contrast: MassTransit consumers auto-recover and re-declare topology when the broker returns (per rabbitmq/docker-stack.yml deploy notes); this host's bespoke
RabbitMQ.Connectionbootstrap does not.Fix: wrap broker connection creation at startup in retry-with-backoff (e.g. Polly exponential, cap ~30-60s, log Warning per attempt and Error only after N failures), or let the host start degraded and connect lazily with the client's built-in automatic recovery. Applies to whatever shares this bootstrap in
SpikerSoft.EventHandlers.Infrastructure— audit the other EventHandler hosts for the same instant-fatal pattern (they were quiet during #885 only because none of them restarted).Repro/observe: stop the broker task; watch
spikersoft-artpipe-model-safetyrestart-loop and Seq fill with the triplet every ~7s (first occurrence today 12:17:22Z, 2026-07-28).Re-verified against
spikersoft-backendorigin/master@ 766cb8e4 (2026-07-29) — still unfixed, and the requested audit of the other EventHandler hosts is done. Notes updated so nobody has to re-derive this:ArtPipeProcessor, confirmed:
SpikerSoft.EventHandlers.ArtPipeProcessor/Services/ArtPipeStageConsumer.cs:133callsfactory.CreateConnectionAsync(...)with no retry wrapper.ExecuteAsyncawaits it at :105 inside atrywhosecatch (Exception ex)logs "Fatal error in ArtPipeStageConsumer" and rethrows (:112) — so a broker-down boot is fatal exactly as described.The
AutomaticRecoveryEnabled = true/NetworkRecoveryInterval/TopologyRecoveryEnabledsettings at :128 are a red herring for this failure mode: the RabbitMQ client only recovers a connection that was established at least once. It does nothing for a failure on the initialCreateConnectionAsync.Audit result — the pattern is fleet-wide, not ArtPipeProcessor-specific. All 24
SpikerSoft.EventHandlers.*/Services/*Consumer.cshosts share it: every one creates its connection with zero retry/backoff, and every one carries the same"Fatal error in …"+throw;inExecuteAsync. Full list:ArtPipeStageConsumer, ArtStudioMetricsConsumer, BlogMediaMetadataConsumer, BlogMediaMoveConsumer, BlogMediaReceivedConsumer, BlogMediaScanConsumer, BlogMediaStripConsumer, BookManagementConsumer, FileMovementConsumer, GameEventsConsumer, LessonVideoMoveConsumer, LessonVideoReceivedConsumer, LessonVideoScanConsumer, LessonVideoTranscodeConsumer, CoverVariantsConsumer, MetadataExtractionConsumer, PhotographPipelineConsumer, ScheduledTaskConsumer, SecurityScanConsumer, BookCreatedConsumer, FileMovedConsumer, MetadataExtractedConsumer, ScanEventsConsumer, UploadReceivedConsumer.
The original note guessed the others "were quiet during #885 only because none of them restarted" — that is confirmed: they are equally fatal, they just happened not to be rescheduled. So the fix should land as a shared bootstrap helper in
SpikerSoft.EventHandlers.Infrastructureapplied to all 24, not a one-off patch to ArtPipeProcessor. The 2026-07-29 outage (#895) is the second occurrence of the same class.Keeping this open — no implementing code exists yet.
Migrated to spikerj/spikersoft-backend#559 as part of the umbrella-tracker breakup.
Verified 2026-08-07 — Code:
spikersoft-backend@9810202— not fixed.SpikerSoft.EventHandlers.Infrastructure/Extensions/ServiceCollectionExtensions.cs,AddEventHandlerRabbitMQ(from:307): the connection is still built as a barefactory.CreateConnectionAsync().GetAwaiter().GetResult()at:340, with:346loggingFailed to create RabbitMQ connectionand rethrowing — no Polly, no retry, no backoff anywhere in the method.SpikerSoft.Api/Extensions/ServiceCollectionExtensions.cscarries the same string. The health checks added at:438/:442detect the condition but do not prevent the fatal bootstrap. Live: Reproduced again in today's umbrella #1000 outage:SpikerSoft.EventHandlers.ArtPipeProcessorloggedFailed to create RabbitMQ connectionevery ~6s for the whole 7.5h broker outage (04:29-12:16Z). Broker is healthy now —curl https://api.spikersoft.com/healthzreports RabbitMQ Healthy, 267/267 queues, 78 consumers — so the crash-loop is quiescent but the defect is unchanged.Status: not started — the bare CreateConnectionAsync bootstrap is unchanged, and it re-fired for 7.5h during today's outage
Closing here. Work now lives in the repo that holds the fix, so
fixes #<N>in a PR willauto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent