The Keycloak PostgreSQL database has a corrupted TOAST value in pg_toast_2619 (the TOAST table backing the pg_statistic planner-statistics catalog). Every query whose plan needs the affected column statistics fails, and Keycloak's ClearExpiredEvents scheduled job fails on every run — so expired Keycloak events are no longer being purged (unbounded EVENT_ENTITY growth).
Symptoms
Postgres side (keycloak_postgres-keycloak, currently on dreamstream4), recurring ~every 60s:
ERROR: missing chunk number 0 for toast value 131293 in pg_toast_2619
Keycloak side (keycloak_keycloak), ~every 7.5 min:
ERROR [org.keycloak.services.scheduled.ScheduledTaskRunner] Failed to run scheduled task ClearExpiredEvents:
org.hibernate.exception.GenericJDBCException: ... [ERROR: missing chunk number 0 for toast value 131293 in pg_toast_2619]
Scope observed over 60 min: a single toast value (131293), single toast table (pg_toast_2619 = pg_statistic), single failing task (ClearExpiredEvents).
What this means
pg_toast_2619 is the TOAST relation for the system catalog pg_statistic. The corruption is in cached planner statistics, not in real Keycloak user/realm data — which makes it safe to remediate by regenerating statistics. But the planner keeps reading the dead toast pointer, so any statement that touches those stats errors out until the stale row is cleared.
Impact
ClearExpiredEvents never completes -> EVENT_ENTITY grows without bound (DB bloat over time).
Any planner path needing the corrupt column stats can fail intermittently.
Proposed remediation (take a backup first)
Because the corruption is confined to pg_statistic, regenerating statistics clears it:
-- as a superuser, against the keycloak DB
-- (optional targeted: ANALYZE each table until the erroring one is found, which overwrites its stat rows)
-- reliable blanket fix:
DELETEFROMpg_statistic;-- may require: SET allow_system_table_mods = on;
ANALYZE;-- regenerates all statistics from scratch
Then confirm the missing chunk errors stop and ClearExpiredEvents succeeds.
Root-cause follow-up
TOAST corruption usually traces to an unclean stop or a storage-layer problem:
Check dmesg / SMART on dreamstream4 (and wherever this PG volume has run) for disk errors.
Verify whether data_checksums is enabled (SHOW data_checksums;) — if off, consider enabling on the next rebuild so future corruption is detected early.
Ensure the Keycloak Postgres has a working backup (relates to the SonarQube-backups gap in a separate issue — worth auditing all DB backup jobs).
Filed proactively by automated swarm health check (docker service log audit).
## Summary
The Keycloak PostgreSQL database has a **corrupted TOAST value** in `pg_toast_2619` (the TOAST table backing the `pg_statistic` planner-statistics catalog). Every query whose plan needs the affected column statistics fails, and Keycloak's `ClearExpiredEvents` scheduled job fails on every run — so **expired Keycloak events are no longer being purged** (unbounded `EVENT_ENTITY` growth).
## Symptoms
Postgres side (`keycloak_postgres-keycloak`, currently on dreamstream4), recurring ~every 60s:
```
ERROR: missing chunk number 0 for toast value 131293 in pg_toast_2619
```
Keycloak side (`keycloak_keycloak`), ~every 7.5 min:
```
ERROR [org.keycloak.services.scheduled.ScheduledTaskRunner] Failed to run scheduled task ClearExpiredEvents:
org.hibernate.exception.GenericJDBCException: ... [ERROR: missing chunk number 0 for toast value 131293 in pg_toast_2619]
```
Scope observed over 60 min: a single toast value (`131293`), single toast table (`pg_toast_2619` = `pg_statistic`), single failing task (`ClearExpiredEvents`).
## What this means
`pg_toast_2619` is the TOAST relation for the system catalog `pg_statistic`. The corruption is in **cached planner statistics**, not in real Keycloak user/realm data — which makes it safe to remediate by regenerating statistics. But the planner keeps reading the dead toast pointer, so any statement that touches those stats errors out until the stale row is cleared.
## Impact
- `ClearExpiredEvents` never completes -> `EVENT_ENTITY` grows without bound (DB bloat over time).
- Any planner path needing the corrupt column stats can fail intermittently.
## Proposed remediation (take a backup first)
Because the corruption is confined to `pg_statistic`, regenerating statistics clears it:
```sql
-- as a superuser, against the keycloak DB
-- (optional targeted: ANALYZE each table until the erroring one is found, which overwrites its stat rows)
-- reliable blanket fix:
DELETE FROM pg_statistic; -- may require: SET allow_system_table_mods = on;
ANALYZE; -- regenerates all statistics from scratch
```
Then confirm the `missing chunk` errors stop and `ClearExpiredEvents` succeeds.
## Root-cause follow-up
TOAST corruption usually traces to an unclean stop or a storage-layer problem:
- Check dmesg / SMART on **dreamstream4** (and wherever this PG volume has run) for disk errors.
- Verify whether `data_checksums` is enabled (`SHOW data_checksums;`) — if `off`, consider enabling on the next rebuild so future corruption is detected early.
- Ensure the Keycloak Postgres has a working backup (relates to the SonarQube-backups gap in a separate issue — worth auditing all DB backup jobs).
---
_Filed proactively by automated swarm health check (docker service log audit)._
Recovery runbook up: infra PR #49 → docs/keycloak-postgres-toast-recovery.md. Summary of the approach (full detail + exact commands in the doc):
Blast-radius gate: SELECT count(*) FROM pg_statistic (reproduces) vs SELECT count(*) FROM event_entity (must be clean). If application tables ALSO throw toast errors → STOP, basebackup, escalate — that would be data corruption, not this runbook.
Fix (~2 min, Keycloak stays up): optional row-by-row identifier for a targeted DELETE FROM pg_statistic WHERE ctid=..., or the equally-safe DELETE FROM pg_statistic; — it's planner-stats cache, fully regenerated by the following VACUUM pg_statistic; ANALYZE;.
Verify: the ~60s missing chunk stream → 0; next ClearExpiredEvents (≤7.5 min) completes; event_entity count shrinks once the purge lands.
Root-cause tail: SHOW data_checksums; (if off, consider checksums at next maintenance — silent-corruption insurance on consumer hardware) + the #503 bad-RAM connection: if this recurs on dreamstream4, memtest it before trusting it with the identity DB.
Will close on your confirmation of step 4's three green checks.
Recovery runbook up: **infra PR #49** → `docs/keycloak-postgres-toast-recovery.md`. Summary of the approach (full detail + exact commands in the doc):
1. **Blast-radius gate**: `SELECT count(*) FROM pg_statistic` (reproduces) vs `SELECT count(*) FROM event_entity` (must be clean). If application tables ALSO throw toast errors → STOP, basebackup, escalate — that would be data corruption, not this runbook.
2. **Fix** (~2 min, Keycloak stays up): optional row-by-row identifier for a targeted `DELETE FROM pg_statistic WHERE ctid=...`, or the equally-safe `DELETE FROM pg_statistic;` — it's planner-stats cache, fully regenerated by the following `VACUUM pg_statistic; ANALYZE;`.
3. **Verify**: the ~60s `missing chunk` stream → 0; next `ClearExpiredEvents` (≤7.5 min) completes; `event_entity` count shrinks once the purge lands.
4. **Root-cause tail**: `SHOW data_checksums;` (if off, consider checksums at next maintenance — silent-corruption insurance on consumer hardware) + the #503 bad-RAM connection: if this recurs on dreamstream4, memtest it before trusting it with the identity DB.
Will close on your confirmation of step 4's three green checks.
Audited against origin/master — the runbook shipped; the repair was never run. This ticket is marked resolved by association with infra PR #49, but that PR merged the documentation, not the fix.
Merged:docs/keycloak-postgres-toast-recovery.md (8ee2a8d, PR #49) — blast-radius gate, DELETE FROM pg_statistic; VACUUM; ANALYZE, verification steps, and the data_checksums root-cause tail.
Not run. The direct evidence is #561, filed 2026-07-14 and still open, which reports the identical error — missing chunk number 0 for toast value 131293 in pg_toast_2619 — still recurring every 15-minute cycle after this ticket's remediation was written, and states plainly that the repair was never executed. Corroborating: git grep -rn 'pg_statistic|toast' origin/master outside the runbook finds nothing — no execution record, no post-repair note anywhere in any repo.
So the state is: we know exactly how to fix this, we wrote it down, and the corruption is still live.
Remaining:
Execute the runbook against the Keycloak database.
Confirm the ~60-second error stream stops.
Confirm ClearExpiredEvents completes and EVENT_ENTITY actually drains.
Run the SHOW data_checksums; tail to settle the root cause — worth doing, because #503 established that this host has confirmed bad RAM, which is a plausible common cause for page-level corruption.
Given #561 tracks the same live corruption, one of these two should probably absorb the other once the repair runs.
Audited against `origin/master` — **the runbook shipped; the repair was never run.** This ticket is marked resolved by association with infra PR #49, but that PR merged the *documentation*, not the fix.
**Merged:** `docs/keycloak-postgres-toast-recovery.md` (`8ee2a8d`, PR #49) — blast-radius gate, `DELETE FROM pg_statistic; VACUUM; ANALYZE`, verification steps, and the `data_checksums` root-cause tail.
**Not run.** The direct evidence is **#561**, filed 2026-07-14 and still open, which reports the *identical* error — `missing chunk number 0 for toast value 131293 in pg_toast_2619` — still recurring every 15-minute cycle **after** this ticket's remediation was written, and states plainly that the repair was never executed. Corroborating: `git grep -rn 'pg_statistic|toast' origin/master` outside the runbook finds nothing — no execution record, no post-repair note anywhere in any repo.
So the state is: we know exactly how to fix this, we wrote it down, and the corruption is still live.
**Remaining:**
1. Execute the runbook against the Keycloak database.
2. Confirm the ~60-second error stream stops.
3. Confirm `ClearExpiredEvents` completes and `EVENT_ENTITY` actually drains.
4. Run the `SHOW data_checksums;` tail to settle the root cause — worth doing, because #503 established that this host has confirmed bad RAM, which is a plausible common cause for page-level corruption.
Given #561 tracks the same live corruption, one of these two should probably absorb the other once the repair runs.
So the pg_toast_2619 errors are still recurring and Keycloak's scheduled task is still erroring ~every 100s, meaning EVENT_ENTITY is still growing unbounded. #561 (which says exactly this — that it persisted after #552 was closed) is the same live condition. Consider closing one of the two as a duplicate and keeping a single ticket for the actual VACUUM FULL / statistics-rebuild execution.
— 2026-08-06 tracker sweep, Opus 5 Agent. Staying open.
**Still failing — re-verified live 2026-08-06.**
The recovery runbook landed (spikersoft-infrastructure#49) but was never executed, or did not hold:
```
docker service logs --since 60m keycloak_postgres-keycloak | grep -ci 'toast|missing chunk' -> 4
docker service logs --since 120m keycloak_keycloak | grep -ci 'ClearExpiredEvents|ScheduledTaskRunner' -> 72
```
So the `pg_toast_2619` errors are still recurring and Keycloak's scheduled task is still erroring ~every 100s, meaning `EVENT_ENTITY` is still growing unbounded. #561 (which says exactly this — that it persisted after #552 was closed) is the same live condition. Consider closing one of the two as a duplicate and keeping a single ticket for the actual `VACUUM FULL` / statistics-rebuild execution.
— 2026-08-06 tracker sweep, Opus 5 Agent. Staying open.
Verified 2026-08-07 live: 4 missing chunk/toast errors in 60m of keycloak_postgres-keycloak, 72 ClearExpiredEvents/ScheduledTaskRunner lines in 120m of keycloak_keycloak, latest 13:42:59Z on dreamstream7 naming the same toast value 131293. Code: spikersoft-infrastructure@86d03ffdocs/keycloak-postgres-toast-recovery.md (commit 8ee2a8d) is the only artifact.
Status: not started — the runbook merged (infra PR #49) but has never been executed; corruption confirmed still live today.
Closing here. Work now lives in the repo that holds the fix, so fixes #194 in a PR will auto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Migrated to **spikerj/spikersoft-infrastructure#194** as part of the umbrella-tracker breakup.
Verified 2026-08-07 live: 4 `missing chunk`/toast errors in 60m of `keycloak_postgres-keycloak`, 72 `ClearExpiredEvents`/ScheduledTaskRunner lines in 120m of `keycloak_keycloak`, latest 13:42:59Z on dreamstream7 naming the same toast value 131293. Code: `spikersoft-infrastructure@86d03ff` `docs/keycloak-postgres-toast-recovery.md` (commit 8ee2a8d) is the only artifact.
Status: not started — the runbook merged (infra PR #49) but has never been executed; corruption confirmed still live today.
Closing here. Work now lives in the repo that holds the fix, so `fixes #194` in a PR will auto-close it on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Summary
The Keycloak PostgreSQL database has a corrupted TOAST value in
pg_toast_2619(the TOAST table backing thepg_statisticplanner-statistics catalog). Every query whose plan needs the affected column statistics fails, and Keycloak'sClearExpiredEventsscheduled job fails on every run — so expired Keycloak events are no longer being purged (unboundedEVENT_ENTITYgrowth).Symptoms
Postgres side (
keycloak_postgres-keycloak, currently on dreamstream4), recurring ~every 60s:Keycloak side (
keycloak_keycloak), ~every 7.5 min:Scope observed over 60 min: a single toast value (
131293), single toast table (pg_toast_2619=pg_statistic), single failing task (ClearExpiredEvents).What this means
pg_toast_2619is the TOAST relation for the system catalogpg_statistic. The corruption is in cached planner statistics, not in real Keycloak user/realm data — which makes it safe to remediate by regenerating statistics. But the planner keeps reading the dead toast pointer, so any statement that touches those stats errors out until the stale row is cleared.Impact
ClearExpiredEventsnever completes ->EVENT_ENTITYgrows without bound (DB bloat over time).Proposed remediation (take a backup first)
Because the corruption is confined to
pg_statistic, regenerating statistics clears it:Then confirm the
missing chunkerrors stop andClearExpiredEventssucceeds.Root-cause follow-up
TOAST corruption usually traces to an unclean stop or a storage-layer problem:
data_checksumsis enabled (SHOW data_checksums;) — ifoff, consider enabling on the next rebuild so future corruption is detected early.Filed proactively by automated swarm health check (docker service log audit).
Recovery runbook up: infra PR #49 →
docs/keycloak-postgres-toast-recovery.md. Summary of the approach (full detail + exact commands in the doc):SELECT count(*) FROM pg_statistic(reproduces) vsSELECT count(*) FROM event_entity(must be clean). If application tables ALSO throw toast errors → STOP, basebackup, escalate — that would be data corruption, not this runbook.DELETE FROM pg_statistic WHERE ctid=..., or the equally-safeDELETE FROM pg_statistic;— it's planner-stats cache, fully regenerated by the followingVACUUM pg_statistic; ANALYZE;.missing chunkstream → 0; nextClearExpiredEvents(≤7.5 min) completes;event_entitycount shrinks once the purge lands.SHOW data_checksums;(if off, consider checksums at next maintenance — silent-corruption insurance on consumer hardware) + the #503 bad-RAM connection: if this recurs on dreamstream4, memtest it before trusting it with the identity DB.Will close on your confirmation of step 4's three green checks.
Board-sweep status (2026-07-22): recovery RUNBOOK merged (infra #49). REMAINING: the actual TOAST repair + ClearExpiredEvents verification — an unexecuted ops procedure.
Audited against
origin/master— the runbook shipped; the repair was never run. This ticket is marked resolved by association with infra PR #49, but that PR merged the documentation, not the fix.Merged:
docs/keycloak-postgres-toast-recovery.md(8ee2a8d, PR #49) — blast-radius gate,DELETE FROM pg_statistic; VACUUM; ANALYZE, verification steps, and thedata_checksumsroot-cause tail.Not run. The direct evidence is #561, filed 2026-07-14 and still open, which reports the identical error —
missing chunk number 0 for toast value 131293 in pg_toast_2619— still recurring every 15-minute cycle after this ticket's remediation was written, and states plainly that the repair was never executed. Corroborating:git grep -rn 'pg_statistic|toast' origin/masteroutside the runbook finds nothing — no execution record, no post-repair note anywhere in any repo.So the state is: we know exactly how to fix this, we wrote it down, and the corruption is still live.
Remaining:
ClearExpiredEventscompletes andEVENT_ENTITYactually drains.SHOW data_checksums;tail to settle the root cause — worth doing, because #503 established that this host has confirmed bad RAM, which is a plausible common cause for page-level corruption.Given #561 tracks the same live corruption, one of these two should probably absorb the other once the repair runs.
Still failing — re-verified live 2026-08-06.
The recovery runbook landed (spikersoft-infrastructure#49) but was never executed, or did not hold:
So the
pg_toast_2619errors are still recurring and Keycloak's scheduled task is still erroring ~every 100s, meaningEVENT_ENTITYis still growing unbounded. #561 (which says exactly this — that it persisted after #552 was closed) is the same live condition. Consider closing one of the two as a duplicate and keeping a single ticket for the actualVACUUM FULL/ statistics-rebuild execution.— 2026-08-06 tracker sweep, Opus 5 Agent. Staying open.
Migrated to spikerj/spikersoft-infrastructure#194 as part of the umbrella-tracker breakup.
Verified 2026-08-07 live: 4
missing chunk/toast errors in 60m ofkeycloak_postgres-keycloak, 72ClearExpiredEvents/ScheduledTaskRunner lines in 120m ofkeycloak_keycloak, latest 13:42:59Z on dreamstream7 naming the same toast value 131293. Code:spikersoft-infrastructure@86d03ffdocs/keycloak-postgres-toast-recovery.md(commit 8ee2a8d) is the only artifact.Status: not started — the runbook merged (infra PR #49) but has never been executed; corruption confirmed still live today.
Closing here. Work now lives in the repo that holds the fix, so
fixes #194in a PR will auto-close it on merge. The umbrella tracker keeps cross-repo epics only.— Opus 5 Agent