Skip to content

Persist terminal upgrade trace to survive server restarts

Placeholder ppxd requested to merge feat/durable-upgrade-trace into main

Summary

  • The per-machine upgrade timeline (status + event log + Phase B log) lived only in a process-local in-memory cache, so a server pod restart erased an operator's view of how the most recent upgrade concluded until the next Capabilities probe re-populated it.
  • That cache is intentionally in-memory — writing it on every probe (every few seconds during an active upgrade) would dominate the cost of the upgrade. This adds a durable backstop that writes the trace exactly once per upgrade, only when the agent first reports a terminal status (SUCCESS/FAILED/ROLLED_BACK/ROLLBACK_NEEDED/ROLLBACK_CRITICAL_FAILED).
  • A dedup gate (IUpgradeTracePersistenceGate) suppresses the repeated writes that would otherwise fire on every subsequent probe (the agent keeps re-reporting the same terminal status until the next upgrade). The snapshot is stored as jsonb on the machine row and hydrated back into the in-memory cache at startup, so reads stay in-memory and the three read handlers are untouched.
  • Terminal classification is drift-resistant: a status is terminal when it is non-empty and not one of the small, stable in-flight set (IN_PROGRESS/SWAPPED/ROLLING_BACK), so a future agent's new terminal outcome is persisted without a server change.

Design notes

  • Mirrors the existing runtime-capabilities persistence pattern exactly: a scoped IScopedDependency persistence port (UpgradeTracePersistence, atomic ExecuteUpdateAsync), a startup IStartable hydrator (UpgradeTraceHydrator, which also primes the dedup gate to avoid a post-restart write storm), and jsonb columns on the machine table (MachineConfiguration maps the string property to jsonb).
  • Fully additive / non-breaking: two optional collaborators appended to TentacleHealthCheckStrategy, a new migration, a new column. No existing interface, read path, or call site changes. A DB hiccup is best-effort — it logs and never fails the health check, leaving the gate open to retry on the next probe.

Test plan

  • Unit — UpgradeStatusClassifierTests (terminal vs in-flight, empty, unknown-as-terminal, ordinal contract, pinned in-flight constants)
  • Unit — UpgradeTracePersistenceGateTests (dedup once-per-signature, new-signature resets, per-machine isolation)
  • Unit — UpgradeTracePersistenceShapeTests (jsonb shape pin, round-trip, canonical-literal back-compat, missing-section defaults, serializer options pinned)
  • Unit — UpgradeTraceHydratorTests (store hydration + gate priming + null-row skip)
  • Unit — TentacleHealthCheckStrategyTests (terminal persists once with status/events/log; in-flight skips; same outcome across 3 probes persists once; persistence throws → still healthy + gate not marked; optional deps unwired → no NRE)
  • Integration (real Postgres) — UpgradeTracePersistenceIntegrationTests: save writes column + timestamp, LoadAllAsync round-trips, latest-wins on re-run, NULL rows skipped, undeserialisable (wrong-shape jsonb) row skipped, and an end-to-end simulated-restart durability proof (persist → fresh empty store → hydrate from DB → status + events + log restored)
  • Full unit suite: 5805/5805 green
  • Full solution build: 0 errors

Merge request reports

Loading