Skip to content

H2: Persist runtime capability cache to DB so server restarts don't wipe it

Summary

H2 of the 1.8.0 Upgrade Hardening initiative (8 PRs total — H1 already merged in #354).

Fixes the operator-facing 1.7.x bug where every server pod restart wiped the in-memory MachineRuntimeCapabilities cache, causing operators to hit the H1 NoOsDetected path on first action until the next scheduled health check populated the cache.

Behaviour changes

Lifecycle event Pre-H2 Post-H2
Successful Capabilities probe Update in-memory cache Update in-memory cache + persist to machine.runtime_capabilities_json
Server pod restart Cache empty until first health check per machine Hydrator (IStartable) reads all DB rows + populates cache at boot
Successful upgrade In-memory invalidate In-memory invalidate + NULL out DB column (next health check overwrites)
Pre-existing machine row (NULL column) n/a Treated as cold — falls to H1's NoOsDetected with health-check hint
DB write failure n/a Logged warning; in-memory cache still fresh; next probe retries

Schema migration

ALTER TABLE machine
    ADD COLUMN runtime_capabilities_json     jsonb       NULL,
    ADD COLUMN runtime_capabilities_updated_at timestamptz NULL;
CREATE INDEX ix_machine_runtime_capabilities_updated_at
    ON machine USING btree (runtime_capabilities_updated_at DESC)
    WHERE runtime_capabilities_updated_at IS NOT NULL;

runtime_capabilities_updated_at is unused in H2 — sets the foundation for a future TTL-invalidation step.

JSON shape (wire contract — pinned by unit test)

{
  "os": "Microsoft Windows NT 10.0.19045.0",
  "osVersion": "10.0.19045.0",
  "defaultShell": "powershell",
  "installedShells": "powershell,cmd",
  "architecture": "X64",
  "agentVersion": "1.8.0",
  "supportedServices": ["IScriptService/v1"]
}

camelCase per existing API convention. New optional fields can be appended; renaming/case-changing breaks the hydrate round-trip on operator's existing rows.

Failure isolation

  • DB write fails on health check → in-memory cache still fresh; warning logged; next successful probe retries
  • Hydrator fails at startup → server starts with empty cache; H1 NoOsDetected path correctly directs operators to health check; doesn't block boot
  • One corrupt JSON row → log + skip during hydration; remaining rows hydrate normally
  • IMachineRuntimeCapabilitiesPersistence == null in constructor (test fixtures, etc.) → write-through is a no-op; in-memory cache works as before

Test plan

  • Pin canonical JSON shape (5 unit tests in MachineRuntimeCapabilitiesPersistenceTests)
  • Round-trip → deserialise from canonical literal → property values match
  • Backward-compat: missing optional fields deserialise to null
  • SerializerOptions.PropertyNamingPolicy = CamelCase literal pin
  • 5545/5545 unit tests green (vs 5540/5540 baseline → +5 net new)
  • Integration test for full Postgres roundtrip — deferred to H8 (avoid duplicating DB roundtrip coverage; H8's E2E matrix includes "server pod restart between dispatch + status poll" which exercises this end-to-end)

What's NOT in this PR

  • H3: Active health-check probe API (UI button currently passive)
  • H4: Upgrade lock TTL + idempotency (cures the "repeated clicks compound" failure)
  • H5: Cross-OS unified upgrade pipeline
  • H6: Agent rollback + abandonment recovery
  • H7: Role/feature capability slots
  • H8: Comprehensive E2E test matrix (includes the integration coverage referenced above)

Merge request reports

Loading