H2: Persist runtime capability cache to DB so server restarts don't wipe it
Summary
H2 of the 1.8.0 Upgrade Hardening initiative (8 PRs total — H1 already merged in #354).
Fixes the operator-facing 1.7.x bug where every server pod restart wiped the in-memory MachineRuntimeCapabilities cache, causing operators to hit the H1 NoOsDetected path on first action until the next scheduled health check populated the cache.
Behaviour changes
| Lifecycle event | Pre-H2 | Post-H2 |
|---|---|---|
| Successful Capabilities probe | Update in-memory cache | Update in-memory cache + persist to machine.runtime_capabilities_json
|
| Server pod restart | Cache empty until first health check per machine | Hydrator (IStartable) reads all DB rows + populates cache at boot |
| Successful upgrade | In-memory invalidate | In-memory invalidate + NULL out DB column (next health check overwrites) |
| Pre-existing machine row (NULL column) | n/a | Treated as cold — falls to H1's NoOsDetected with health-check hint |
| DB write failure | n/a | Logged warning; in-memory cache still fresh; next probe retries |
Schema migration
ALTER TABLE machine
ADD COLUMN runtime_capabilities_json jsonb NULL,
ADD COLUMN runtime_capabilities_updated_at timestamptz NULL;
CREATE INDEX ix_machine_runtime_capabilities_updated_at
ON machine USING btree (runtime_capabilities_updated_at DESC)
WHERE runtime_capabilities_updated_at IS NOT NULL;
runtime_capabilities_updated_at is unused in H2 — sets the foundation for a future TTL-invalidation step.
JSON shape (wire contract — pinned by unit test)
{
"os": "Microsoft Windows NT 10.0.19045.0",
"osVersion": "10.0.19045.0",
"defaultShell": "powershell",
"installedShells": "powershell,cmd",
"architecture": "X64",
"agentVersion": "1.8.0",
"supportedServices": ["IScriptService/v1"]
}
camelCase per existing API convention. New optional fields can be appended; renaming/case-changing breaks the hydrate round-trip on operator's existing rows.
Failure isolation
- DB write fails on health check → in-memory cache still fresh; warning logged; next successful probe retries
-
Hydrator fails at startup → server starts with empty cache; H1
NoOsDetectedpath correctly directs operators to health check; doesn't block boot - One corrupt JSON row → log + skip during hydration; remaining rows hydrate normally
-
IMachineRuntimeCapabilitiesPersistence == nullin constructor (test fixtures, etc.) → write-through is a no-op; in-memory cache works as before
Test plan
-
Pin canonical JSON shape (5 unit tests in MachineRuntimeCapabilitiesPersistenceTests) -
Round-trip → deserialise from canonical literal → property values match -
Backward-compat: missing optional fields deserialise to null -
SerializerOptions.PropertyNamingPolicy = CamelCaseliteral pin -
5545/5545unit tests green (vs5540/5540baseline →+5net new) -
Integration test for full Postgres roundtrip — deferred to H8 (avoid duplicating DB roundtrip coverage; H8's E2E matrix includes "server pod restart between dispatch + status poll" which exercises this end-to-end)
What's NOT in this PR
- H3: Active health-check probe API (UI button currently passive)
- H4: Upgrade lock TTL + idempotency (cures the "repeated clicks compound" failure)
- H5: Cross-OS unified upgrade pipeline
- H6: Agent rollback + abandonment recovery
- H7: Role/feature capability slots
- H8: Comprehensive E2E test matrix (includes the integration coverage referenced above)