Skip to content

H6: Distinguish 'RolledBack' from 'Failed' in upgrade outcome

Placeholder ppxd requested to merge feat/upgrade-hardening-h6-agent-rollback into main

Summary

H6 of the 1.8.0 Upgrade Hardening initiative (H1-H5 already merged).

Pre-H6, when an upgrade failed its post-install health check and the agent-side install script's rollback path successfully restored the previous binary, the server reported MachineUpgradeStatus.Failed — indistinguishable from "upgrade left machine broken". Operators couldn't tell whether their machine was:

  • Actually broken → red badge, investigate immediately
  • Safely rolled back to baseline → yellow/amber badge, OK to retry once root cause is understood

Both install scripts ALREADY implement comprehensive rollback (.bak/.previous backup → atomic swap → post-restart health check → on failure, restore baseline → write ROLLED_BACK status). H6 just exposes the outcome to the server-facing API.

Behaviour change

Scenario Pre-H6 server response Post-H6 server response
Binary swap succeeded, agent healthy on new version Upgraded Upgraded (unchanged)
Binary swap committed but post-install health check failed AND rollback restored baseline AND baseline is healthy Failed ("Upgrade script failed exit 4. Last log: rollback succeeded") RolledBack ("Upgrade to X failed post-install health check; rolled back to the previous binary. Agent is healthy on the baseline version — safe to retry once root cause is understood.")
Binary swap committed, agent broken on new version, rollback ALSO failed Failed (correct) Failed (unchanged — ROLLBACK_CRITICAL_FAILED is genuinely broken)
Halibut RPC failed before script ran Failed (correct) Failed (unchanged)

Why Windows needs no code change

The Windows strategy's outer Halibut wrapper exits 0 after scheduling the detached Task Scheduler task. The actual rollback outcome reaches the server via the agent's last-upgrade.json on the next health check, surfaced verbatim through /api/machines/{id}/upgrade-status. The existing ROLLED_BACK status string from the Windows script is already operator-visible there.

The Linux strategy IS the only place that needs InterpretScriptResult adjustment, because its outer Halibut script completes synchronously with the rollback outcome (exit code 4).

Wire contract pin

  • New MachineUpgradeStatus.RolledBack = 5 — appended at end per wire-stability convention
  • 13-row integrity test pins both numeric ordinals AND name literals for all 6 status values + a total-count drift detector
  • LinuxTentacleUpgradeStrategy.LinuxExitRolledBack constant (= 4) pinned to the script's exit 4 contract — if either side changes, the test fails

Test plan

  • +13 wire-stability pins on MachineUpgradeStatus enum (6 ordinal + 6 name + 1 count)
  • +1 strategy test: exit 4 → RolledBack with operator-actionable Detail
  • +1 constant pin: LinuxExitRolledBack == 4
  • Existing ScriptNonZeroExit_ReturnsFailedWithLastLogLine still covers non-rollback failures (preserves regression coverage)
  • 5578/5578 unit tests green (+15 net new)
  • E2E: deferred to H8 (E2E matrix includes rollback-triggered-by-failing-binary scenarios)

Operator UX after H6

FE can render three distinct upgrade-status badges based on MachineUpgradeStatus:

  • 🟢 Upgraded / AlreadyUpToDate — green, no action
  • 🟡 RolledBack — yellow, "Safely restored to baseline. Review upgrade.log + last-upgrade.json before retry."
  • 🔴 Failed / NotSupported — red, "Investigation needed."
  • 🔵 Initiated — blue, "In progress, check back."

What's NOT in this PR

  • H7: Role/feature capability slots (next — catches "IIS not installed" at plan-time)
  • H8: Comprehensive E2E test matrix (validates the full upgrade chain end-to-end)

Merge request reports

Loading