H6: Distinguish 'RolledBack' from 'Failed' in upgrade outcome
Summary
H6 of the 1.8.0 Upgrade Hardening initiative (H1-H5 already merged).
Pre-H6, when an upgrade failed its post-install health check and the agent-side install script's rollback path successfully restored the previous binary, the server reported MachineUpgradeStatus.Failed — indistinguishable from "upgrade left machine broken". Operators couldn't tell whether their machine was:
- Actually broken → red badge, investigate immediately
- Safely rolled back to baseline → yellow/amber badge, OK to retry once root cause is understood
Both install scripts ALREADY implement comprehensive rollback (.bak/.previous backup → atomic swap → post-restart health check → on failure, restore baseline → write ROLLED_BACK status). H6 just exposes the outcome to the server-facing API.
Behaviour change
| Scenario | Pre-H6 server response | Post-H6 server response |
|---|---|---|
| Binary swap succeeded, agent healthy on new version | Upgraded |
Upgraded (unchanged) |
| Binary swap committed but post-install health check failed AND rollback restored baseline AND baseline is healthy |
Failed ("Upgrade script failed exit 4. Last log: rollback succeeded") |
RolledBack ("Upgrade to X failed post-install health check; rolled back to the previous binary. Agent is healthy on the baseline version — safe to retry once root cause is understood.") |
| Binary swap committed, agent broken on new version, rollback ALSO failed |
Failed (correct) |
Failed (unchanged — ROLLBACK_CRITICAL_FAILED is genuinely broken) |
| Halibut RPC failed before script ran |
Failed (correct) |
Failed (unchanged) |
Why Windows needs no code change
The Windows strategy's outer Halibut wrapper exits 0 after scheduling the detached Task Scheduler task. The actual rollback outcome reaches the server via the agent's last-upgrade.json on the next health check, surfaced verbatim through /api/machines/{id}/upgrade-status. The existing ROLLED_BACK status string from the Windows script is already operator-visible there.
The Linux strategy IS the only place that needs InterpretScriptResult adjustment, because its outer Halibut script completes synchronously with the rollback outcome (exit code 4).
Wire contract pin
- New
MachineUpgradeStatus.RolledBack = 5— appended at end per wire-stability convention - 13-row integrity test pins both numeric ordinals AND name literals for all 6 status values + a total-count drift detector
-
LinuxTentacleUpgradeStrategy.LinuxExitRolledBackconstant (= 4) pinned to the script'sexit 4contract — if either side changes, the test fails
Test plan
-
+13 wire-stability pins on MachineUpgradeStatusenum (6 ordinal + 6 name + 1 count) -
+1 strategy test: exit 4 → RolledBackwith operator-actionable Detail -
+1 constant pin: LinuxExitRolledBack == 4 -
Existing ScriptNonZeroExit_ReturnsFailedWithLastLogLinestill covers non-rollback failures (preserves regression coverage) -
5578/5578 unit tests green (+15 net new) -
E2E: deferred to H8 (E2E matrix includes rollback-triggered-by-failing-binary scenarios)
Operator UX after H6
FE can render three distinct upgrade-status badges based on MachineUpgradeStatus:
-
🟢 Upgraded/AlreadyUpToDate— green, no action -
🟡 RolledBack— yellow, "Safely restored to baseline. Review upgrade.log + last-upgrade.json before retry." -
🔴 Failed/NotSupported— red, "Investigation needed." -
🔵 Initiated— blue, "In progress, check back."
What's NOT in this PR
- H7: Role/feature capability slots (next — catches "IIS not installed" at plan-time)
- H8: Comprehensive E2E test matrix (validates the full upgrade chain end-to-end)