feat(LinuxTentacleE2E): Phase 12.L.E.13 — ROLLBACK_CRITICAL_FAILED worst-case path
Summary
Pins the .sh's exit-9 + ROLLBACK_CRITICAL_FAILED contract: when Phase B fails AND the .bak rollback ALSO fails, the agent goes BINARYLESS and the operator UI MUST surface this as a manual-intervention-required state.
Coverage layering completed
| Phase | Scenario | Exit | Status |
|---|---|---|---|
| J.L.E.7 | Happy path | 0 | SUCCESS |
| J.L.E.9 | Phase B fail → rollback succeeds | 4 | ROLLED_BACK |
| This PR | Phase B fail → rollback ALSO fails | 9 | ROLLBACK_CRITICAL_FAILED |
Without this pin, a regression that breaks the exit-9 path (e.g. someone removes the write_status call so operators see ROLLING_BACK forever instead of ROLLBACK_CRITICAL_FAILED) ships silently. Operators would get no clear "your agent is binaryless, run install-tentacle.sh to recover" signal.
Test mechanism (composed)
-
failHealthz=truev2 (J.L.E.9 mechanism) triggers Phase B's healthz fail → rollback path. -
Watchdog Task polls for
INSTALL_DIR.bak(= Phase B'ssudo mv $INSTALL_DIR $BAK_DIRjust completed) and overwrites.bak's service script with#!/bin/bash\nexit 1. Mode0755keeps it executable so systemd CAN run it and observe the immediate failure. - Rollback's
sudo mv $BAK_DIR $INSTALL_DIRsucceeds (rename), but contents are broken. - Rollback's
sudo systemctl startfires the broken script → bash exits 1 → unit failed →is-activenever returns active. -
.sh's 30-iterationis-activewait times out →ROLLBACK_OK=0→write_status ROLLBACK_CRITICAL_FAILED→ exit 9.
Why ownership works
sudo mv preserves file ownership. Test process owns INSTALL_DIR (under /tmp). After Phase B's mv, .bak's contents are still test-process-owned → we write to them without sudo.
Why a watchdog (vs pre-corruption)
Pre-corrupting INSTALL_DIR's script would crash the running v1 service before the test reaches the v1-marker pre-condition assertion. Watchdog waits for .bak to exist (= Phase B's mv just completed) THEN corrupts.
Assertions
-
exitCode == 9(per .sh line 719) - Watchdog actually fired within 60s deadline (sanity: without this, test would silently fall back to J.L.E.9's path)
last-upgrade.json.Status == "ROLLBACK_CRITICAL_FAILED"-
last-upgrade.json.Detailcontains"manual intervention" - stdout contains
"CRITICAL: rollback also failed"(journalctl signal) -
.bakdirectory consumed by rollback's mv (failure is post-mv)
Infrastructure additions
LinuxLifecycleContext:
-
StartBakCorruptionWatchdog(TimeSpan deadline)→ spawns Task that polls for.bakthen corrupts its embedded service script. ReturnsTask<bool>for sanity check after.shexits.
Fidelity tier
.sh + real systemd + real sudo + real bash + real watchdog corruption + real systemctl-start failure on a deliberately-broken v1.
The heaviest test in the suite (~50s expected runtime) but worst-case path coverage value justifies it.
Test plan
-
Linux E2E workflow runs Squid.LinuxTentacleE2ETests(manualworkflow_dispatchafter merge) -
E1uRollbackCritical_PhaseBHealthzFailAndRollbackV1Broken_ExitsNineWithCriticalFailedStatuspasses within ~60s -
No regression on existing 13 Linux E2E tests