Skip to content

feat(LinuxTentacleE2E): Phase 12.L.E.13 — ROLLBACK_CRITICAL_FAILED worst-case path

Summary

Pins the .sh's exit-9 + ROLLBACK_CRITICAL_FAILED contract: when Phase B fails AND the .bak rollback ALSO fails, the agent goes BINARYLESS and the operator UI MUST surface this as a manual-intervention-required state.

Coverage layering completed

Phase Scenario Exit Status
J.L.E.7 Happy path 0 SUCCESS
J.L.E.9 Phase B fail → rollback succeeds 4 ROLLED_BACK
This PR Phase B fail → rollback ALSO fails 9 ROLLBACK_CRITICAL_FAILED

Without this pin, a regression that breaks the exit-9 path (e.g. someone removes the write_status call so operators see ROLLING_BACK forever instead of ROLLBACK_CRITICAL_FAILED) ships silently. Operators would get no clear "your agent is binaryless, run install-tentacle.sh to recover" signal.

Test mechanism (composed)

  1. failHealthz=true v2 (J.L.E.9 mechanism) triggers Phase B's healthz fail → rollback path.
  2. Watchdog Task polls for INSTALL_DIR.bak (= Phase B's sudo mv $INSTALL_DIR $BAK_DIR just completed) and overwrites .bak's service script with #!/bin/bash\nexit 1. Mode 0755 keeps it executable so systemd CAN run it and observe the immediate failure.
  3. Rollback's sudo mv $BAK_DIR $INSTALL_DIR succeeds (rename), but contents are broken.
  4. Rollback's sudo systemctl start fires the broken script → bash exits 1 → unit failed → is-active never returns active.
  5. .sh's 30-iteration is-active wait times out → ROLLBACK_OK=0 → write_status ROLLBACK_CRITICAL_FAILED → exit 9.

Why ownership works

sudo mv preserves file ownership. Test process owns INSTALL_DIR (under /tmp). After Phase B's mv, .bak's contents are still test-process-owned → we write to them without sudo.

Why a watchdog (vs pre-corruption)

Pre-corrupting INSTALL_DIR's script would crash the running v1 service before the test reaches the v1-marker pre-condition assertion. Watchdog waits for .bak to exist (= Phase B's mv just completed) THEN corrupts.

Assertions

  • exitCode == 9 (per .sh line 719)
  • Watchdog actually fired within 60s deadline (sanity: without this, test would silently fall back to J.L.E.9's path)
  • last-upgrade.json.Status == "ROLLBACK_CRITICAL_FAILED"
  • last-upgrade.json.Detail contains "manual intervention"
  • stdout contains "CRITICAL: rollback also failed" (journalctl signal)
  • .bak directory consumed by rollback's mv (failure is post-mv)

Infrastructure additions

LinuxLifecycleContext:

  • StartBakCorruptionWatchdog(TimeSpan deadline) → spawns Task that polls for .bak then corrupts its embedded service script. Returns Task<bool> for sanity check after .sh exits.

Fidelity tier

🟢 High (Rule 12.4): real prod .sh + real systemd + real sudo + real bash + real watchdog corruption + real systemctl-start failure on a deliberately-broken v1.

The heaviest test in the suite (~50s expected runtime) but worst-case path coverage value justifies it.

Test plan

  • Linux E2E workflow runs Squid.LinuxTentacleE2ETests (manual workflow_dispatch after merge)
  • E1uRollbackCritical_PhaseBHealthzFailAndRollbackV1Broken_ExitsNineWithCriticalFailedStatus passes within ~60s
  • No regression on existing 13 Linux E2E tests

🤖 Generated with Claude Code

Merge request reports

Loading