Skip to content

feat(LinuxTentacleE2E): Phase 12.L.E.9 — Phase B healthz fail → .bak rollback (Linux)

Summary

Linux mirror of Windows E7.u1 (E7u1_NewBinaryOnStartCrashes_TriggersAutoRollbackToV1) but at the healthz-failure axis — Linux's systemctl-start succeeds even when the new binary's healthz responder is broken, so the failure surfaces in the .sh's curl -fsS retry loop, not at service-start.

This is the highest-severity unguarded path on Linux until this test landed. Without this pin, a future polish that breaks rollback ships silently → next bad upgrade leaves the agent BINARYLESS → operator intervention required.

What the test pins

The .sh's rollback contract (lines ~620–720 of upgrade-linux-tentacle.sh):

1. systemctl restart $SERVICE         ← v2 active
2. for i in 1..HEALTHCHECK_RETRIES:
     curl -fsS $HEALTHCHECK_URL       ← v2 returns 503 (this test)
3. retries exhausted → HEALTH_OK=0
4. Rollback path:
     - systemctl stop $SERVICE
     - rm -rf $INSTALL_DIR
     - mv $BAK_DIR $INSTALL_DIR       ← v1 binary back
     - systemctl start $SERVICE        ← v1 active
     - wait is-active up to 30s
     - exit 4 + write_status ROLLED_BACK

Test mechanism — surgical 1-line mutation

ctx.BuildV2BundleTarGz(targetVersion: "2.0.0-test", failHealthz: true);

Injects a 200→503 swap in the embedded python3 healthz responder. Same script binds the same port, returns a real HTTP response, just an unhealthy one. .sh's curl -fsS rejects 5xx as exit 22 → HEALTH_OK stays 0 → retry loop exhausts → rollback fires.

The mutation guards against silent script drift: throws InvalidOperationException if the sentinel self.send_response(200) is no longer present, forcing the test author to update both sides in lockstep.

Assertions

  • exitCode == 4 (rollback succeeded)
  • last-upgrade.json.Status == "ROLLED_BACK"
  • last-upgrade.json.Detail contains "New binary failed health check" (operator-actionable)
  • Marker eventually = "1.0.0" (proves v1 systemctl-started AND wrote its marker post-rollback)
  • .bak directory consumed by rollback's mv (reverse-assert; if it still exists, mv silently failed)

Why this isn't already covered

E1.h-Linux (J.L.E.7): healthz passes → SUCCESS path E1.u1-Linux: download 404 → Phase A bails before any swap E12.u1-Linux: SHA mismatch → Phase A bails before any swap E15.h-Linux (J.L.E.8): preservation across SUCCESS upgrade

None of these exercise Phase B's failure-then-restore path. This PR is the first Linux test to actually run rollback end-to-end.

Fidelity tier

🟢 High (Rule 12.4): real production .sh + real systemd-run + real sudo + real bash + real LocalReleaseMirror + real python3 returning 503 + real mv rollback + real v1 systemctl-start + real marker rewrite. No mocks at any layer.

Test plan

  • Linux E2E workflow runs Squid.LinuxTentacleE2ETests to completion (manual workflow_dispatch after merge)
  • E1uRollback_PhaseBHealthcheckFails_RestoresV1FromBak passes within ~30s
  • Pre-existing 9 Linux E2E tests stay green (no regression on shared infra)

🤖 Generated with Claude Code

Merge request reports

Loading