Add P0-#2: Server-restart polling-reconnect E2E + StubSquidServer.RestartHalibutAsync
Summary
- Pins the production server-restart polling-reconnect contract that no E2E previously covered
- Adds
StubSquidServer.RestartHalibutAsyncto Linux + Windows infrastructure (same-port + same-cert + replayed-trust restart simulation) - New Linux E2E
R4h_RealBinary_ServerPodRestart_AgentReconnectsAndDispatchesResumeexercises real binary's Halibut polling reconnect through a simulated server pod restart - Tier
🟢 H — only the server-side restart event is simulated; rest is production code
Production scenario this catches
Squid server pod restarts during routine operations (rolling deploy, OOM-kill, k8s reschedule). All polling-agent TCP connections receive RST. Halibut polling client should:
- Detect the connection drop
- Retry with backoff + jitter (production has
SQUID_TENTACLE_POLLING_STARTUP_JITTER_MS) - Reconnect to the same host:port once the new pod's listener is up
- Resume dispatches normally
Untested risks pre-PR:
- Halibut polling client backoff regresses to 5min retry → ops sees 5min of "agent unreachable" alerts after every routine deploy
- Trust list isn't preserved across simulated pod restart → reconnects fail TLS handshake
- Agent process restarts (systemd Restart=on-failure) instead of just reconnecting → operator's in-flight scripts get killed
Why no E2E covered this before
StubSquidServer's pre-PR API only supported StartAsync + DisposeAsync. A new stub got:
- New ports (would break agent's persisted
ServerCommsUrlconfig) - New self-signed cert (would break agent's pinned thumbprint)
This PR adds RestartHalibutAsync to both Linux and Windows StubSquidServer copies:
- Disposes current Halibut runtime → closes listener → all polling-agent TCP connections receive RST
- Rebuilds new runtime on the same port + same cert
- Replays the trust list onto the new runtime
- Faithfully simulates production pod-restart from the agent's perspective
Test mechanism (10 steps, ~60-90s)
- R1h-style setup: register, service install, RUNNING, polling up
- Verify polling channel via capabilities probe (initial)
- Dispatch BEFORE restart (sanity baseline — proves baseline works)
-
SIMULATE POD RESTART:
stub.RestartHalibutAsync() - Brief simulated downtime (1s; production is 5-30s)
- Wait for capabilities probe to reconnect (60s timeout — Halibut backoff)
- Dispatch AFTER restart (proves dispatches resume)
- Re-probe to confirm channel stability
- Dispatch STABLE post-reconnect (proves it's not transient)
- Cleanup:
service uninstall --purge
Key assertion: capability version stable across restart
postRestartCapabilities.AgentVersion.ShouldBe(initialAgentVersion,
customMessage: "If different: the agent process restarted instead of just reconnecting...");
This catches the difference between "agent crashed + systemd restarted" vs "polling reconnected" — operators care about the latter.
Test plan
-
dotnet buildboth projects — 0 errors -
CI on ubuntu-latestrunsCategory=LinuxTentacleBinaryE2E:-
R4h_RealBinary_ServerPodRestart_AgentReconnectsAndDispatchesResumepasses -
All other R1h/R2h tests still green
-
-
Re-run 2-3 times to confirm reconnect timing is deterministic across runs
Phase status — P0/P1 hardening sweep after 1.6.4 release
| PR | Priority | Title | Status |
|---|---|---|---|
| #275 | P0-#3 | LocalScriptService async-flush race fix |
|
| #276 | P0-#1 | SCM-launched real-binary E2E (verifies #274) |
|
| this | P0-#2 | Server-restart polling-reconnect E2E |
|
| (next) | P1-#4 | Tentacle upgrade + polling composite | pending |
| (next) | P1-#5 | Long-soak (5min+ probe + dispatch loop) | pending |
| (next) | P1-#6 | Capabilities cache TTL invalidation | pending |