Skip to content

Add P0-#2: Server-restart polling-reconnect E2E + StubSquidServer.RestartHalibutAsync

Placeholder ppxd requested to merge feat/server-restart-polling-reconnect into main

Summary

  • Pins the production server-restart polling-reconnect contract that no E2E previously covered
  • Adds StubSquidServer.RestartHalibutAsync to Linux + Windows infrastructure (same-port + same-cert + replayed-trust restart simulation)
  • New Linux E2E R4h_RealBinary_ServerPodRestart_AgentReconnectsAndDispatchesResume exercises real binary's Halibut polling reconnect through a simulated server pod restart
  • Tier 🟢 H — only the server-side restart event is simulated; rest is production code

Production scenario this catches

Squid server pod restarts during routine operations (rolling deploy, OOM-kill, k8s reschedule). All polling-agent TCP connections receive RST. Halibut polling client should:

  1. Detect the connection drop
  2. Retry with backoff + jitter (production has SQUID_TENTACLE_POLLING_STARTUP_JITTER_MS)
  3. Reconnect to the same host:port once the new pod's listener is up
  4. Resume dispatches normally

Untested risks pre-PR:

  • Halibut polling client backoff regresses to 5min retry → ops sees 5min of "agent unreachable" alerts after every routine deploy
  • Trust list isn't preserved across simulated pod restart → reconnects fail TLS handshake
  • Agent process restarts (systemd Restart=on-failure) instead of just reconnecting → operator's in-flight scripts get killed

Why no E2E covered this before

StubSquidServer's pre-PR API only supported StartAsync + DisposeAsync. A new stub got:

  • New ports (would break agent's persisted ServerCommsUrl config)
  • New self-signed cert (would break agent's pinned thumbprint)

This PR adds RestartHalibutAsync to both Linux and Windows StubSquidServer copies:

  • Disposes current Halibut runtime → closes listener → all polling-agent TCP connections receive RST
  • Rebuilds new runtime on the same port + same cert
  • Replays the trust list onto the new runtime
  • Faithfully simulates production pod-restart from the agent's perspective

Test mechanism (10 steps, ~60-90s)

  1. R1h-style setup: register, service install, RUNNING, polling up
  2. Verify polling channel via capabilities probe (initial)
  3. Dispatch BEFORE restart (sanity baseline — proves baseline works)
  4. SIMULATE POD RESTART: stub.RestartHalibutAsync()
  5. Brief simulated downtime (1s; production is 5-30s)
  6. Wait for capabilities probe to reconnect (60s timeout — Halibut backoff)
  7. Dispatch AFTER restart (proves dispatches resume)
  8. Re-probe to confirm channel stability
  9. Dispatch STABLE post-reconnect (proves it's not transient)
  10. Cleanup: service uninstall --purge

Key assertion: capability version stable across restart

postRestartCapabilities.AgentVersion.ShouldBe(initialAgentVersion,
    customMessage: "If different: the agent process restarted instead of just reconnecting...");

This catches the difference between "agent crashed + systemd restarted" vs "polling reconnected" — operators care about the latter.

Test plan

  • dotnet build both projects — 0 errors
  • CI on ubuntu-latest runs Category=LinuxTentacleBinaryE2E:
    • R4h_RealBinary_ServerPodRestart_AgentReconnectsAndDispatchesResume passes
    • All other R1h/R2h tests still green
  • Re-run 2-3 times to confirm reconnect timing is deterministic across runs

Phase status — P0/P1 hardening sweep after 1.6.4 release

PR Priority Title Status
#275 P0-#3 LocalScriptService async-flush race fix 🟡 CI in flight
#276 P0-#1 SCM-launched real-binary E2E (verifies #274) 🟡 CI in flight
this P0-#2 Server-restart polling-reconnect E2E 🟡 CI pending
(next) P1-#4 Tentacle upgrade + polling composite pending
(next) P1-#5 Long-soak (5min+ probe + dispatch loop) pending
(next) P1-#6 Capabilities cache TTL invalidation pending

🤖 Generated with Claude Code

Merge request reports

Loading