Skip to content

Add Phase 13 PR-4: Real-binary polling agent soak stability (Linux + Windows)

Summary

  • Adds R2h soak tests on both Linux + Windows that mirror R1h's setup but exercise the polling channel over a 30s window across 5 dispatches + 5 capability probes interleaved
  • Closes a Phase 13 gap: R1h proves the polling channel comes up + accepts a single dispatch, but doesn't prove it STAYS up across multiple dispatches or that the agent process survives the soak
  • Tier 🟢 H per Rule 12.4 — same fidelity as R1h, just over a longer window

Stacking

PR-4 is stacked on top of #271 (PR-3) and #269 (PR-1), both of which add the per-OS R1h infrastructure that R2h extends. Once those merge, GitHub will retarget this PR's base branch to main. R2h-Linux extends the test class added in PR-1; R2h-Windows extends the test class added in PR-3.

Why R2h matters (the gap R1h leaves open)

Operator scenario R1h covers R2h adds
First deployment dispatch after agent registers ✅ ✅
Multi-step deployment pipeline against the same agent (5 steps in a row) ❌ ✅
Long-running soak / keep-alive (agent stays connected for minutes) ❌ ✅
Capabilities probe consistency across the agent's lifetime ❌ ✅

A regression where the agent crashes after a single dispatch OR Halibut's polling client fails to retry on transient drops would ship silently — ALL of R1h's assertions still pass.

What R2h catches that R1h doesn't

  • Agent process crash after first dispatch (Halibut error handler regression, scriptbackend leaks, FD leaks)
  • Polling loop's Task.Delay backoff drifts to a value beyond the test's deadline
  • LocalScriptService stale state between dispatches (workdir not cleaned, output capture confused with prior iteration's data)
  • Capabilities reporting flap (different version reported across probes — would indicate version-string mutation regression)

Mechanism (mirrored Linux + Windows)

  1. Setup identical to R1h.
  2. Loop 5 iterations spanning ~25s wall-clock:
    • Probe capabilities → assert version stays same as initial
    • Dispatch a per-iteration script with unique marker
    • Assert exit 0 + marker round-trips
    • Sleep 4-5s before next iteration
  3. Final pin: agent process still alive (Linux: systemctl is-active; Windows: Process.HasExited == false).
  4. Cleanup: same as R1h.

Why 5 iterations / ~25s soak

Balances coverage (multi-dispatch + agent staying alive) against CI cost. Each iter ~3s + 4s spacing = ~7s × 5 = ~35s ceiling. Real wall-clock typically 25-30s.

Test plan

  • CI on ubuntu-latest runs Category=LinuxTentacleBinaryE2E:
    • R1h_RealBinary_PollingAgent_ScriptDispatchRoundTripsThroughHalibut (PR-1) still green
    • R2h_RealBinary_PollingAgent_SoakStability_5DispatchesOver25Seconds passes
  • CI on windows-latest runs Category=WindowsTentacleBinaryE2E:
    • R1h_RealBinary_PollingAgent_ScriptDispatchRoundTripsThroughHalibut (PR-3) still green
    • R2h_RealBinary_PollingAgent_SoakStability_5DispatchesOver25Seconds passes
  • dotnet build both test projects — 0 errors

Phase 13 status after this lands

PR Status Coverage
#269 (PR-1) 🟢 GREEN Linux real-binary as polling agent (single dispatch)
#270 (PR-2) 🟢 GREEN WindowsTentacleBinaryFixture + smoke
#271 (PR-3) 🟢 GREEN Windows real-binary as polling agent (single dispatch)
PR-4 (this) 🟡 awaiting CI Soak stability for both OSes (multi-dispatch + agent survives)

🤖 Generated with Claude Code

Merge request reports

Loading