Add Phase 13 PR-4: Real-binary polling agent soak stability (Linux + Windows)
Summary
- Adds R2h soak tests on both Linux + Windows that mirror R1h's setup but exercise the polling channel over a 30s window across 5 dispatches + 5 capability probes interleaved
- Closes a Phase 13 gap: R1h proves the polling channel comes up + accepts a single dispatch, but doesn't prove it STAYS up across multiple dispatches or that the agent process survives the soak
- Tier
🟢 H per Rule 12.4 — same fidelity as R1h, just over a longer window
Stacking
PR-4 is stacked on top of #271 (PR-3) and #269 (PR-1), both of which add the per-OS R1h infrastructure that R2h extends. Once those merge, GitHub will retarget this PR's base branch to main. R2h-Linux extends the test class added in PR-1; R2h-Windows extends the test class added in PR-3.
Why R2h matters (the gap R1h leaves open)
| Operator scenario | R1h covers | R2h adds |
|---|---|---|
| First deployment dispatch after agent registers | ||
| Multi-step deployment pipeline against the same agent (5 steps in a row) | ||
| Long-running soak / keep-alive (agent stays connected for minutes) | ||
| Capabilities probe consistency across the agent's lifetime |
A regression where the agent crashes after a single dispatch OR Halibut's polling client fails to retry on transient drops would ship silently — ALL of R1h's assertions still pass.
What R2h catches that R1h doesn't
- Agent process crash after first dispatch (Halibut error handler regression, scriptbackend leaks, FD leaks)
- Polling loop's
Task.Delaybackoff drifts to a value beyond the test's deadline -
LocalScriptServicestale state between dispatches (workdir not cleaned, output capture confused with prior iteration's data) - Capabilities reporting flap (different version reported across probes — would indicate version-string mutation regression)
Mechanism (mirrored Linux + Windows)
- Setup identical to R1h.
- Loop 5 iterations spanning ~25s wall-clock:
- Probe capabilities → assert version stays same as initial
- Dispatch a per-iteration script with unique marker
- Assert exit 0 + marker round-trips
- Sleep 4-5s before next iteration
- Final pin: agent process still alive (Linux:
systemctl is-active; Windows:Process.HasExited == false). - Cleanup: same as R1h.
Why 5 iterations / ~25s soak
Balances coverage (multi-dispatch + agent staying alive) against CI cost. Each iter ~3s + 4s spacing = ~7s × 5 = ~35s ceiling. Real wall-clock typically 25-30s.
Test plan
-
CI on ubuntu-latestrunsCategory=LinuxTentacleBinaryE2E:-
R1h_RealBinary_PollingAgent_ScriptDispatchRoundTripsThroughHalibut(PR-1) still green -
R2h_RealBinary_PollingAgent_SoakStability_5DispatchesOver25Secondspasses
-
-
CI on windows-latestrunsCategory=WindowsTentacleBinaryE2E:-
R1h_RealBinary_PollingAgent_ScriptDispatchRoundTripsThroughHalibut(PR-3) still green -
R2h_RealBinary_PollingAgent_SoakStability_5DispatchesOver25Secondspasses
-
-
dotnet buildboth test projects — 0 errors
Phase 13 status after this lands
| PR | Status | Coverage |
|---|---|---|
| #269 (PR-1) |
|
Linux real-binary as polling agent (single dispatch) |
| #270 (PR-2) |
|
WindowsTentacleBinaryFixture + smoke |
| #271 (PR-3) |
|
Windows real-binary as polling agent (single dispatch) |
| PR-4 (this) |
|
Soak stability for both OSes (multi-dispatch + agent survives) |