Skip to content

Add P1-#5: Long-soak E2E with VmRSS + FD leak detection (R6h-Linux)

Placeholder ppxd requested to merge feat/long-soak-e2e into main

Summary

  • New E2E R6h_PollingAgent_LongSoak_NoMemoryOrFdLeaks runs the polling agent for a configurable window (default 120s, up to 30 min via env var) with periodic dispatches + probes interleaved
  • Captures /proc/<pid>/status (VmRSS) + /proc/<pid>/fd (FD count) at baseline and final, asserts growth stays bounded
  • Catches leak-class regressions invisible to short tests

Coverage delta

Test Soak duration Catches
R2h (PR #272) 25s, 5 dispatches Multi-dispatch + agent-still-alive
R6h (this PR) 120s default (configurable up to 30 min), ~12 dispatches + ~8 probes Memory leaks, FD leaks, polling-loop drift, capabilities-reporting flap

What this catches that prior tests don't

  • Memory leak class: each dispatch allocates LogWriter / workdir / Process. R2h's 5 dispatches over 25s isn't enough cycles to surface. R6h's 12+ dispatches over 2+ minutes amplify the signal.
  • File descriptor leak class: each dispatch opens stdin/stdout/stderr pipes that should close on completion. ulimit (default 1024) caps the agent at ~341 leaked dispatches before connections fail. R6h pins FD count growth ≤ 20 across the soak.
  • Polling-loop drift: a regression that stretches Task.Delay backoff beyond bounds wouldn't surface in 25s but would in 2-min.
  • Capabilities reporting flap: probe values varying mid-soak indicates non-determinism in agent's CapabilitiesService.

Configurable via env var (Rule 8)

SQUID_TENTACLE_E2E_SOAK_SECONDS=300   # operator investigating a leak
  • Default 120s — short enough to keep CI cycle reasonable, long enough to catch most leak-class regressions
  • Min 30 / max 1800 (30 min) — clamped via Math.Clamp
  • Workflow_dispatch can override via input for diagnostic runs

Assertions

  1. All dispatches succeeded — tracked as failedDispatches list with timestamps
  2. All probes consistent — probeVersions.Distinct().Count() == 1
  3. Process still alive — systemctl is-active exits 0
  4. VmRSS growth ≤ 50MB — production startup ~150MB; 50MB growth in 2min indicates runaway leak
  5. FD count growth ≤ 20 — each dispatch opens ~3 pipes that close on completion; net should be 0 + small tolerance for transient runtime

Test plan

  • dotnet build — 0 errors
  • CI on ubuntu-latest runs Category=LinuxTentacleBinaryE2E:
    • R6h_PollingAgent_LongSoak_NoMemoryOrFdLeaks passes with default 120s soak (~145s total runtime)
    • All other R1h/R2h/R4h/R5h tests still green
  • (manual) operator can run with SQUID_TENTACLE_E2E_SOAK_SECONDS=600 for diagnostic 10-min soak

Phase status — P0/P1 sweep after 1.6.4 release

PR Priority Title Status
#275 P0-#3 LocalScriptService async-flush 🟡 CI
#276 P0-#1 SCM-launched real-binary E2E 🟡 CI
#277 P0-#2 Server-restart polling-reconnect E2E 🟡 CI
#278 P1-#4 Upgrade + polling composite E2E 🟡 CI
this P1-#5 Long-soak E2E with leak detection 🟡 CI pending
(next) P1-#6 Capabilities cache TTL invalidation pending

🤖 Generated with Claude Code

Merge request reports

Loading