Skip to content

P0-4: Bounded sync wait on log-stream task during ScriptPod teardown

Placeholder ppxd requested to merge fix/p0-4-task-run-exception-wrapper into main

Summary

  • Real production race: pre-fix CompleteScript / CancelScript cancelled the log-stream CTS and immediately deleted the pod + wiped the workspace; the background Task.Run(() => StreamLogsAsync(...)) was still running. Result: log lines lost AFTER the final drain, and read-after-delete exceptions landing on TaskScheduler.UnobservedTaskException (NOT Serilog — we don't subscribe).
  • Fix: new ScriptPodService.StopLogStream(ctx) cancels the CTS AND waits synchronously on the task with a 5 s bound. Sync-wait is intentional — IScriptService is a synchronous Halibut RPC contract; making the callers async would break deployed servers.

Test plan

  • Unit (StopLogStreamTests): 5 cases
    • StopLogStream_NullTask_Returns — graceful no-op when StartLogStream wasn't called
    • StopLogStream_NullCts_Returns — graceful no-op when CTS missing
    • StopLogStream_RunningTask_CancelsAndAwaits — pin: task is IsCompleted=true after StopLogStream returns
    • StopLogStream_TaskThatThrowsSync_ObservesException — pin: synchronous throw observed, not propagated
    • StopLogStream_TaskThatNeverFinishes_TimesOutAndContinues — pin: stopwatch < 8 s (5 s timeout + scheduler jitter)
  • Tentacle.Tests full suite: 1423/1423 (initial run had one unrelated flaky timing test that passed on rerun)

Rollback

Single-PR revert is clean. If reverted, the only behaviour delta is the orphan-task race re-opens — log lines may be lost on busy nodes or pods that produce output during teardown. No data loss.

🤖 Generated with Claude Code

Merge request reports

Loading