P0-4: Bounded sync wait on log-stream task during ScriptPod teardown
Summary
-
Real production race: pre-fix
CompleteScript/CancelScriptcancelled the log-stream CTS and immediately deleted the pod + wiped the workspace; the backgroundTask.Run(() => StreamLogsAsync(...))was still running. Result: log lines lost AFTER the final drain, and read-after-delete exceptions landing onTaskScheduler.UnobservedTaskException(NOT Serilog — we don't subscribe). -
Fix: new
ScriptPodService.StopLogStream(ctx)cancels the CTS AND waits synchronously on the task with a 5 s bound. Sync-wait is intentional —IScriptServiceis a synchronous Halibut RPC contract; making the callers async would break deployed servers.
Test plan
-
Unit (StopLogStreamTests): 5 cases -
StopLogStream_NullTask_Returns— graceful no-op when StartLogStream wasn't called -
StopLogStream_NullCts_Returns— graceful no-op when CTS missing -
StopLogStream_RunningTask_CancelsAndAwaits— pin: task isIsCompleted=trueafter StopLogStream returns -
StopLogStream_TaskThatThrowsSync_ObservesException— pin: synchronous throw observed, not propagated -
StopLogStream_TaskThatNeverFinishes_TimesOutAndContinues— pin: stopwatch < 8 s (5 s timeout + scheduler jitter)
-
-
Tentacle.Tests full suite: 1423/1423 (initial run had one unrelated flaky timing test that passed on rerun)
Rollback
Single-PR revert is clean. If reverted, the only behaviour delta is the orphan-task race re-opens — log lines may be lost on busy nodes or pods that produce output during teardown. No data loss.