Pause and preserve checkpoint on deployment timeout, add manual resume
Summary
- On wall-clock timeout, a deployment now pauses (
Paused) with its checkpoint preserved instead of transitioning toFailedand deleting the checkpoint. The already-completed batches are no longer thrown away. - Adds a manual resume trigger —
POST /api/tasks/{id}/resume→ResumeServerTaskCommand→ServerTaskControlService.ResumeTaskAsync— that re-dispatches the paused deployment. It reuses the existing, proven resume machinery:StartExecutingAsyncalready transitionsPaused → Executing(IsResumed), and the state-agnosticResumeCheckpointPhase(Order 50) restores progress, so resume skips completed batches. - New default is the safe behaviour. The historical fail-fast path (
Failed+ checkpoint deleted) stays available behindSQUID_DEPLOYMENT_TIMEOUT_RESUMABLE=falsefor operators who alert on the terminalFailedstate.
Design notes
-
Pausedis the only resumable state (Paused → Executingis the lone valid resume transition;Failed/TimedOutare terminal). So a resumable timeout must land inPaused— henceOnTimedOutAsynctransitionsExecuting → Paused, keeps the checkpoint (it is the resume point), and writes noDeploymentCompletionrecord (a paused deployment hasn't completed). -
ResumeTaskAsyncis the strict, operator-facing sibling of the opportunisticTryAutoResumeAsync(which fires after an interruption response). Both now share oneEnqueueResumeAsync. Unlike the auto path,ResumeTaskAsyncsurfaces every precondition failure as a typed exception —ServerTaskNotFoundException,ServerTaskStateTransitionException(task not paused),ServerTaskAwaitingInterruptionException(resume a guided-failure pause by submitting the interruption, not here) — so the API returns an actionable error rather than a silent 200. - Non-breaking surface: the new completion-handler method has a single implementer (updated); the endpoint/command/exception are purely additive. The one intentional behaviour change (timeout terminal-state
Failed → Paused) is the requested safe-default flip and is reversible via the env var.
Test plan
-
Unit — DeploymentCompletionHandlerTests:OnTimedOutAsynctransitionsExecuting → Paused, does not delete the checkpoint, does not record a completion, does not trigger auto-deploys. -
Unit — DeploymentPipelineRunnerTimeoutTests: timeout routes toOnTimedOutAsyncby default andOnFailureAsyncwhen fail-fast;DeploymentTimedOutEventemitted in both modes; cancel still wins over timeout;SQUID_DEPLOYMENT_TIMEOUT_RESUMABLEconst pinned + 15-case parse matrix + env→property wiring. -
Unit — ServerTaskControlServiceResumeTests: paused+no-interruption enqueues + persists new job id; not-found / not-paused (7 states) / pending-interruption all throw and never enqueue;TryAutoResumeAsyncregression after the shared-enqueue refactor. -
Integration (real Postgres) — IntegrationTimeoutResume:OnTimedOutAsync→ taskPaused+ checkpoint row preserved (with progress); fail-fast contrast →Failed+ checkpoint deleted. -
E2E — KubernetesResumeCheckpointE2ETests.ResumeAfterTimeoutPause_CompletesDeploymentFromCheckpoint: plant checkpoint at batch 0 + setPaused(post-timeout state) →ProcessAsyncruns only batches 1+2, endsSuccess, checkpoint cleaned up. (Runs on the CI Kind-cluster E2E gate.) -
Full DeploymentExecution + Deployments + Handlers unit sweep: 4355 passed, 0 failed.
Follow-up (out of scope)
- A timeout currently records the
DeploymentFailedaudit-event category (pre-existing mapping;EventCategoryhas noTimedOut/Pausedvalue). With the resumable default this is mildly imprecise — a dedicated category + the operator-facing "Resume" button belong with the frontend Events work.