Skip to content

Pause and preserve checkpoint on deployment timeout, add manual resume

Placeholder ppxd requested to merge fix/checkpoint-preserving-timeout into main

Summary

  • On wall-clock timeout, a deployment now pauses (Paused) with its checkpoint preserved instead of transitioning to Failed and deleting the checkpoint. The already-completed batches are no longer thrown away.
  • Adds a manual resume trigger — POST /api/tasks/{id}/resume → ResumeServerTaskCommand → ServerTaskControlService.ResumeTaskAsync — that re-dispatches the paused deployment. It reuses the existing, proven resume machinery: StartExecutingAsync already transitions Paused → Executing (IsResumed), and the state-agnostic ResumeCheckpointPhase (Order 50) restores progress, so resume skips completed batches.
  • New default is the safe behaviour. The historical fail-fast path (Failed + checkpoint deleted) stays available behind SQUID_DEPLOYMENT_TIMEOUT_RESUMABLE=false for operators who alert on the terminal Failed state.

Design notes

  • Paused is the only resumable state (Paused → Executing is the lone valid resume transition; Failed/TimedOut are terminal). So a resumable timeout must land in Paused — hence OnTimedOutAsync transitions Executing → Paused, keeps the checkpoint (it is the resume point), and writes no DeploymentCompletion record (a paused deployment hasn't completed).
  • ResumeTaskAsync is the strict, operator-facing sibling of the opportunistic TryAutoResumeAsync (which fires after an interruption response). Both now share one EnqueueResumeAsync. Unlike the auto path, ResumeTaskAsync surfaces every precondition failure as a typed exception — ServerTaskNotFoundException, ServerTaskStateTransitionException (task not paused), ServerTaskAwaitingInterruptionException (resume a guided-failure pause by submitting the interruption, not here) — so the API returns an actionable error rather than a silent 200.
  • Non-breaking surface: the new completion-handler method has a single implementer (updated); the endpoint/command/exception are purely additive. The one intentional behaviour change (timeout terminal-state Failed → Paused) is the requested safe-default flip and is reversible via the env var.

Test plan

  • Unit — DeploymentCompletionHandlerTests: OnTimedOutAsync transitions Executing → Paused, does not delete the checkpoint, does not record a completion, does not trigger auto-deploys.
  • Unit — DeploymentPipelineRunnerTimeoutTests: timeout routes to OnTimedOutAsync by default and OnFailureAsync when fail-fast; DeploymentTimedOutEvent emitted in both modes; cancel still wins over timeout; SQUID_DEPLOYMENT_TIMEOUT_RESUMABLE const pinned + 15-case parse matrix + env→property wiring.
  • Unit — ServerTaskControlServiceResumeTests: paused+no-interruption enqueues + persists new job id; not-found / not-paused (7 states) / pending-interruption all throw and never enqueue; TryAutoResumeAsync regression after the shared-enqueue refactor.
  • Integration (real Postgres) — IntegrationTimeoutResume: OnTimedOutAsync → task Paused + checkpoint row preserved (with progress); fail-fast contrast → Failed + checkpoint deleted.
  • E2E — KubernetesResumeCheckpointE2ETests.ResumeAfterTimeoutPause_CompletesDeploymentFromCheckpoint: plant checkpoint at batch 0 + set Paused (post-timeout state) → ProcessAsync runs only batches 1+2, ends Success, checkpoint cleaned up. (Runs on the CI Kind-cluster E2E gate.)
  • Full DeploymentExecution + Deployments + Handlers unit sweep: 4355 passed, 0 failed.

Follow-up (out of scope)

  • A timeout currently records the DeploymentFailed audit-event category (pre-existing mapping; EventCategory has no TimedOut/Paused value). With the resumable default this is mildly imprecise — a dedicated category + the operator-facing "Resume" button belong with the frontend Events work.

Merge request reports

Loading