Skip to content

P0-5: Retry checkpoint persistence with exponential backoff before swallow

Placeholder ppxd requested to merge fix/p0-5-checkpoint-persist-retry into main

Summary

  • Real production hazard: pre-fix, PersistCheckpointAsync did a single DB write attempt; on failure the catch logged at Warning level and continued. A transient blip silently lost the checkpoint. The Warning level meant operators running at Error threshold never saw the loss → server restart at the wrong moment would replay completed batches → double kubectl applies / helm upgrades / etc.
  • Fix: 3 attempts × exponential backoff (200ms / 600ms / 1800ms; total ≤2.6s). On all-fail escalate to Log.Error with explicit "resume hazard" message — visible to alerting at Error threshold. Deploy continues; only the resume story is degraded.

Test plan

  • Unit (CheckpointPersistRetryTests): 6 cases
    • MaxAttempts_PinnedAt3 + InitialDelay_IsAt200ms — Rule 8 drift detectors
    • TransientFailureThenSuccess_RetriesAndSucceeds — pin: exactly one retry
    • AllAttemptsFail_LogsErrorAndContinues — pin: full attempt count, no throw
    • CancellationDuringRetry_StopsImmediately — pin: OCE propagates
    • RetryDelays_AreBoundedAndExponential — stopwatch ≥700ms, < 5s
  • Full unit suite: 5114/5114

Rollback

Single-PR revert is clean. If reverted, the only behaviour delta is the silent loss reopens — no data corruption, no schema change.

🤖 Generated with Claude Code

Merge request reports

Loading