P0-5: Retry checkpoint persistence with exponential backoff before swallow
Summary
-
Real production hazard: pre-fix,
PersistCheckpointAsyncdid a single DB write attempt; on failure the catch logged atWarninglevel and continued. A transient blip silently lost the checkpoint. The Warning level meant operators running at Error threshold never saw the loss → server restart at the wrong moment would replay completed batches → double kubectl applies / helm upgrades / etc. -
Fix: 3 attempts × exponential backoff (200ms / 600ms / 1800ms; total ≤2.6s). On all-fail escalate to
Log.Errorwith explicit "resume hazard" message — visible to alerting at Error threshold. Deploy continues; only the resume story is degraded.
Test plan
-
Unit (CheckpointPersistRetryTests): 6 cases -
MaxAttempts_PinnedAt3+InitialDelay_IsAt200ms— Rule 8 drift detectors -
TransientFailureThenSuccess_RetriesAndSucceeds— pin: exactly one retry -
AllAttemptsFail_LogsErrorAndContinues— pin: full attempt count, no throw -
CancellationDuringRetry_StopsImmediately— pin: OCE propagates -
RetryDelays_AreBoundedAndExponential— stopwatch ≥700ms, < 5s
-
-
Full unit suite: 5114/5114
Rollback
Single-PR revert is clean. If reverted, the only behaviour delta is the silent loss reopens — no data corruption, no schema change.