Skip to content

Harden in-flight ledger: bounded lock stripes + retried ensure-row

Placeholder ppxd requested to merge chore/inflight-store-hardening into main

Summary

Two non-blocking hardening follow-ups from the milestone 1.8.4 deep review of resume-by-ticket (#384):

  • M1 — bounded locking: InFlightScriptStore used a static ConcurrentDictionary<int, SemaphoreSlim> keyed by serverTaskId that never evicted → one semaphore leaked per deployment task for the server's lifetime. Replaced with a fixed set of lock stripes (serverTaskId % 64): a task always maps to the same stripe (so its read-modify-write still serialises), distinct tasks colliding on a stripe serialise harmlessly, and the set is bounded.
  • M2 — retried ensure-row: EnsureCheckpointRowAsync was a single un-retried write while the sibling PersistCheckpointAsync had a 3-attempt backoff — so one transient DB blip silently disabled resume-by-ticket for the whole run. Extracted the retry+backoff into a shared helper used by both.

Behaviour-preserving and non-breaking: per-task serialisation + persist-retry semantics are unchanged.

Test plan

  • CheckpointPersistRetry suite green incl. new EnsureCheckpointRow_TransientFailureThenSuccess_Retries (M2)
  • InFlightScriptStore integration green incl. the 25-machine concurrent-record test (M1 striping preserves per-task serialisation)
  • Full unit suite green (5682); Squid.Core builds clean

Merge request reports

Loading