Harden in-flight ledger: bounded lock stripes + retried ensure-row
Summary
Two non-blocking hardening follow-ups from the milestone 1.8.4 deep review of resume-by-ticket (#384):
-
M1 — bounded locking:
InFlightScriptStoreused astatic ConcurrentDictionary<int, SemaphoreSlim>keyed byserverTaskIdthat never evicted → one semaphore leaked per deployment task for the server's lifetime. Replaced with a fixed set of lock stripes (serverTaskId % 64): a task always maps to the same stripe (so its read-modify-write still serialises), distinct tasks colliding on a stripe serialise harmlessly, and the set is bounded. -
M2 — retried ensure-row:
EnsureCheckpointRowAsyncwas a single un-retried write while the siblingPersistCheckpointAsynchad a 3-attempt backoff — so one transient DB blip silently disabled resume-by-ticket for the whole run. Extracted the retry+backoff into a shared helper used by both.
Behaviour-preserving and non-breaking: per-task serialisation + persist-retry semantics are unchanged.
Test plan
-
CheckpointPersistRetrysuite green incl. newEnsureCheckpointRow_TransientFailureThenSuccess_Retries(M2) -
InFlightScriptStoreintegration green incl. the 25-machine concurrent-record test (M1 striping preserves per-task serialisation) -
Full unit suite green (5682); Squid.Core builds clean