Serialize deploy and upgrade per machine across pods
Summary
Second of two PRs hardening multi-pod concurrency correctness (companion to the atomic task-slot in #441).
Problem. A tentacle upgrade restarts a machine's agent, so it must never run while a deployment's script is executing on that machine. Upgrades already take a per-machine Redis lock before dispatch (MachineUpgradeService.DispatchUnderLockAsync); deployments took none — so on a horizontally-scaled multi-pod server, a deploy and an upgrade for the same machine could run concurrently and the upgrade would restart the agent mid-script. (The #427 agent-side script-isolation mutex only serializes once both scripts reach the agent — it can't stop the server-side dispatch decision.)
Fix — the deployment dispatch acquires the SAME per-machine lock the upgrade uses:
- New
IMachineDispatchLockwraps the existingIRedisSafeRunner, running an action under the per-machine lock keyed exactly asUpgradeDispatchLockReconciler.BuildLockKey(machineId)— pinned by a unit test so deploy and upgrade can never drift onto different keys. It maps the runner's two non-success outcomes to a typedMachineLockUnavailableException: contention (held by an upgrade or another deploy) and Redis-unreachable. -
ExecuteStepsPhaseruns each per-machine script under the lock, held for the whole script (RedLock auto-extends) — so an upgrade for that machine observes contention and defers. - On
MachineLockUnavailableExceptionthe deployment pauses (resumable, checkpoint preserved) — never proceeds unguarded, never fails — and re-attempts the lock on resume.TargetCatchClassifierleaves the locked target non-terminal (resume retries it) and does not fail-fast peers; the runner routes it to the transient-pause path. A racing cancel/timeout still wins.
The upgrade side is unchanged (it already contends on the same key). Untagged ad-hoc script runs (health checks etc.) are unaffected. Reuses the proven Redis lock rather than inventing new machinery — consistent with the existing upgrade serialization.
Test plan
-
Unit (6034/6034): MachineDispatchLockmaps success/contention/Redis-down; deploy key == upgrade key pin;TargetCatchClassifier→ non-terminal + no fail-fast (and parent-cancel still fails); runner pauses (not fails) onMachineLockUnavailableException -
Integration (real Redis): a held lock blocks a concurrent same-machine claim with contention and frees on release; different machines don't block each other -
Solution build: 0 errors