Skip to content

Serialize deploy and upgrade per machine across pods

Placeholder ppxd requested to merge fix/p3-deploy-upgrade-machine-mutex into main

Summary

Second of two PRs hardening multi-pod concurrency correctness (companion to the atomic task-slot in #441).

Problem. A tentacle upgrade restarts a machine's agent, so it must never run while a deployment's script is executing on that machine. Upgrades already take a per-machine Redis lock before dispatch (MachineUpgradeService.DispatchUnderLockAsync); deployments took none — so on a horizontally-scaled multi-pod server, a deploy and an upgrade for the same machine could run concurrently and the upgrade would restart the agent mid-script. (The #427 agent-side script-isolation mutex only serializes once both scripts reach the agent — it can't stop the server-side dispatch decision.)

Fix — the deployment dispatch acquires the SAME per-machine lock the upgrade uses:

  • New IMachineDispatchLock wraps the existing IRedisSafeRunner, running an action under the per-machine lock keyed exactly as UpgradeDispatchLockReconciler.BuildLockKey(machineId) — pinned by a unit test so deploy and upgrade can never drift onto different keys. It maps the runner's two non-success outcomes to a typed MachineLockUnavailableException: contention (held by an upgrade or another deploy) and Redis-unreachable.
  • ExecuteStepsPhase runs each per-machine script under the lock, held for the whole script (RedLock auto-extends) — so an upgrade for that machine observes contention and defers.
  • On MachineLockUnavailableException the deployment pauses (resumable, checkpoint preserved) — never proceeds unguarded, never fails — and re-attempts the lock on resume. TargetCatchClassifier leaves the locked target non-terminal (resume retries it) and does not fail-fast peers; the runner routes it to the transient-pause path. A racing cancel/timeout still wins.

The upgrade side is unchanged (it already contends on the same key). Untagged ad-hoc script runs (health checks etc.) are unaffected. Reuses the proven Redis lock rather than inventing new machinery — consistent with the existing upgrade serialization.

Test plan

  • Unit (6034/6034): MachineDispatchLock maps success/contention/Redis-down; deploy key == upgrade key pin; TargetCatchClassifier → non-terminal + no fail-fast (and parent-cancel still fails); runner pauses (not fails) on MachineLockUnavailableException
  • Integration (real Redis): a held lock blocks a concurrent same-machine claim with contention and frees on release; different machines don't block each other
  • Solution build: 0 errors

Merge request reports

Loading