Skip to content

feat(LinuxTentacleE2E): Phase 12.L.E.10 — E11.u1-Linux flock contention (concurrent dispatch no-op)

Summary

Linux mirror of Windows E11.u1 with the platform-correct lock mechanism: kernel BSD flock (not file-content detection).

Why ship-blocking

An operator manually triggering an upgrade while a scheduled upgrade is mid-flight (or two operators racing) MUST NOT cause two parallel Phase A's competing on the same INSTALL_DIR + STATE_DIR. The .sh's flock -n guard at line 288 is the only thing preventing this; without it, two simultaneous mv swaps would corrupt the install tree → unrecoverable agent.

Linux vs Windows mechanism

Platform Lock primitive Lockfile content Pre-stage sufficient?
Windows File-existence + PID liveness check Holds the lock state ✅
Linux Kernel BSD flock (held by an active process owning the FD) Informational only ❌ needs an active holder

So the Linux test must spawn a process that genuinely holds the kernel lock during the test window. Pre-staging an unattended file does nothing — the kernel auto-releases flocks on process death.

Test mechanism

using var flockHolder = ctx.StartFlockHolder(holdSeconds: 30);

Spawns a backgrounded flock -n -x $LOCK_FILE sleep 30 child process. Disposing the holder kills the child; kernel auto-releases the flock.

Expected behaviour (.sh line 288–295)

if ! flock -n "$LOCK_FD"; then
  echo "An upgrade is already in progress on this host..."
  exit 0
fi

Exit 0 as a NO-OP. Intentional: server-side dispatch is at-least-once, so duplicates are normal, not errors. Distinct from .sh's other exit codes (2/6/7/13/14) which are all fail-codes.

Assertions

  • exitCode == 0
  • stdout contains "upgrade is already in progress on this host" (operator visibility)
  • stdout does NOT contain "Detaching to systemd scope" (Phase A never spawned)
  • Marker stays at "1.0.0" (Phase B never ran)
  • .bak directory does NOT exist (mv-swap never executed)
  • flock holder still alive at end (sanity: lock genuinely held throughout)

Infrastructure additions

LinuxLifecycleContext:

  • StartFlockHolder(int holdSeconds) → spawns backgrounded flock child holding kernel lock
  • FlockHolder class wraps Process; Dispose kills + reaps; kernel auto-releases flock on process death even on test panic

Fidelity tier

🟢 High (Rule 12.4): real prod .sh + real systemd + real bash + real util-linux flock + real backgrounded sleep holder. Skipped on macOS via LinuxLifecycleContext.IsAvailable (no flock CLI on darwin).

Test plan

  • Linux E2E workflow runs Squid.LinuxTentacleE2ETests to completion (manual workflow_dispatch after merge)
  • E11u1_PreExistingFlock_LockHeldExitsZeroAsNoOp_Linux passes within ~5s
  • No regression on existing 11 Linux E2E tests

🤖 Generated with Claude Code

Merge request reports

Loading