How Magnitude Handles Concurrent ACN Launches and Process Coordination
Magnitude guarantees that only a single ACN (agent container) is active for a given target at any moment by serializing all launch attempts through a finite-state machine supervised by an Effect-based semaphore.
The Magnitude daemon orchestrates ACN (agent container) lifecycle management through a carefully designed coordination pipeline. Understanding how concurrent ACN launches are handled requires examining four interconnected components that together eliminate race conditions while maintaining fault tolerance.
The Four Core Coordination Components
Magnitude's ACN coordination architecture separates concerns across four specialized modules in packages/daemon-management/src/acn-jit/:
AcnOwnerObserver: Continuous State Inspection
The AcnOwnerObserver (acn-owner-observer.ts) continuously reads the current ACN owner record from the owner store and probes the actual process group via ProcessGroupController. It normalizes the observed state into one of four typed classifications:
AcnRecordedOwnerAbsent– no owner record existsAcnRecordedOwnerProcessGroupSurvives– record points to a living process groupAcnRecordedOwnerLiveWithoutHealth– process responds but health check pendingAcnRecordedOwnerLiveWithHealth– fully verified healthy owner
This abstraction layer isolates transient failures and provides clean data for downstream decisions.
AcnCandidateLaunchSupervisor: FSM-Driven Process Control
The AcnCandidateLaunchSupervisor (acn-candidate-launch-supervisor.ts) encapsulates a finite-state machine tracking candidate lifecycle:
NotLaunched → Spawned → Admitted → Ready
↘
Failed
Critical to concurrent safety, the supervisor protects state transitions with Effect.makeSemaphore(1):
// From acn-candidate-launch-supervisor.ts
const lock = yield* Effect.makeSemaphore(1)
// All mutate operations acquire this exclusive lock
const launch = (cmd: LaunchCommand) =>
lock.withPermits(1)(Effect.gen(function* () {
// FSM state check, spawn, record identity, monitor...
}))
This semaphore ensures that concurrent calls to launch() or reconcile() from multiple sources queue sequentially rather than execute in parallel.
AcnConvergenceDecider: Pure Functional Decision Logic
The AcnConvergenceDecider (acn-convergence-decider.ts) implements decideAcnConvergence()—a pure function that receives a snapshot containing:
- Current owner observation
- Candidate launch state
- Timing information (
elapsedMs,now)
It returns deterministic actions:
| Action | Trigger Condition |
|---|---|
Wait |
State still converging; defer to next reconciliation |
PrepareLaunch |
Resolve command but defer process spawn |
LaunchCandidate |
Start new candidate process |
ShutdownDaemon / ShutdownDaemonThenFail |
Terminate outdated/unhealthy daemon |
ConfirmReady |
Promote admitted candidate to active owner |
FailCandidate |
Abort on unrecoverable errors |
Because this logic is pure (no side effects), concurrent observations with identical snapshots always yield identical decisions.
AcnEnsuranceCoordinator: The Central Reconciliation Loop
The AcnEnsuranceCoordinator (acn-ensurance-coordinator.ts) orchestrates the complete lifecycle by executing this loop every RECONCILIATION_INTERVAL (1 second):
// Simplified coordination loop structure
const loop = Effect.gen(function* () {
const ownerObservation = yield* ownerObserver.observe(target)
yield* updateConvergenceMemory(ownerObservation)
const candidateState = yield* candidateSupervisor.reconcile(
ownerObservation,
convergenceMemory
)
const action = decideAcnConvergence({
ownerObservation,
candidateState,
timing: { elapsedMs, now }
})
return yield* executeAction(action)
}).pipe(
Effect.timeout(ACN_ENSURE_TIMEOUT),
Effect.repeat({ schedule: Schedule.fixed(RECONCILIATION_INTERVAL) })
)
All state modifications use Ref and SubscriptionRef primitives from Effect, making concurrent updates safe by construction.
Process Coordination Mechanics
Process Group Validation
The ProcessGroupController (from @magnitudedev/acn-protocol/coordination) tracks both the leader PID and process-start identity. The observer uses this to verify that a running process truly belongs to the recorded owner—not a stale PID that has been recycled by the OS.
The sameAcnOwner() function performs strict equality checks on both identity components before confirming ownership.
Ownership Record Integrity
Owner records in the store contain:
- Target identifier (pid, revision, port)
- Process start identity (unique per spawn)
- Health check endpoint
This prevents the ABA problem where a process dies and a new process coincidentally receives the same PID.
Complete Setup Example
import { makeAcnOwnerObserver } from "@magnitudedev/daemon-management/src/acn-jit/acn-owner-observer";
import { makeAcnCandidateLaunchSupervisor } from "@magnitudedev/daemon-management/src/acn-jit/acn-candidate-launch-supervisor";
import { makeAcnEnsuranceCoordinator } from "@magnitudedev/daemon-management/src/acn-jit/acn-ensurance-coordinator";
// Initialize required dependencies
const ownerStore = /* AcnOwnerStore implementation */;
const procController = /* ProcessGroupController implementation */;
const httpClient = /* Effect Platform HttpClient */;
const childSpawner = /* ChildProcessSpawner implementation */;
const launchResolver = /* AcnDaemonLaunchCommandResolver implementation */;
const shutdownSup = /* AcnDaemonShutdownSupervisor implementation */;
// Build coordination components
const ownerObserver = makeAcnOwnerObserver(ownerStore, procController, httpClient);
const candidateSup = await makeAcnCandidateLaunchSupervisor(childSpawner, procController);
// Assemble coordinator with target configuration
const coordinator = await makeAcnEnsuranceCoordinator({
target: { pid: 0, revision: 1, port: 8080 },
emit: (e) => console.log("Ensurance event:", e),
debug: true,
dataDirectory: "/var/magnitude/acn",
ownerObserver,
shutdownSupervisor: shutdownSup,
candidateSupervisor: candidateSup,
launchCommandResolver: launchResolver,
});
// Start coordination loop—handles concurrent attempts automatically
coordinator.run.runPromise()
.then((readyInstance) => console.log("ACN ready:", readyInstance))
.catch((err) => console.error("Ensurance failed:", err));
Key Source Files
| Component | Path |
|---|---|
| Owner observation | packages/daemon-management/src/acn-jit/acn-owner-observer.ts |
| Candidate supervisor | packages/daemon-management/src/acn-jit/acn-candidate-launch-supervisor.ts |
| Decision logic | packages/daemon-management/src/acn-jit/acn-convergence-decider.ts |
| Main coordinator | packages/daemon-management/src/acn-jit/acn-ensurance-coordinator.ts |
| Protocol definitions | packages/acn-protocol/ |
Summary
- Single ownership guarantee: The semaphore in
AcnCandidateLaunchSupervisorserializes all launch operations, preventing duplicate ACN processes. - Deterministic convergence: Pure
decideAcnConvergence()function ensures consistent decisions across concurrent reconciliation loops. - Process identity verification:
ProcessGroupControllerwith start-identity tracking eliminates PID recycling hazards. - Fault isolation: Typed observations (
AcnRecordedOwner*) separate transient failures from permanent states, preventing unnecessary process kills. - Graceful degradation: Timeouts and staged FSM transitions (
Spawned→Admitted→Ready) allow health verification before promoting candidates.
Frequently Asked Questions
What happens if two parts of the system try to launch an ACN simultaneously?
Both attempts queue on the AcnCandidateLaunchSupervisor semaphore. The first acquires the lock and proceeds through the FSM; the second waits, then re-evaluates on its turn. By that point, ownerObserver.observe() will likely detect the first candidate, causing the second attempt to yield or retry rather than spawn a duplicate.
How does Magnitude prevent a stale PID from being mistaken for the current owner?
The sameAcnOwner() function validates both the PID and the process-start identity—a unique token generated at spawn time. Even if the OS recycles a PID, the identity mismatch prevents false ownership confirmation.
What is the reconciliation interval and can it be tuned?
The default RECONCILIATION_INTERVAL is 1 second, defined in acn-ensurance-coordinator.ts. This balances responsiveness against system load. The constant can be modified at compile time; shorter intervals detect failures faster but increase polling overhead.
How does the system recover from a failed candidate launch?
The FSM transitions to Failed state, which decideAcnConvergence() detects. Depending on timing and retry policy, the decider may emit PrepareLaunch or LaunchCandidate to start a fresh attempt. The ACN_ENSURE_TIMEOUT prevents infinite retry loops by surfacing terminal failures to the caller.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →