How Magnitude Handles ACN Daemon Recovery and Health Checks
Magnitude implements a layered recovery system that monitors ACN daemon health through periodic HTTP /health requests, applies grace periods to tolerate transient failures, and automatically recovers or gracefully shuts down daemons based on configurable RPC policies like ReplaySafe or AtMostOnce.
The magnitudedev/magnitude repository orchestrates AI agent workflows through a persistent Agent Core Node (ACN) daemon that must remain reachable and in a consistent state. Understanding how Magnitude handles ACN daemon recovery and health checks reveals a sophisticated architecture that balances rapid failure detection with tolerance for transient network issues. The system combines proactive health polling, policy-driven client resilience, and automated lifecycle management to ensure continuous operation.
Health Monitoring Infrastructure
The /health Endpoint and Response Schema
Each ACN process exposes an HTTP /health endpoint that returns a JSON payload validated by AcnHealthResponseSchema. The response contains a state tag—such as Ready, Starting, or Stopping—along with a revision number that tracks the daemon's internal version state.
The AcnLifecycle class consumes this health observation stream in packages/sdk/src/acn-jit/lifecycle.ts. When health checks repeatedly fail, it emits diagnostic errors like AcnHealthResponseInvalid that surface in the UI without internal error tags, ensuring clear user-facing messages while maintaining type safety.
Polling Strategy in AcnOwnerObserver
The AcnOwnerObserver class, located in packages/sdk/src/acn-jit/acn-owner-observer.ts, implements a dual-request polling strategy to eliminate false positives from transient network failures. It issues a primary health request followed by a fallback request if the first fails.
The observer decodes responses using AcnHealthResponseSchema. Only when the status code is 200 and the state equals Ready does it return Option.some containing the health data; otherwise, it returns Option.none to signal an unhealthy state.
// Observing health with fallback logic in the SDK
const healthAttempt = (owner: AcnOwner) =>
http
.execute(HttpClientRequest.get(`http://127.0.0.1:${owner.port}/health`))
.pipe(Effect.map((response) => ({ status: response.status, health: response })));
const observed = yield* healthAttempt(owner).pipe(
Effect.either, // first attempt
Effect.flatMap((first) => first._tag === "Right" ? Effect.succeed(first) : healthAttempt(owner))
);
Grace Periods and Convergence Logic
To prevent premature recovery actions, Magnitude implements time-based tolerance policies in packages/sdk/src/acn-jit/acn-convergence-decider.ts. The system defines HEALTH_GRACE and STOPPING_GRACE constants that dictate how long the daemon can remain in non-Ready or Stopping states before being declared unhealthy.
The AcnEnsuranceCoordinator in acn-ensurance-coordinator.ts tracks the timestamp of the last healthy observation in healthStateObservedAt. It uses this value to drive convergence decisions about whether to launch a new daemon, maintain the current one, or initiate retirement.
Recovery Policies and Client Resilience
RPC Recovery Policy Resolution
RPC declarations in the ACN protocol boundary (packages/acn-protocol/src/boundary/rpc.ts) specify recovery behaviors through a policy object containing a recovery field. The recoveryPolicy function in packages/sdk/src/acn-jit/recovering-protocol.ts maps these declarations to AcnRpcRecoveryPolicy values.
Queries default to "ReplaySafe", while mutations use the explicitly declared policy. This distinction ensures that idempotent operations can be retried transparently while non-idempotent operations fail fast.
AcnRecoveringClient State Machine
The AcnRecoveringClient class in packages/sdk/src/acn-jit/acn-recovering-client.ts wraps RPC calls with a SubscriptionRef tracking AcnRecoveryState (AcnRecoveryInactive, AcnRecovering, or AcnRecovered). When a call fails with a recoverable error, the client consults the policy:
- ReplaySafe: The request retries transparently, preserving idempotence.
- AtMostOnce: The error propagates immediately without retry.
The client publishes state changes to the recovery.changes stream, allowing UI components to react to recovery attempts in real time.
// Recovery state handling in AcnRecoveringClient
const recoveryState = yield* SubscriptionRef.make<AcnRecoveryState>(new AcnRecoveryInactive({}));
yield* client.someRpc(request).pipe(
Effect.catchAll((error) =>
// Apply the recovery policy (ReplaySafe vs AtMostOnce)
Option.isSome(error.recoveryOccurrence)
? Effect.succeed(error) // no retry for AtMostOnce
: Effect.retry(error) // retry for ReplaySafe
),
Effect.tap(() => SubscriptionRef.set(recoveryState, new AcnRecovered({ occurrence: 1 })))
);
Daemon Lifecycle Management
Launch Supervision with AcnCandidateLaunchSupervisor
When health observations indicate the absence of a Ready daemon, the AcnCandidateLaunchSupervisor in packages/sdk/src/acn-jit/acn-candidate-launch-supervisor.ts evaluates convergence criteria to decide when to spawn a new ACN process. This component ensures the system maintains desired capacity without thrashing during brief health blips.
Graceful Shutdown via AcnDaemonShutdownSupervisor
The AcnDaemonShutdownSupervisor in packages/sdk/src/acn-jit/acn-daemon-shutdown-supervisor.ts monitors for the Stopping state and enforces the STOPPING_GRACE period before invoking the /shutdown endpoint. This allows the daemon to clean up resources and transition cleanly before termination.
UI Integration and Recovery Feedback
Recovery State Atoms
The client-side state management in packages/client-common/src/state/acn-recovery.ts creates an Atom that subscribes to the recovery.changes stream from the SDK. React components consume this atom to display real-time ACN daemon recovery status without direct coupling to the polling logic.
User Notification Rendering
packages/client-common/src/state/notification-area-state.ts transforms internal AcnRecoveryState values into human-readable messages using the recoveryActivity helper. This ensures users receive clear notifications like "ACN restarting..." or "ACN recovered after X attempts" without exposure to internal error tags.
// UI: Subscribing to recovery notifications
const recovery = useAcnStartup().recovery;
useEffect(() => {
const dispose = recovery.changes.runForEach((state) => {
const msg = recoveryActivity(state);
if (msg) showNotification(msg);
});
return () => dispose();
}, []);
Release-Time Health Validation
Before promoting a new ACN binary, the scripts/accept-release-candidate.ts script performs rigorous health validation. It verifies that the health response contains:
health.serviceequal to"magnitude-acn"health.versionmatching the target manifest versionhealth.revisionmatching the manifest'sacnRevisionhealth.state._tagequal to"Ready"
Any mismatch aborts the release, ensuring only healthy daemons reach production.
Summary
- Dual-request health polling via
AcnOwnerObserverinpackages/sdk/src/acn-jit/acn-owner-observer.tseliminates false positives by validatingAcnHealthResponseSchemawith fallback logic. - Grace periods (
HEALTH_GRACEandSTOPPING_GRACE) inacn-convergence-decider.tsprevent premature recovery actions during transient failures. - Policy-driven resilience through
AcnRecoveringClientappliesReplaySaferetries for idempotent operations while respectingAtMostOnceconstraints for mutations. - Automated lifecycle management via
AcnCandidateLaunchSupervisorandAcnDaemonShutdownSupervisorhandles daemon spawning and graceful termination. - Reactive UI integration using
Atomsubscriptions inacn-recovery.tsprovides real-time visibility into recovery states without tight coupling to polling internals. - Pre-deployment validation in
accept-release-candidate.tsguarantees that only Ready-state daemons with matching versions and revisions enter production.
Frequently Asked Questions
What happens if the ACN daemon returns a non-Ready state temporarily?
Magnitude tolerates transient unhealthy states using configurable grace periods defined in acn-convergence-decider.ts. The AcnEnsuranceCoordinator tracks the last healthy observation timestamp in healthStateObservedAt and only declares the daemon unhealthy if it remains in a non-Ready state longer than the HEALTH_GRACE period.
How does Magnitude distinguish between replayable and non-replayable RPC failures?
The system inspects the recovery field in RPC declarations from packages/acn-protocol/src/boundary/rpc.ts. The recoveryPolicy function in recovering-protocol.ts maps these to AcnRpcRecoveryPolicy values, where ReplaySafe allows transparent retries and AtMostOnce propagates errors immediately without retry to prevent duplicate mutations.
Can users see when the ACN daemon is recovering in the UI?
Yes. The SDK exposes recovery state changes through a SubscriptionRef in AcnRecoveringClient, which packages/client-common/src/state/acn-recovery.ts mirrors into a reactive Atom. UI components subscribe to this state and display human-readable notifications via the recoveryActivity helper in notification-area-state.ts.
What prevents an unhealthy daemon from being deployed in production?
The scripts/accept-release-candidate.ts validation script performs pre-release health checks, verifying that the daemon's service, version, revision, and state fields match the target manifest. If any check fails, the script aborts the release pipeline, ensuring only healthy daemons reach production environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →