How Magnitude Handles ACN Daemon Recovery and Health Checks

Magnitude implements a layered recovery system that monitors ACN daemon health through periodic HTTP /health requests, applies grace periods to tolerate transient failures, and automatically recovers or gracefully shuts down daemons based on configurable RPC policies like ReplaySafe or AtMostOnce.

The magnitudedev/magnitude repository orchestrates AI agent workflows through a persistent Agent Core Node (ACN) daemon that must remain reachable and in a consistent state. Understanding how Magnitude handles ACN daemon recovery and health checks reveals a sophisticated architecture that balances rapid failure detection with tolerance for transient network issues. The system combines proactive health polling, policy-driven client resilience, and automated lifecycle management to ensure continuous operation.

Health Monitoring Infrastructure

The /health Endpoint and Response Schema

Each ACN process exposes an HTTP /health endpoint that returns a JSON payload validated by AcnHealthResponseSchema. The response contains a state tag—such as Ready, Starting, or Stopping—along with a revision number that tracks the daemon's internal version state.

The AcnLifecycle class consumes this health observation stream in packages/sdk/src/acn-jit/lifecycle.ts. When health checks repeatedly fail, it emits diagnostic errors like AcnHealthResponseInvalid that surface in the UI without internal error tags, ensuring clear user-facing messages while maintaining type safety.

Polling Strategy in AcnOwnerObserver

The AcnOwnerObserver class, located in packages/sdk/src/acn-jit/acn-owner-observer.ts, implements a dual-request polling strategy to eliminate false positives from transient network failures. It issues a primary health request followed by a fallback request if the first fails.

The observer decodes responses using AcnHealthResponseSchema. Only when the status code is 200 and the state equals Ready does it return Option.some containing the health data; otherwise, it returns Option.none to signal an unhealthy state.

// Observing health with fallback logic in the SDK
const healthAttempt = (owner: AcnOwner) =>
  http
    .execute(HttpClientRequest.get(`http://127.0.0.1:${owner.port}/health`))
    .pipe(Effect.map((response) => ({ status: response.status, health: response })));

const observed = yield* healthAttempt(owner).pipe(
  Effect.either, // first attempt
  Effect.flatMap((first) => first._tag === "Right" ? Effect.succeed(first) : healthAttempt(owner))
);

Grace Periods and Convergence Logic

To prevent premature recovery actions, Magnitude implements time-based tolerance policies in packages/sdk/src/acn-jit/acn-convergence-decider.ts. The system defines HEALTH_GRACE and STOPPING_GRACE constants that dictate how long the daemon can remain in non-Ready or Stopping states before being declared unhealthy.

The AcnEnsuranceCoordinator in acn-ensurance-coordinator.ts tracks the timestamp of the last healthy observation in healthStateObservedAt. It uses this value to drive convergence decisions about whether to launch a new daemon, maintain the current one, or initiate retirement.

Recovery Policies and Client Resilience

RPC Recovery Policy Resolution

RPC declarations in the ACN protocol boundary (packages/acn-protocol/src/boundary/rpc.ts) specify recovery behaviors through a policy object containing a recovery field. The recoveryPolicy function in packages/sdk/src/acn-jit/recovering-protocol.ts maps these declarations to AcnRpcRecoveryPolicy values.

Queries default to "ReplaySafe", while mutations use the explicitly declared policy. This distinction ensures that idempotent operations can be retried transparently while non-idempotent operations fail fast.

AcnRecoveringClient State Machine

The AcnRecoveringClient class in packages/sdk/src/acn-jit/acn-recovering-client.ts wraps RPC calls with a SubscriptionRef tracking AcnRecoveryState (AcnRecoveryInactive, AcnRecovering, or AcnRecovered). When a call fails with a recoverable error, the client consults the policy:

  • ReplaySafe: The request retries transparently, preserving idempotence.
  • AtMostOnce: The error propagates immediately without retry.

The client publishes state changes to the recovery.changes stream, allowing UI components to react to recovery attempts in real time.

// Recovery state handling in AcnRecoveringClient
const recoveryState = yield* SubscriptionRef.make<AcnRecoveryState>(new AcnRecoveryInactive({}));
yield* client.someRpc(request).pipe(
  Effect.catchAll((error) =>
    // Apply the recovery policy (ReplaySafe vs AtMostOnce)
    Option.isSome(error.recoveryOccurrence)
      ? Effect.succeed(error) // no retry for AtMostOnce
      : Effect.retry(error)   // retry for ReplaySafe
  ),
  Effect.tap(() => SubscriptionRef.set(recoveryState, new AcnRecovered({ occurrence: 1 })))
);

Daemon Lifecycle Management

Launch Supervision with AcnCandidateLaunchSupervisor

When health observations indicate the absence of a Ready daemon, the AcnCandidateLaunchSupervisor in packages/sdk/src/acn-jit/acn-candidate-launch-supervisor.ts evaluates convergence criteria to decide when to spawn a new ACN process. This component ensures the system maintains desired capacity without thrashing during brief health blips.

Graceful Shutdown via AcnDaemonShutdownSupervisor

The AcnDaemonShutdownSupervisor in packages/sdk/src/acn-jit/acn-daemon-shutdown-supervisor.ts monitors for the Stopping state and enforces the STOPPING_GRACE period before invoking the /shutdown endpoint. This allows the daemon to clean up resources and transition cleanly before termination.

UI Integration and Recovery Feedback

Recovery State Atoms

The client-side state management in packages/client-common/src/state/acn-recovery.ts creates an Atom that subscribes to the recovery.changes stream from the SDK. React components consume this atom to display real-time ACN daemon recovery status without direct coupling to the polling logic.

User Notification Rendering

packages/client-common/src/state/notification-area-state.ts transforms internal AcnRecoveryState values into human-readable messages using the recoveryActivity helper. This ensures users receive clear notifications like "ACN restarting..." or "ACN recovered after X attempts" without exposure to internal error tags.

// UI: Subscribing to recovery notifications
const recovery = useAcnStartup().recovery;
useEffect(() => {
  const dispose = recovery.changes.runForEach((state) => {
    const msg = recoveryActivity(state);
    if (msg) showNotification(msg);
  });
  return () => dispose();
}, []);

Release-Time Health Validation

Before promoting a new ACN binary, the scripts/accept-release-candidate.ts script performs rigorous health validation. It verifies that the health response contains:

  • health.service equal to "magnitude-acn"
  • health.version matching the target manifest version
  • health.revision matching the manifest's acnRevision
  • health.state._tag equal to "Ready"

Any mismatch aborts the release, ensuring only healthy daemons reach production.

Summary

  • Dual-request health polling via AcnOwnerObserver in packages/sdk/src/acn-jit/acn-owner-observer.ts eliminates false positives by validating AcnHealthResponseSchema with fallback logic.
  • Grace periods (HEALTH_GRACE and STOPPING_GRACE) in acn-convergence-decider.ts prevent premature recovery actions during transient failures.
  • Policy-driven resilience through AcnRecoveringClient applies ReplaySafe retries for idempotent operations while respecting AtMostOnce constraints for mutations.
  • Automated lifecycle management via AcnCandidateLaunchSupervisor and AcnDaemonShutdownSupervisor handles daemon spawning and graceful termination.
  • Reactive UI integration using Atom subscriptions in acn-recovery.ts provides real-time visibility into recovery states without tight coupling to polling internals.
  • Pre-deployment validation in accept-release-candidate.ts guarantees that only Ready-state daemons with matching versions and revisions enter production.

Frequently Asked Questions

What happens if the ACN daemon returns a non-Ready state temporarily?

Magnitude tolerates transient unhealthy states using configurable grace periods defined in acn-convergence-decider.ts. The AcnEnsuranceCoordinator tracks the last healthy observation timestamp in healthStateObservedAt and only declares the daemon unhealthy if it remains in a non-Ready state longer than the HEALTH_GRACE period.

How does Magnitude distinguish between replayable and non-replayable RPC failures?

The system inspects the recovery field in RPC declarations from packages/acn-protocol/src/boundary/rpc.ts. The recoveryPolicy function in recovering-protocol.ts maps these to AcnRpcRecoveryPolicy values, where ReplaySafe allows transparent retries and AtMostOnce propagates errors immediately without retry to prevent duplicate mutations.

Can users see when the ACN daemon is recovering in the UI?

Yes. The SDK exposes recovery state changes through a SubscriptionRef in AcnRecoveringClient, which packages/client-common/src/state/acn-recovery.ts mirrors into a reactive Atom. UI components subscribe to this state and display human-readable notifications via the recoveryActivity helper in notification-area-state.ts.

What prevents an unhealthy daemon from being deployed in production?

The scripts/accept-release-candidate.ts validation script performs pre-release health checks, verifying that the daemon's service, version, revision, and state fields match the target manifest. If any check fails, the script aborts the release pipeline, ensuring only healthy daemons reach production environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →