# How Magnitude Handles ACN Daemon Recovery and Health Checks

> Learn how Magnitude handles ACN daemon recovery and health checks with its layered system, grace periods, and configurable RPC policies for reliable operation.

- Repository: [Magnitude/magnitude](https://github.com/magnitudedev/magnitude)
- Tags: internals
- Published: 2026-09-05

---

**Magnitude implements a layered recovery system that monitors ACN daemon health through periodic HTTP `/health` requests, applies grace periods to tolerate transient failures, and automatically recovers or gracefully shuts down daemons based on configurable RPC policies like `ReplaySafe` or `AtMostOnce`.**

The magnitudedev/magnitude repository orchestrates AI agent workflows through a persistent Agent Core Node (ACN) daemon that must remain reachable and in a consistent state. Understanding how Magnitude handles ACN daemon recovery and health checks reveals a sophisticated architecture that balances rapid failure detection with tolerance for transient network issues. The system combines proactive health polling, policy-driven client resilience, and automated lifecycle management to ensure continuous operation.

## Health Monitoring Infrastructure

### The `/health` Endpoint and Response Schema

Each ACN process exposes an HTTP `/health` endpoint that returns a JSON payload validated by `AcnHealthResponseSchema`. The response contains a `state` tag—such as `Ready`, `Starting`, or `Stopping`—along with a `revision` number that tracks the daemon's internal version state.

The `AcnLifecycle` class consumes this health observation stream in [`packages/sdk/src/acn-jit/lifecycle.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/acn-jit/lifecycle.ts). When health checks repeatedly fail, it emits diagnostic errors like `AcnHealthResponseInvalid` that surface in the UI without internal error tags, ensuring clear user-facing messages while maintaining type safety.

### Polling Strategy in AcnOwnerObserver

The `AcnOwnerObserver` class, located in [`packages/sdk/src/acn-jit/acn-owner-observer.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/acn-jit/acn-owner-observer.ts), implements a dual-request polling strategy to eliminate false positives from transient network failures. It issues a primary health request followed by a fallback request if the first fails.

The observer decodes responses using `AcnHealthResponseSchema`. Only when the status code is 200 and the state equals `Ready` does it return `Option.some` containing the health data; otherwise, it returns `Option.none` to signal an unhealthy state.

```typescript
// Observing health with fallback logic in the SDK
const healthAttempt = (owner: AcnOwner) =>
  http
    .execute(HttpClientRequest.get(`http://127.0.0.1:${owner.port}/health`))
    .pipe(Effect.map((response) => ({ status: response.status, health: response })));

const observed = yield* healthAttempt(owner).pipe(
  Effect.either, // first attempt
  Effect.flatMap((first) => first._tag === "Right" ? Effect.succeed(first) : healthAttempt(owner))
);

```

### Grace Periods and Convergence Logic

To prevent premature recovery actions, Magnitude implements time-based tolerance policies in [`packages/sdk/src/acn-jit/acn-convergence-decider.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/acn-jit/acn-convergence-decider.ts). The system defines `HEALTH_GRACE` and `STOPPING_GRACE` constants that dictate how long the daemon can remain in non-Ready or Stopping states before being declared unhealthy.

The `AcnEnsuranceCoordinator` in [`acn-ensurance-coordinator.ts`](https://github.com/magnitudedev/magnitude/blob/main/acn-ensurance-coordinator.ts) tracks the timestamp of the last healthy observation in `healthStateObservedAt`. It uses this value to drive convergence decisions about whether to launch a new daemon, maintain the current one, or initiate retirement.

## Recovery Policies and Client Resilience

### RPC Recovery Policy Resolution

RPC declarations in the ACN protocol boundary ([`packages/acn-protocol/src/boundary/rpc.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn-protocol/src/boundary/rpc.ts)) specify recovery behaviors through a `policy` object containing a `recovery` field. The `recoveryPolicy` function in [`packages/sdk/src/acn-jit/recovering-protocol.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/acn-jit/recovering-protocol.ts) maps these declarations to `AcnRpcRecoveryPolicy` values.

Queries default to `"ReplaySafe"`, while mutations use the explicitly declared policy. This distinction ensures that idempotent operations can be retried transparently while non-idempotent operations fail fast.

### AcnRecoveringClient State Machine

The `AcnRecoveringClient` class in [`packages/sdk/src/acn-jit/acn-recovering-client.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/acn-jit/acn-recovering-client.ts) wraps RPC calls with a `SubscriptionRef` tracking `AcnRecoveryState` (`AcnRecoveryInactive`, `AcnRecovering`, or `AcnRecovered`). When a call fails with a recoverable error, the client consults the policy:

- **ReplaySafe**: The request retries transparently, preserving idempotence.
- **AtMostOnce**: The error propagates immediately without retry.

The client publishes state changes to the `recovery.changes` stream, allowing UI components to react to recovery attempts in real time.

```typescript
// Recovery state handling in AcnRecoveringClient
const recoveryState = yield* SubscriptionRef.make<AcnRecoveryState>(new AcnRecoveryInactive({}));
yield* client.someRpc(request).pipe(
  Effect.catchAll((error) =>
    // Apply the recovery policy (ReplaySafe vs AtMostOnce)
    Option.isSome(error.recoveryOccurrence)
      ? Effect.succeed(error) // no retry for AtMostOnce
      : Effect.retry(error)   // retry for ReplaySafe
  ),
  Effect.tap(() => SubscriptionRef.set(recoveryState, new AcnRecovered({ occurrence: 1 })))
);

```

## Daemon Lifecycle Management

### Launch Supervision with AcnCandidateLaunchSupervisor

When health observations indicate the absence of a Ready daemon, the `AcnCandidateLaunchSupervisor` in [`packages/sdk/src/acn-jit/acn-candidate-launch-supervisor.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/acn-jit/acn-candidate-launch-supervisor.ts) evaluates convergence criteria to decide when to spawn a new ACN process. This component ensures the system maintains desired capacity without thrashing during brief health blips.

### Graceful Shutdown via AcnDaemonShutdownSupervisor

The `AcnDaemonShutdownSupervisor` in [`packages/sdk/src/acn-jit/acn-daemon-shutdown-supervisor.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/acn-jit/acn-daemon-shutdown-supervisor.ts) monitors for the `Stopping` state and enforces the `STOPPING_GRACE` period before invoking the `/shutdown` endpoint. This allows the daemon to clean up resources and transition cleanly before termination.

## UI Integration and Recovery Feedback

### Recovery State Atoms

The client-side state management in [`packages/client-common/src/state/acn-recovery.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/client-common/src/state/acn-recovery.ts) creates an `Atom` that subscribes to the `recovery.changes` stream from the SDK. React components consume this atom to display real-time ACN daemon recovery status without direct coupling to the polling logic.

### User Notification Rendering

[`packages/client-common/src/state/notification-area-state.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/client-common/src/state/notification-area-state.ts) transforms internal `AcnRecoveryState` values into human-readable messages using the `recoveryActivity` helper. This ensures users receive clear notifications like "ACN restarting..." or "ACN recovered after X attempts" without exposure to internal error tags.

```typescript
// UI: Subscribing to recovery notifications
const recovery = useAcnStartup().recovery;
useEffect(() => {
  const dispose = recovery.changes.runForEach((state) => {
    const msg = recoveryActivity(state);
    if (msg) showNotification(msg);
  });
  return () => dispose();
}, []);

```

## Release-Time Health Validation

Before promoting a new ACN binary, the [`scripts/accept-release-candidate.ts`](https://github.com/magnitudedev/magnitude/blob/main/scripts/accept-release-candidate.ts) script performs rigorous health validation. It verifies that the health response contains:
- `health.service` equal to `"magnitude-acn"`
- `health.version` matching the target manifest version
- `health.revision` matching the manifest's `acnRevision`
- `health.state._tag` equal to `"Ready"`

Any mismatch aborts the release, ensuring only healthy daemons reach production.

## Summary

- **Dual-request health polling** via `AcnOwnerObserver` in [`packages/sdk/src/acn-jit/acn-owner-observer.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/acn-jit/acn-owner-observer.ts) eliminates false positives by validating `AcnHealthResponseSchema` with fallback logic.
- **Grace periods** (`HEALTH_GRACE` and `STOPPING_GRACE`) in [`acn-convergence-decider.ts`](https://github.com/magnitudedev/magnitude/blob/main/acn-convergence-decider.ts) prevent premature recovery actions during transient failures.
- **Policy-driven resilience** through `AcnRecoveringClient` applies `ReplaySafe` retries for idempotent operations while respecting `AtMostOnce` constraints for mutations.
- **Automated lifecycle management** via `AcnCandidateLaunchSupervisor` and `AcnDaemonShutdownSupervisor` handles daemon spawning and graceful termination.
- **Reactive UI integration** using `Atom` subscriptions in [`acn-recovery.ts`](https://github.com/magnitudedev/magnitude/blob/main/acn-recovery.ts) provides real-time visibility into recovery states without tight coupling to polling internals.
- **Pre-deployment validation** in [`accept-release-candidate.ts`](https://github.com/magnitudedev/magnitude/blob/main/accept-release-candidate.ts) guarantees that only Ready-state daemons with matching versions and revisions enter production.

## Frequently Asked Questions

### What happens if the ACN daemon returns a non-Ready state temporarily?

Magnitude tolerates transient unhealthy states using configurable grace periods defined in [`acn-convergence-decider.ts`](https://github.com/magnitudedev/magnitude/blob/main/acn-convergence-decider.ts). The `AcnEnsuranceCoordinator` tracks the last healthy observation timestamp in `healthStateObservedAt` and only declares the daemon unhealthy if it remains in a non-`Ready` state longer than the `HEALTH_GRACE` period.

### How does Magnitude distinguish between replayable and non-replayable RPC failures?

The system inspects the `recovery` field in RPC declarations from [`packages/acn-protocol/src/boundary/rpc.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn-protocol/src/boundary/rpc.ts). The `recoveryPolicy` function in [`recovering-protocol.ts`](https://github.com/magnitudedev/magnitude/blob/main/recovering-protocol.ts) maps these to `AcnRpcRecoveryPolicy` values, where `ReplaySafe` allows transparent retries and `AtMostOnce` propagates errors immediately without retry to prevent duplicate mutations.

### Can users see when the ACN daemon is recovering in the UI?

Yes. The SDK exposes recovery state changes through a `SubscriptionRef` in `AcnRecoveringClient`, which [`packages/client-common/src/state/acn-recovery.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/client-common/src/state/acn-recovery.ts) mirrors into a reactive `Atom`. UI components subscribe to this state and display human-readable notifications via the `recoveryActivity` helper in [`notification-area-state.ts`](https://github.com/magnitudedev/magnitude/blob/main/notification-area-state.ts).

### What prevents an unhealthy daemon from being deployed in production?

The [`scripts/accept-release-candidate.ts`](https://github.com/magnitudedev/magnitude/blob/main/scripts/accept-release-candidate.ts) validation script performs pre-release health checks, verifying that the daemon's `service`, `version`, `revision`, and `state` fields match the target manifest. If any check fails, the script aborts the release pipeline, ensuring only healthy daemons reach production environments.