# How Prime Agent's Daemon Supervisor Manages Worker Health: A Deep Dive into Multi-Layered Process Monitoring

> Discover how Prime Agent's daemon supervisor ensures worker health with multi-layered process monitoring, including heartbeats, OS probes, and idle eviction. Learn more now.

- Repository: [Prime Intellect/prime-agent](https://github.com/PrimeIntellect-ai/prime-agent)
- Tags: deep-dive
- Published: 2026-09-06

---

**Prime Agent's daemon supervisor maintains worker health through a three-tier monitoring system combining protocol-level heartbeats, OS-level process probes, and policy-based idle eviction.**

The supervisor orchestrates session workers via Unix-domain sockets, ensuring processes stay alive, responsive, and properly cleaned up when they become unhealthy. This article examines the implementation in `PrimeIntellect-ai/prime-agent`, focusing on the health management logic found in [`packages/coding-agent/src/modes/daemon/daemon-supervisor.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-supervisor.ts) and its companion modules.

## Worker Registration and Initial Connection

Before health monitoring begins, workers must register with the supervisor. On startup, each worker writes a JSON descriptor file to the supervisor's descriptor directory. The supervisor loads these entries via `loadWorkerDescriptors()` (lines 13,099–13,142), creating a `ResidentWorker` record for each process.

Connection establishment follows this registration:

```typescript
// Supervisor-side worker initialization
const client = new DaemonWorkerClient(workerSocketPath);
await client.connect();                     // Establish socket connection
const hello = await client.waitForHello();  // Wait for daemon_hello frame
worker.helloMessage = hello;

```

The `DaemonWorkerClient` class in [`daemon-worker-client.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-worker-client.ts) (lines 79–99) handles the TCP connection and hello handshake. Once connected, the client routes inbound frames to `handleFrame()` (lines 35–70), which parses protocol messages and updates worker state.

## Protocol-Level Health: Heartbeat Exchange

Workers emit **roster frames** containing heartbeat arrays as defined in [`daemon-worker-protocol.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-worker-protocol.ts). The supervisor extracts these in `handleWorkerFrame()`:

```typescript
// daemon-supervisor.ts
private async handleWorkerFrame(frame: PrivateFrame<DaemonWorkerFrameHeader>) {
  if (frame.header.outboundType === "roster") {
    const roster = JSON.parse(frame.payload.toString()) as DaemonRosterFrame;
    const worker = this.workers.get(frame.header.workerId);
    if (worker) {
      worker.heartbeatSnapshot = roster.heartbeat ?? [];
      worker.heartbeatSnapshotStale = false;
    }
  }
}

```

Each successful receipt clears `worker.heartbeatSnapshotStale` and caches the latest heartbeat data in `worker.heartbeatSnapshot`. This provides the first layer of health visibility—protocol responsiveness.

## Detecting Stale Workers: The Watchdog Timer

The supervisor runs a **roster watchdog timer** (`rosterWatchdogTimer`) started in `start()`, firing every `ROSTER_WATCHDOG_INTERVAL_MS` (15,000ms). Each tick invokes `sweepRosterStaleness()` to identify unresponsive workers:

```typescript
private sweepRosterStaleness() {
  const now = Date.now();
  for (const worker of this.workers.values()) {
    if (worker.heartbeatSnapshotStale ||
        (worker.lastFrameAt && now - worker.lastFrameAt > ROSTER_STALE_AFTER_MS)) {
      this.probeWorkerLiveness(worker);
    }
  }
}

```

A worker becomes **stale** when either:
- `heartbeatSnapshotStale` remains true (no recent roster frame received)
- `lastFrameAt` timestamp exceeds `ROSTER_STALE_AFTER_MS` (3× the heartbeat interval)

The `scheduleIdleEvictionSweep()` logic (lines 46–58) and `rosterWatchdogTimer` definition (lines 41–43) coordinate these checks.

## Process-Level Health: Liveness Probing

Stale workers trigger **OS-level verification** via `isProcessAlive()` imported from [`packages/coding-agent/src/utils/child-process.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/utils/child-process.ts) (line 64). The `probeWorkerLiveness()` method wraps this with retry logic using `WORKER_RETRY_DELAYS_MS`:

- **Dead process confirmed** → Immediate eviction via `stopWorker()`
- **Process alive but unresponsive** → Deferred recovery tracking begins

The `deferredRecoveryRounds` counter (line 1,556) tracks how many recovery attempts a living-but-unresponsive worker receives before forced termination at `MAX_DEFERRED_RECOVERY_ROUNDS`.

## Policy-Level Health: Idle Eviction

Beyond liveness, the supervisor evaluates **business-level health** through idle detection. The `runIdleEvictionSweep()` method (lines 90–124) periodically checks whether workers have exceeded configured idle limits:

```typescript
private async runIdleEvictionSweep(now = Date.now()) {
  const idleMinutes = this.settingsManager.getIdleEvictionMinutes();
  if (idleMinutes === "off") return;
  
  const refreshed = new Set<ResidentWorker>();
  // Refresh summaries via refreshWorkerSummaries()
  
  const candidates = [...refreshed].filter(w =>
    canEvictWorker(this.workerEvictionSnapshot(w), idleMinutes, now));
  await Promise.all(candidates.map(w => this.stopWorker(w, true)));
}

```

The `canEvictWorker()` function in [`packages/coding-agent/src/core/session-action-store.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/core/session-action-store.ts) evaluates session activity against the `idleMinutes` threshold. This allows administrators to set policy-driven cleanup independent of process health.

## Graceful Shutdown and Update-Restart

During daemon-wide restarts, workers pass through a **startup gate** (`WORKER_STARTUP_GATE_FD`, line 176). The supervisor coordinates draining via `mutationDrain.waitForDrain` (lines 191–197), ensuring in-flight operations complete before workers restart. Workers failing to acknowledge the gate are forcibly stopped.

## Summary

Prime Agent's daemon supervisor implements **comprehensive worker health management** through:

- **Registration and socket-based communication** via JSON descriptors and `DaemonWorkerClient`
- **Protocol heartbeats** embedded in roster frames with stale-snapshot detection
- **OS-level liveness probes** using `kill(0)` with retry logic
- **Deferred recovery rounds** for temporarily unresponsive but living processes
- **Idle eviction sweeps** driven by configurable session policies
- **Graceful shutdown coordination** through startup gates and mutation draining

These mechanisms operate in [`packages/coding-agent/src/modes/daemon/daemon-supervisor.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-supervisor.ts), supported by protocol definitions in [`daemon-worker-protocol.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-worker-protocol.ts) and process utilities in [`child-process.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/child-process.ts).

## Frequently Asked Questions

### How does Prime Agent's daemon supervisor detect an unresponsive worker?

The supervisor uses a 15-second watchdog timer (`rosterWatchdogTimer`) that checks `heartbeatSnapshotStale` flags and `lastFrameAt` timestamps. Workers exceeding `ROSTER_STALE_AFTER_MS` (3× heartbeat interval) trigger `probeWorkerLiveness()`, which verifies the process with OS-level `kill(0)` calls and retry delays.

### What happens when a worker process is alive but not sending heartbeats?

The supervisor increments `deferredRecoveryRounds` for that worker. After reaching `MAX_DEFERRED_RECOVERY_ROUNDS`, the worker is forcibly stopped via `stopWorker()`. This prevents zombie workers from consuming resources while allowing temporary network or load-related hiccups to resolve.

### Can idle workers be terminated even if their heartbeats are healthy?

Yes. The `runIdleEvictionSweep()` operates independently of heartbeat monitoring. It checks session activity via `canEvictWorker()` in [`session-action-store.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/session-action-store.ts) against `settingsManager.getIdleEvictionMinutes()`. Workers idle beyond the configured threshold are stopped regardless of protocol responsiveness.

### Where is the worker health state stored in the supervisor?

Each worker's state lives in a `ResidentWorker` entry maintained in the supervisor's `workers` Map. Key fields include `helloMessage` (initial handshake), `heartbeatSnapshot` (latest roster data), `heartbeatSnapshotStale` (flag for watchdog checks), `lastFrameAt` (timestamp tracking), and `deferredRecoveryRounds` (recovery attempt counter).