How Prime Agent's Daemon Supervisor Manages Worker Health: A Deep Dive into Multi-Layered Process Monitoring

Prime Agent's daemon supervisor maintains worker health through a three-tier monitoring system combining protocol-level heartbeats, OS-level process probes, and policy-based idle eviction.

The supervisor orchestrates session workers via Unix-domain sockets, ensuring processes stay alive, responsive, and properly cleaned up when they become unhealthy. This article examines the implementation in PrimeIntellect-ai/prime-agent, focusing on the health management logic found in packages/coding-agent/src/modes/daemon/daemon-supervisor.ts and its companion modules.

Worker Registration and Initial Connection

Before health monitoring begins, workers must register with the supervisor. On startup, each worker writes a JSON descriptor file to the supervisor's descriptor directory. The supervisor loads these entries via loadWorkerDescriptors() (lines 13,099–13,142), creating a ResidentWorker record for each process.

Connection establishment follows this registration:

// Supervisor-side worker initialization
const client = new DaemonWorkerClient(workerSocketPath);
await client.connect();                     // Establish socket connection
const hello = await client.waitForHello();  // Wait for daemon_hello frame
worker.helloMessage = hello;

The DaemonWorkerClient class in daemon-worker-client.ts (lines 79–99) handles the TCP connection and hello handshake. Once connected, the client routes inbound frames to handleFrame() (lines 35–70), which parses protocol messages and updates worker state.

Protocol-Level Health: Heartbeat Exchange

Workers emit roster frames containing heartbeat arrays as defined in daemon-worker-protocol.ts. The supervisor extracts these in handleWorkerFrame():

// daemon-supervisor.ts
private async handleWorkerFrame(frame: PrivateFrame<DaemonWorkerFrameHeader>) {
  if (frame.header.outboundType === "roster") {
    const roster = JSON.parse(frame.payload.toString()) as DaemonRosterFrame;
    const worker = this.workers.get(frame.header.workerId);
    if (worker) {
      worker.heartbeatSnapshot = roster.heartbeat ?? [];
      worker.heartbeatSnapshotStale = false;
    }
  }
}

Each successful receipt clears worker.heartbeatSnapshotStale and caches the latest heartbeat data in worker.heartbeatSnapshot. This provides the first layer of health visibility—protocol responsiveness.

Detecting Stale Workers: The Watchdog Timer

The supervisor runs a roster watchdog timer (rosterWatchdogTimer) started in start(), firing every ROSTER_WATCHDOG_INTERVAL_MS (15,000ms). Each tick invokes sweepRosterStaleness() to identify unresponsive workers:

private sweepRosterStaleness() {
  const now = Date.now();
  for (const worker of this.workers.values()) {
    if (worker.heartbeatSnapshotStale ||
        (worker.lastFrameAt && now - worker.lastFrameAt > ROSTER_STALE_AFTER_MS)) {
      this.probeWorkerLiveness(worker);
    }
  }
}

A worker becomes stale when either:

  • heartbeatSnapshotStale remains true (no recent roster frame received)
  • lastFrameAt timestamp exceeds ROSTER_STALE_AFTER_MS (3× the heartbeat interval)

The scheduleIdleEvictionSweep() logic (lines 46–58) and rosterWatchdogTimer definition (lines 41–43) coordinate these checks.

Process-Level Health: Liveness Probing

Stale workers trigger OS-level verification via isProcessAlive() imported from packages/coding-agent/src/utils/child-process.ts (line 64). The probeWorkerLiveness() method wraps this with retry logic using WORKER_RETRY_DELAYS_MS:

  • Dead process confirmed → Immediate eviction via stopWorker()
  • Process alive but unresponsive → Deferred recovery tracking begins

The deferredRecoveryRounds counter (line 1,556) tracks how many recovery attempts a living-but-unresponsive worker receives before forced termination at MAX_DEFERRED_RECOVERY_ROUNDS.

Policy-Level Health: Idle Eviction

Beyond liveness, the supervisor evaluates business-level health through idle detection. The runIdleEvictionSweep() method (lines 90–124) periodically checks whether workers have exceeded configured idle limits:

private async runIdleEvictionSweep(now = Date.now()) {
  const idleMinutes = this.settingsManager.getIdleEvictionMinutes();
  if (idleMinutes === "off") return;
  
  const refreshed = new Set<ResidentWorker>();
  // Refresh summaries via refreshWorkerSummaries()
  
  const candidates = [...refreshed].filter(w =>
    canEvictWorker(this.workerEvictionSnapshot(w), idleMinutes, now));
  await Promise.all(candidates.map(w => this.stopWorker(w, true)));
}

The canEvictWorker() function in packages/coding-agent/src/core/session-action-store.ts evaluates session activity against the idleMinutes threshold. This allows administrators to set policy-driven cleanup independent of process health.

Graceful Shutdown and Update-Restart

During daemon-wide restarts, workers pass through a startup gate (WORKER_STARTUP_GATE_FD, line 176). The supervisor coordinates draining via mutationDrain.waitForDrain (lines 191–197), ensuring in-flight operations complete before workers restart. Workers failing to acknowledge the gate are forcibly stopped.

Summary

Prime Agent's daemon supervisor implements comprehensive worker health management through:

  • Registration and socket-based communication via JSON descriptors and DaemonWorkerClient
  • Protocol heartbeats embedded in roster frames with stale-snapshot detection
  • OS-level liveness probes using kill(0) with retry logic
  • Deferred recovery rounds for temporarily unresponsive but living processes
  • Idle eviction sweeps driven by configurable session policies
  • Graceful shutdown coordination through startup gates and mutation draining

These mechanisms operate in packages/coding-agent/src/modes/daemon/daemon-supervisor.ts, supported by protocol definitions in daemon-worker-protocol.ts and process utilities in child-process.ts.

Frequently Asked Questions

How does Prime Agent's daemon supervisor detect an unresponsive worker?

The supervisor uses a 15-second watchdog timer (rosterWatchdogTimer) that checks heartbeatSnapshotStale flags and lastFrameAt timestamps. Workers exceeding ROSTER_STALE_AFTER_MS (3× heartbeat interval) trigger probeWorkerLiveness(), which verifies the process with OS-level kill(0) calls and retry delays.

What happens when a worker process is alive but not sending heartbeats?

The supervisor increments deferredRecoveryRounds for that worker. After reaching MAX_DEFERRED_RECOVERY_ROUNDS, the worker is forcibly stopped via stopWorker(). This prevents zombie workers from consuming resources while allowing temporary network or load-related hiccups to resolve.

Can idle workers be terminated even if their heartbeats are healthy?

Yes. The runIdleEvictionSweep() operates independently of heartbeat monitoring. It checks session activity via canEvictWorker() in session-action-store.ts against settingsManager.getIdleEvictionMinutes(). Workers idle beyond the configured threshold are stopped regardless of protocol responsiveness.

Where is the worker health state stored in the supervisor?

Each worker's state lives in a ResidentWorker entry maintained in the supervisor's workers Map. Key fields include helloMessage (initial handshake), heartbeatSnapshot (latest roster data), heartbeatSnapshotStale (flag for watchdog checks), lastFrameAt (timestamp tracking), and deferredRecoveryRounds (recovery attempt counter).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →