# How the Daemon Supervisor Recovers Workers After a Crash in Prime Agent

> Discover how the daemon supervisor in Prime Agent recovers crashed workers using its three-stage pipeline: JSON descriptors, live process adoption, and journal replay or respawning.

- Repository: [Prime Intellect/prime-agent](https://github.com/PrimeIntellect-ai/prime-agent)
- Tags: internals
- Published: 2026-09-05

---

**The daemon supervisor implements a three-stage recovery pipeline—loading persisted JSON descriptors, attempting to adopt live worker processes, and replaying recovery journals or respawning failed workers—to guarantee session continuity after unexpected crashes.**

When the PrimeIntellect-ai/prime-agent daemon supervisor restarts after a crash, it must reconstruct the state of all previously running session workers without losing active sessions. According to the source code in [`packages/coding-agent/src/modes/daemon/daemon-supervisor.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-supervisor.ts), the supervisor achieves this through a deterministic "adopt-or-recover-or-respawn" strategy that relies on disk-persisted worker descriptors and command journals rather than volatile in-memory state.

## The Three-Stage Worker Recovery Pipeline

The recovery process executes sequentially during supervisor startup. Each stage is designed to maximize the chance of restoring the exact pre-crash state while providing a clean fallback to fresh workers when necessary.

### Stage 1: Loading Persisted Worker Descriptors

On initialization, the supervisor reads every `*.json` file from its `descriptorDir` directory. These files contain the last known state of each worker, including the **PID**, socket path, `createCommand` that launched the process, and paths to recovery journals. The `loadWorkerDescriptors()` method (lines 13,000–13,040) deserializes these files into `ResidentWorker` objects and stores them in `this.workers`, ensuring the supervisor never depends on ephemeral memory to recover worker identity.

```typescript
private loadWorkerDescriptors(): void {
    for (const name of readdirSync(this.descriptorDir)) {
        if (name === SUPERVISOR_CONFIG_FILE_NAME || !name.endsWith(".json")) continue;
        const path = join(this.descriptorDir, name);
        const descriptor: unknown = JSON.parse(readFileSync(path, "utf8"));
        if (!isDaemonWorkerDescriptor(descriptor, this.socketPath)) continue;
        // …populate ResidentWorker and store in this.workers
    }
}

```

### Stage 2: Adoption of Live Workers

For each loaded descriptor, the supervisor calls `adoptOrRecoverWorker()` (lines 9,000–9,150) to determine if the original process survived the crash. The helper first invokes `isProcessAlive()` from [`utils/child-process.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/utils/child-process.ts) to verify the recorded PID. If the process is alive and the socket can be opened, the supervisor attaches to the existing worker using `DaemonWorkerClient`, repopulating the worker’s roster and snapshot cache from the descriptor data.

```typescript
await Promise.all(
    workersToAdopt.map(async (worker) => {
        try {
            await this.adoptOrRecoverWorker(worker);   // core recovery logic
        } catch (error) {
            // If any worker cannot be adopted, abort startup
        }
    })
);

```

### Stage 3: Journal Replay and Respawn

If the PID is dead or the socket is unreachable, the supervisor enters recovery mode. It instantiates a `WorkerRecoveryJournal` using the path stored in the descriptor and replays any pending commands to restore the worker to its last consistent state. The replay is bounded by `MAX_DEFERRED_RECOVERY_ROUNDS` (10 rounds) and `DEFERRED_RECOVERY_RECHECK_MS` (5 seconds) to prevent infinite loops.

If journal replay fails or the worker never responds, the supervisor executes `createOrReuseWorker()` (lines 9,210–9,260) to spawn a fresh process using the stored `createCommand`. This ensures the session directory layout remains unchanged even when the original process cannot be restored.

```typescript
private async adoptOrRecoverWorker(worker: ResidentWorker): Promise<void> {
    // 1. Check if the recorded PID is still alive
    if (isProcessAlive(worker.descriptor.pid)) {
        // 2. Try to open the worker’s socket
        await this.connectToWorker(worker);
        return;
    }

    // 3. Worker is dead – replay its recovery journal
    const journal = new WorkerRecoveryJournal(worker.descriptor.recoveryJournalPath);
    await journal.replayPendingCommands();

    // 4. If replay fails, start a fresh worker
    await this.createOrReuseWorker(
        worker.descriptor.ownerClientId ?? "recovered",
        worker.descriptor.createCommand
    );
}

```

## Recovery Constants and Timing Controls

The supervisor enforces strict timeouts and retry limits to prevent recovery operations from hanging indefinitely:

- **`DEFERRED_RECOVERY_RECHECK_MS = 5000`** – The supervisor waits 5 seconds between each round of probing a dead worker.
- **`MAX_DEFERRED_RECOVERY_ROUNDS = 10`** – Recovery attempts cease after approximately 2.5 minutes, triggering a respawn.
- **`WORKER_RETRY_DELAYS_MS = [250, 1000, 5000]`** – Back-off delays applied when reconnecting to temporarily unresponsive live workers.
- **`WORKER_CONNECT_TIMEOUT_MS = 30000`** – Hard timeout for any single worker connection attempt.

These constants ensure that the daemon supervisor recover workers after a crash efficiently without waiting indefinitely on zombie processes.

## Key Source Files

The recovery mechanism spans several modules across the codebase:

- **[`packages/coding-agent/src/modes/daemon/daemon-supervisor.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-supervisor.ts)** – Contains `loadWorkerDescriptors()`, `adoptOrRecoverWorker()`, and `createOrReuseWorker()` implementing the core recovery logic.
- **[`packages/coding-agent/src/modes/daemon/daemon-supervisor-ownership.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-supervisor-ownership.ts)** – Manages the exclusive ownership lease for the supervisor socket, validated before recovery begins.
- **[`packages/coding-agent/src/modes/daemon/daemon-worker-client.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-worker-client.ts)** – Provides the client wrapper used to communicate with live workers during the adoption phase.
- **[`packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts)** – Persists processed commands to disk, enabling the supervisor to replay state during recovery.
- **[`packages/coding-agent/src/utils/child-process.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/utils/child-process.ts)** – Implements `isProcessAlive()` and other utilities used to verify worker process status.

## Summary

- The supervisor reconstructs worker state from JSON descriptors stored in `descriptorDir` on startup, eliminating dependency on in-memory data.
- Process liveness is verified using `isProcessAlive()` before attempting to adopt existing workers via `DaemonWorkerClient`.
- Dead workers are restored by replaying their `WorkerRecoveryJournal`, which re-executes uncommitted commands to reach a consistent state.
- If recovery fails after `MAX_DEFERRED_RECOVERY_ROUNDS` attempts, the supervisor spawns a fresh worker using the original `createCommand`, preserving session continuity.
- All recovery logic is centralized in [`daemon-supervisor.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-supervisor.ts) with strict timing constants to prevent infinite recovery loops.

## Frequently Asked Questions

### How does the supervisor determine if a worker process survived the crash?

The supervisor calls `isProcessAlive()` from [`utils/child-process.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/utils/child-process.ts) to check if the PID stored in the worker descriptor is still running. If the PID exists and the worker’s socket is reachable, the supervisor adopts the process; otherwise, it initiates journal-based recovery or respawns the worker.

### What happens if the worker recovery journal is corrupted?

If `WorkerRecoveryJournal.replayPendingCommands()` fails or the journal is unreadable, the supervisor falls back to spawning a new worker via `createOrReuseWorker()`. The new process uses the same `createCommand` and session directory as the original, ensuring filesystem consistency even when state replay is impossible.

### How long does the supervisor attempt to recover a dead worker before giving up?

The supervisor attempts recovery for up to `MAX_DEFERRED_RECOVERY_ROUNDS` (10 rounds), waiting `DEFERRED_RECOVERY_RECHECK_MS` (5 seconds) between each attempt. After approximately 2.5 minutes, it abandons recovery and respawns the worker to restore availability.

### Where is worker state stored to survive supervisor restarts?

Worker state is persisted as JSON files in the `descriptorDir` directory. Each descriptor contains the worker PID, socket path, launch command, and recovery journal path. This persistent storage strategy ensures that the daemon supervisor recover workers after a crash even when the original process is terminated.