How the Daemon Supervisor Recovers Workers After a Crash in Prime Agent
The daemon supervisor implements a three-stage recovery pipeline—loading persisted JSON descriptors, attempting to adopt live worker processes, and replaying recovery journals or respawning failed workers—to guarantee session continuity after unexpected crashes.
When the PrimeIntellect-ai/prime-agent daemon supervisor restarts after a crash, it must reconstruct the state of all previously running session workers without losing active sessions. According to the source code in packages/coding-agent/src/modes/daemon/daemon-supervisor.ts, the supervisor achieves this through a deterministic "adopt-or-recover-or-respawn" strategy that relies on disk-persisted worker descriptors and command journals rather than volatile in-memory state.
The Three-Stage Worker Recovery Pipeline
The recovery process executes sequentially during supervisor startup. Each stage is designed to maximize the chance of restoring the exact pre-crash state while providing a clean fallback to fresh workers when necessary.
Stage 1: Loading Persisted Worker Descriptors
On initialization, the supervisor reads every *.json file from its descriptorDir directory. These files contain the last known state of each worker, including the PID, socket path, createCommand that launched the process, and paths to recovery journals. The loadWorkerDescriptors() method (lines 13,000–13,040) deserializes these files into ResidentWorker objects and stores them in this.workers, ensuring the supervisor never depends on ephemeral memory to recover worker identity.
private loadWorkerDescriptors(): void {
for (const name of readdirSync(this.descriptorDir)) {
if (name === SUPERVISOR_CONFIG_FILE_NAME || !name.endsWith(".json")) continue;
const path = join(this.descriptorDir, name);
const descriptor: unknown = JSON.parse(readFileSync(path, "utf8"));
if (!isDaemonWorkerDescriptor(descriptor, this.socketPath)) continue;
// …populate ResidentWorker and store in this.workers
}
}
Stage 2: Adoption of Live Workers
For each loaded descriptor, the supervisor calls adoptOrRecoverWorker() (lines 9,000–9,150) to determine if the original process survived the crash. The helper first invokes isProcessAlive() from utils/child-process.ts to verify the recorded PID. If the process is alive and the socket can be opened, the supervisor attaches to the existing worker using DaemonWorkerClient, repopulating the worker’s roster and snapshot cache from the descriptor data.
await Promise.all(
workersToAdopt.map(async (worker) => {
try {
await this.adoptOrRecoverWorker(worker); // core recovery logic
} catch (error) {
// If any worker cannot be adopted, abort startup
}
})
);
Stage 3: Journal Replay and Respawn
If the PID is dead or the socket is unreachable, the supervisor enters recovery mode. It instantiates a WorkerRecoveryJournal using the path stored in the descriptor and replays any pending commands to restore the worker to its last consistent state. The replay is bounded by MAX_DEFERRED_RECOVERY_ROUNDS (10 rounds) and DEFERRED_RECOVERY_RECHECK_MS (5 seconds) to prevent infinite loops.
If journal replay fails or the worker never responds, the supervisor executes createOrReuseWorker() (lines 9,210–9,260) to spawn a fresh process using the stored createCommand. This ensures the session directory layout remains unchanged even when the original process cannot be restored.
private async adoptOrRecoverWorker(worker: ResidentWorker): Promise<void> {
// 1. Check if the recorded PID is still alive
if (isProcessAlive(worker.descriptor.pid)) {
// 2. Try to open the worker’s socket
await this.connectToWorker(worker);
return;
}
// 3. Worker is dead – replay its recovery journal
const journal = new WorkerRecoveryJournal(worker.descriptor.recoveryJournalPath);
await journal.replayPendingCommands();
// 4. If replay fails, start a fresh worker
await this.createOrReuseWorker(
worker.descriptor.ownerClientId ?? "recovered",
worker.descriptor.createCommand
);
}
Recovery Constants and Timing Controls
The supervisor enforces strict timeouts and retry limits to prevent recovery operations from hanging indefinitely:
DEFERRED_RECOVERY_RECHECK_MS = 5000– The supervisor waits 5 seconds between each round of probing a dead worker.MAX_DEFERRED_RECOVERY_ROUNDS = 10– Recovery attempts cease after approximately 2.5 minutes, triggering a respawn.WORKER_RETRY_DELAYS_MS = [250, 1000, 5000]– Back-off delays applied when reconnecting to temporarily unresponsive live workers.WORKER_CONNECT_TIMEOUT_MS = 30000– Hard timeout for any single worker connection attempt.
These constants ensure that the daemon supervisor recover workers after a crash efficiently without waiting indefinitely on zombie processes.
Key Source Files
The recovery mechanism spans several modules across the codebase:
packages/coding-agent/src/modes/daemon/daemon-supervisor.ts– ContainsloadWorkerDescriptors(),adoptOrRecoverWorker(), andcreateOrReuseWorker()implementing the core recovery logic.packages/coding-agent/src/modes/daemon/daemon-supervisor-ownership.ts– Manages the exclusive ownership lease for the supervisor socket, validated before recovery begins.packages/coding-agent/src/modes/daemon/daemon-worker-client.ts– Provides the client wrapper used to communicate with live workers during the adoption phase.packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts– Persists processed commands to disk, enabling the supervisor to replay state during recovery.packages/coding-agent/src/utils/child-process.ts– ImplementsisProcessAlive()and other utilities used to verify worker process status.
Summary
- The supervisor reconstructs worker state from JSON descriptors stored in
descriptorDiron startup, eliminating dependency on in-memory data. - Process liveness is verified using
isProcessAlive()before attempting to adopt existing workers viaDaemonWorkerClient. - Dead workers are restored by replaying their
WorkerRecoveryJournal, which re-executes uncommitted commands to reach a consistent state. - If recovery fails after
MAX_DEFERRED_RECOVERY_ROUNDSattempts, the supervisor spawns a fresh worker using the originalcreateCommand, preserving session continuity. - All recovery logic is centralized in
daemon-supervisor.tswith strict timing constants to prevent infinite recovery loops.
Frequently Asked Questions
How does the supervisor determine if a worker process survived the crash?
The supervisor calls isProcessAlive() from utils/child-process.ts to check if the PID stored in the worker descriptor is still running. If the PID exists and the worker’s socket is reachable, the supervisor adopts the process; otherwise, it initiates journal-based recovery or respawns the worker.
What happens if the worker recovery journal is corrupted?
If WorkerRecoveryJournal.replayPendingCommands() fails or the journal is unreadable, the supervisor falls back to spawning a new worker via createOrReuseWorker(). The new process uses the same createCommand and session directory as the original, ensuring filesystem consistency even when state replay is impossible.
How long does the supervisor attempt to recover a dead worker before giving up?
The supervisor attempts recovery for up to MAX_DEFERRED_RECOVERY_ROUNDS (10 rounds), waiting DEFERRED_RECOVERY_RECHECK_MS (5 seconds) between each attempt. After approximately 2.5 minutes, it abandons recovery and respawns the worker to restore availability.
Where is worker state stored to survive supervisor restarts?
Worker state is persisted as JSON files in the descriptorDir directory. Each descriptor contains the worker PID, socket path, launch command, and recovery journal path. This persistent storage strategy ensures that the daemon supervisor recover workers after a crash even when the original process is terminated.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →