Retry Strategy for Worker Recovery in Prime Agent: Bounded Retries and Persistent Journaling
Prime Agent implements a deterministic retry strategy for worker recovery that combines a persistent WorkerRecoveryJournal with configurable retry limits and fixed back-off delays to ensure crashed workers restart reliably without entering infinite loops.
The PrimeIntellect-ai/prime-agent repository uses a daemon architecture designed to maintain long-running coding sessions across worker processes. When a worker crashes or becomes unresponsive, the system follows a strict recovery protocol defined in packages/coding-agent/src/modes/daemon/daemon-supervisor.js and packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts. This article explains how the retry strategy works, what configuration options are available, and how persistent state tracking prevents data loss during recovery.
How WorkerRecoveryJournal Persists Recovery State
The foundation of the retry strategy is the WorkerRecoveryJournal class, which maintains an append-only log of worker states. Before attempting any restart operation, the DaemonSupervisor creates a journal instance at a specified recovery path and writes an atomic record.
Each WorkerRecoveryRecord contains four critical fields:
activeSessionId– Identifies the parent session owning the workerbusy– Boolean flag indicating whether the worker is processing a requestoperation– String describing the current step (e.g.,"launch"or"attach")recordedAt– ISO 8601 timestamp for ordering and debugging
The journal retains only the most recent record per activeSessionId. When the supervisor detects that all recorded workers have busy: false, it triggers compact(), rewriting the file to a minimal state to prevent unbounded growth.
import { WorkerRecoveryJournal } from "./worker-recovery-journal.js";
const recoveryJournalPath = "/var/run/prime-agent/worker.recovery.jsonl";
const journal = new WorkerRecoveryJournal(recoveryJournalPath);
// Record a launch attempt before starting the process
journal.record({
activeSessionId: "sess-42",
sessionId: "w-42",
sessionFile: "/tmp/session-42.json",
busy: true,
operation: "launch",
});
Bounded Retry Logic in DaemonSupervisor
The DaemonSupervisor orchestrates the actual retry loop. When a worker fails to launch or crashes, the supervisor consults the journal and enters a deterministic retry cycle with two configurable parameters:
maxRetries– The maximum number of restart attempts (default: 5)retryDelay– The fixed back-off delay between attempts in milliseconds (default: 5000, or 5 seconds)
The supervisor instantiates a WorkerRecoveryJournal for each managed worker, then wraps the launch logic in a while loop that increments an internal counter until maxRetries is exhausted. If the counter exceeds the limit, the worker is marked as permanently failed and the error propagates upward.
async function launchWorker(workerSpec) {
const maxRetries = workerSpec.settings.maxRetries ?? 5;
const retryDelay = workerSpec.settings.retryDelay ?? 5000;
let attempt = 0;
while (attempt <= maxRetries) {
try {
journal.record({
activeSessionId: workerSpec.sessionId,
sessionId: workerSpec.id,
busy: true,
operation: "launch",
});
await startWorkerProcess(workerSpec);
// Success: mark as not busy and exit loop
journal.record({
activeSessionId: workerSpec.sessionId,
sessionId: workerSpec.id,
busy: false,
operation: "launch",
});
return;
} catch (e) {
attempt++;
if (attempt > maxRetries) throw e;
await new Promise(r => setTimeout(r, retryDelay));
}
}
}
Record-Driven Recovery Workflow
The retry strategy follows a strict state machine driven by journal records:
- Pre-launch – Supervisor writes a
busy: truerecord with operation"launch" - Attempt – Supervisor executes
startWorkerProcess() - Success path – On clean start, writes
busy: falseand callsjournal.compact()if no other workers are busy - Failure path – On exception, increments retry counter, waits
retryDelay, and returns to step 1 - Permanent failure – When
maxRetriesis exceeded, the supervisor stops attempting recovery and retains the last record for audit purposes
This workflow ensures that even if the daemon itself crashes during a retry attempt, the next daemon startup can read the journal and resume recovery exactly where it left off.
Graceful Shutdown and Journal Compaction
When the daemon receives a shutdown signal, it immediately calls compact() on the WorkerRecoveryJournal. This guarantees that a subsequent daemon start sees a clean state, avoiding unnecessary retries of workers that were intentionally stopped. The compaction logic filters out all historical records, keeping only the current state snapshot.
Summary
- Prime Agent's retry strategy for worker recovery uses a persistent
WorkerRecoveryJournalto maintain state across process restarts - Default configuration allows 5 retry attempts with a 5-second fixed back-off delay, both configurable via
maxRetriesandretryDelaysettings - Atomic records track session ID, busy status, operation type, and timestamp for every recovery step
- Automatic compaction occurs when all workers report
busy: falseor when the daemon shuts down gracefully - Source implementation resides in
packages/coding-agent/src/modes/daemon/daemon-supervisor.jsandpackages/coding-agent/src/modes/daemon/worker-recovery-journal.ts
Frequently Asked Questions
What is the default retry limit for worker recovery in Prime Agent?
The default maxRetries value is 5. After five failed attempts to restart a worker, the DaemonSupervisor marks the worker as permanently failed and stops the retry loop. This limit can be overridden in the daemon configuration settings.
How does Prime Agent prevent infinite retry loops when a worker keeps crashing?
The system enforces a bounded retry policy using the maxRetries parameter. Unlike exponential back-off strategies that can theoretically run forever, Prime Agent uses a fixed retry ceiling. Once the ceiling is reached, the supervisor explicitly throws the last error and ceases recovery attempts for that specific worker.
What information does the WorkerRecoveryJournal store during recovery attempts?
According to the source code in packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts, each record stores the activeSessionId, a busy boolean flag, the operation being performed (such as "launch" or "attach"), and a recordedAt timestamp. This data allows the supervisor to resume recovery exactly where it left off after a daemon restart.
Where can I find the integration tests for worker recovery?
The retry behavior is verified in two key test files: packages/coding-agent/test/worker-recovery-journal.test.ts contains unit tests for the journal's record handling and compaction logic, while packages/coding-agent/test/suite/4603-worker-recovery.test.ts provides integration coverage for the supervisor's five-second retry behavior and end-to-end recovery flows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →