# Retry Strategy for Worker Recovery in Prime Agent: Bounded Retries and Persistent Journaling

> Learn about Prime Agent's retry strategy for worker recovery. Discover how bounded retries and persistent journaling ensure reliable restarts without infinite loops.

- Repository: [Prime Intellect/prime-agent](https://github.com/PrimeIntellect-ai/prime-agent)
- Tags: deep-dive
- Published: 2026-09-05

---

**Prime Agent implements a deterministic retry strategy for worker recovery that combines a persistent `WorkerRecoveryJournal` with configurable retry limits and fixed back-off delays to ensure crashed workers restart reliably without entering infinite loops.**

The PrimeIntellect-ai/prime-agent repository uses a daemon architecture designed to maintain long-running coding sessions across worker processes. When a worker crashes or becomes unresponsive, the system follows a strict recovery protocol defined in [`packages/coding-agent/src/modes/daemon/daemon-supervisor.js`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-supervisor.js) and [`packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts). This article explains how the retry strategy works, what configuration options are available, and how persistent state tracking prevents data loss during recovery.

## How WorkerRecoveryJournal Persists Recovery State

The foundation of the retry strategy is the **`WorkerRecoveryJournal`** class, which maintains an append-only log of worker states. Before attempting any restart operation, the `DaemonSupervisor` creates a journal instance at a specified recovery path and writes an atomic record.

Each **`WorkerRecoveryRecord`** contains four critical fields:

- **`activeSessionId`** – Identifies the parent session owning the worker
- **`busy`** – Boolean flag indicating whether the worker is processing a request
- **`operation`** – String describing the current step (e.g., `"launch"` or `"attach"`)
- **`recordedAt`** – ISO 8601 timestamp for ordering and debugging

The journal retains only the most recent record per `activeSessionId`. When the supervisor detects that all recorded workers have `busy: false`, it triggers **`compact()`**, rewriting the file to a minimal state to prevent unbounded growth.

```typescript
import { WorkerRecoveryJournal } from "./worker-recovery-journal.js";

const recoveryJournalPath = "/var/run/prime-agent/worker.recovery.jsonl";
const journal = new WorkerRecoveryJournal(recoveryJournalPath);

// Record a launch attempt before starting the process
journal.record({
  activeSessionId: "sess-42",
  sessionId: "w-42",
  sessionFile: "/tmp/session-42.json",
  busy: true,
  operation: "launch",
});

```

## Bounded Retry Logic in DaemonSupervisor

The **`DaemonSupervisor`** orchestrates the actual retry loop. When a worker fails to launch or crashes, the supervisor consults the journal and enters a deterministic retry cycle with two configurable parameters:

1. **`maxRetries`** – The maximum number of restart attempts (default: **5**)
2. **`retryDelay`** – The fixed back-off delay between attempts in milliseconds (default: **5000**, or 5 seconds)

The supervisor instantiates a `WorkerRecoveryJournal` for each managed worker, then wraps the launch logic in a `while` loop that increments an internal counter until `maxRetries` is exhausted. If the counter exceeds the limit, the worker is marked as permanently failed and the error propagates upward.

```typescript
async function launchWorker(workerSpec) {
  const maxRetries = workerSpec.settings.maxRetries ?? 5;
  const retryDelay = workerSpec.settings.retryDelay ?? 5000;
  let attempt = 0;

  while (attempt <= maxRetries) {
    try {
      journal.record({
        activeSessionId: workerSpec.sessionId,
        sessionId: workerSpec.id,
        busy: true,
        operation: "launch",
      });
      await startWorkerProcess(workerSpec);
      
      // Success: mark as not busy and exit loop
      journal.record({
        activeSessionId: workerSpec.sessionId,
        sessionId: workerSpec.id,
        busy: false,
        operation: "launch",
      });
      return;
    } catch (e) {
      attempt++;
      if (attempt > maxRetries) throw e;
      await new Promise(r => setTimeout(r, retryDelay));
    }
  }
}

```

## Record-Driven Recovery Workflow

The retry strategy follows a strict state machine driven by journal records:

1. **Pre-launch** – Supervisor writes a `busy: true` record with operation `"launch"`
2. **Attempt** – Supervisor executes `startWorkerProcess()`
3. **Success path** – On clean start, writes `busy: false` and calls `journal.compact()` if no other workers are busy
4. **Failure path** – On exception, increments retry counter, waits `retryDelay`, and returns to step 1
5. **Permanent failure** – When `maxRetries` is exceeded, the supervisor stops attempting recovery and retains the last record for audit purposes

This workflow ensures that even if the daemon itself crashes during a retry attempt, the next daemon startup can read the journal and resume recovery exactly where it left off.

## Graceful Shutdown and Journal Compaction

When the daemon receives a shutdown signal, it immediately calls **`compact()`** on the `WorkerRecoveryJournal`. This guarantees that a subsequent daemon start sees a clean state, avoiding unnecessary retries of workers that were intentionally stopped. The compaction logic filters out all historical records, keeping only the current state snapshot.

## Summary

- **Prime Agent's retry strategy for worker recovery** uses a persistent `WorkerRecoveryJournal` to maintain state across process restarts
- **Default configuration** allows **5 retry attempts** with a **5-second fixed back-off delay**, both configurable via `maxRetries` and `retryDelay` settings
- **Atomic records** track session ID, busy status, operation type, and timestamp for every recovery step
- **Automatic compaction** occurs when all workers report `busy: false` or when the daemon shuts down gracefully
- **Source implementation** resides in [`packages/coding-agent/src/modes/daemon/daemon-supervisor.js`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-supervisor.js) and [`packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts)

## Frequently Asked Questions

### What is the default retry limit for worker recovery in Prime Agent?

The default `maxRetries` value is **5**. After five failed attempts to restart a worker, the `DaemonSupervisor` marks the worker as permanently failed and stops the retry loop. This limit can be overridden in the daemon configuration settings.

### How does Prime Agent prevent infinite retry loops when a worker keeps crashing?

The system enforces a **bounded retry policy** using the `maxRetries` parameter. Unlike exponential back-off strategies that can theoretically run forever, Prime Agent uses a fixed retry ceiling. Once the ceiling is reached, the supervisor explicitly throws the last error and ceases recovery attempts for that specific worker.

### What information does the WorkerRecoveryJournal store during recovery attempts?

According to the source code in [`packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts), each record stores the `activeSessionId`, a `busy` boolean flag, the `operation` being performed (such as `"launch"` or `"attach"`), and a `recordedAt` timestamp. This data allows the supervisor to resume recovery exactly where it left off after a daemon restart.

### Where can I find the integration tests for worker recovery?

The retry behavior is verified in two key test files: [`packages/coding-agent/test/worker-recovery-journal.test.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/test/worker-recovery-journal.test.ts) contains unit tests for the journal's record handling and compaction logic, while [`packages/coding-agent/test/suite/4603-worker-recovery.test.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/test/suite/4603-worker-recovery.test.ts) provides integration coverage for the supervisor's five-second retry behavior and end-to-end recovery flows.