# Worker Crash Recovery and Session Restoration in Prime Agent: Process Lifecycle Explained

> Understand Prime Agent's worker crash recovery and session restoration. Learn how it uses a JSON-L journal to automatically restore active sessions without data loss.

- Repository: [Prime Intellect/prime-agent](https://github.com/PrimeIntellect-ai/prime-agent)
- Tags: internals
- Published: 2026-08-18

---

**Prime Agent implements a deterministic crash recovery protocol that captures worker state in a JSON-L recovery journal, marks crashed sessions with a "crash" status, and automatically spawns replacement workers to restore active sessions without data loss.**

The Prime Agent system from the PrimeIntellect-ai/prime-agent repository uses a daemon-worker architecture to manage long-running LLM sessions. When the worker process crashes, the system follows a rigorous lifecycle to detect the failure, preserve session data, and restore execution continuity. This article examines the complete process lifecycle, referencing the actual source implementation to demonstrate how crash recovery and session restoration maintain data integrity.

## Architecture Overview: Daemon-Worker Model

Prime Agent operates through a **daemon** process that spawns long-living **worker** processes to execute LLM-driven sessions. The daemon handles I/O multiplexing, manages concurrent sessions, and persists critical state, while the worker performs the actual computation.

When launched via `DaemonWorker.start()` in [`packages/coding-agent/src/modes/daemon/daemon-mode.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-mode.ts), the worker receives optional initialization parameters. If restoring a previous session, the daemon passes `restoreActiveSessionId` in the worker launch options, instructing the worker to load existing state from the session ledger rather than creating a fresh session.

## The Recovery Journal: Capturing Worker State

The cornerstone of crash recovery is the **Worker Recovery Journal**, implemented in [`packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts). This persistent JSON-L append-only log records every critical operation during the worker's lifecycle.

### Journal Format and Tombstone Entries

The journal captures:
- Session creation events
- Snapshot generation timestamps
- Ledger modifications
- A final **tombstone** entry with `status: "crash"` when the process exits abnormally

This incremental write pattern ensures that even if the worker terminates unexpectedly, the daemon can determine exactly where execution stopped by scanning for the last consistent entry.

### Safe Ledger Persistence with RLM

Supporting the journal is the **Resource Ledger Manager (RLM)**, which flushes pending entries to disk after each successful operation. In [`packages/coding-agent/src/modes/daemon/rlm-ledger.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/rlm-ledger.ts) (lines 704-789), the `safeAppend` method ensures atomic writes, preventing torn writes that could corrupt session state during a power loss or sudden termination.

## Crash Detection and Analysis

The daemon monitors worker health through two primary mechanisms defined in the test suite and daemon implementation.

### Detecting Socket Errors and Process Exits

Crash detection triggers through:
1. **Socket error emissions** - When the worker socket emits an error event, as tested in [`packages/coding-agent/test/daemon-client.test.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/test/daemon-client.test.ts) (line 710)
2. **Process exit handling** - The daemon's `workerExitHandler` catches non-zero exit statuses, indicating abnormal termination

Upon detection, the daemon immediately halts I/O operations for that worker and initiates the recovery sequence.

### Analyzing the Recovery Journal

The daemon instantiates a `WorkerRecoveryJournal` for the crashed worker's log file. The analysis method scans entries to determine consistency:

- If the final entry is a crash tombstone, the session requires restoration
- If the journal ends cleanly, the worker completed successfully

This logic in [`worker-recovery-journal.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/worker-recovery-journal.ts) (lines 45-78) differentiates between recoverable crashes and clean shutdowns, preventing unnecessary restoration overhead.

## Session State Transitions

Once a crash is confirmed, the system updates session metadata to reflect the failure state.

### Marking Sessions as Crashed

The daemon updates the **Session Registry** in [`packages/coding-agent/src/modes/daemon/daemon-session-list.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-session-list.ts), explicitly setting the session `status` field to `"crash"`. This atomic state change ensures that the session is no longer considered active for new operations, preventing data corruption from concurrent access.

### Registry Filtering and Session Isolation

The registry implementation filters out `status: "crash"` entries from the active session view (lines 470-492 in [`daemon-session-list.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-session-list.ts)). This isolation prevents crashed sessions from appearing in health checks or user listings while maintaining the session data for restoration purposes.

## Worker Replacement and Session Restoration

With the crashed session isolated and analyzed, the daemon proceeds to restore execution continuity.

### Spawning Replacement Workers with restoreActiveSessionId

The daemon spawns a fresh worker through the logic in [`packages/coding-agent/src/modes/daemon/daemon-mode.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-mode.ts) (lines 1088-1105). The spawn configuration includes:

```typescript
{
  restoreActiveSessionId: "session-abc123",  // ID from crashed session
  authenticationToken: "worker-token"
}

```

This parameter instructs the new worker to bypass initialization and load existing state.

### Resuming from Snapshots

The replacement worker performs the following restoration sequence:
1. Loads the session ledger from disk using the provided ID
2. Reconstructs conversation state from the last valid snapshot
3. Replays journal entries up to the crash point (excluding partial operations)
4. Resumes the LLM conversation loop

The `WorkerRecoveryJournal` ensures the new worker does not replay operations that were partially written before the crash, maintaining exactly-once semantics for session modifications.

## Implementation Code Examples

### Launching a Daemon with Session Restoration

```typescript
import { Daemon } from "prime-agent";
import { join } from "path";
import { tmpdir } from "os";

const socketPath = join(tmpdir(), "daemon.sock");

// Initialize daemon with crash recovery enabled
const daemon = new Daemon(socketPath, {
  worker: {
    authenticationToken: "secure-token",
    // Restore session from previous crash
    restoreActiveSessionId: "session-abc123",
  },
});

await daemon.start();

```

**Source**: [`packages/coding-agent/test/daemon-launch.test.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/test/daemon-launch.test.ts) (lines 49-57)

### Handling Worker Crash Events

```typescript
daemon.on("workerCrash", async (workerId: string) => {
  // Create recovery journal analyzer for crashed worker
  const journal = daemon.createRecoveryJournal(workerId);
  
  // Analyze crash state
  const crashInfo = await journal.analyze();
  // Returns: { status: "crash", latestSnapshot: "snap-001", sessionId: "..." }
  
  // Update registry to mark session as crashed
  daemon.sessionList.updateState(workerId, { status: "crash" });
  
  // Spawn replacement worker attached to same session
  await daemon.spawnWorker({
    restoreActiveSessionId: crashInfo.sessionId,
  });
});

```

**Source**: [`packages/coding-agent/src/modes/daemon/daemon-mode.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-mode.ts) (lines 1088-1105)

### Session State File Structure

The persisted session state uses this JSON structure:

```json
{
  "sessionId": "session-abc123",
  "status": "crash",
  "ledger": {
    "entries": [
      {"op": "snapshot", "id": "snap-001", "timestamp": 1699123456},
      {"op": "modify", "path": "/workspace/file.ts", "checksum": "abc..."}
    ]
  },
  "snapshotId": "snap-001"
}

```

When `restoreActiveSessionId` is provided, the daemon reads this file from [`packages/coding-agent/src/modes/daemon/session-state.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/session-state.ts) and rebuilds the execution context.

## Summary

- **Worker Recovery Journal** in [`worker-recovery-journal.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/worker-recovery-journal.ts) provides append-only crash telemetry using JSON-L format with tombstone entries for abnormal exits
- **Crash Detection** occurs through socket error monitoring and process exit handlers in the daemon, triggering immediate isolation of the failed worker
- **Session State Management** updates the registry in [`daemon-session-list.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-session-list.ts) to mark sessions as `"crash"`, filtering them from active listings while preserving data
- **Worker Replacement** uses `restoreActiveSessionId` to attach new processes to existing sessions, ensuring continuity without user intervention
- **RLM Safe Appends** in [`rlm-ledger.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/rlm-ledger.ts) guarantee ledger integrity through atomic writes, preventing data corruption during unexpected termination

## Frequently Asked Questions

### How does Prime Agent detect a worker crash?

The daemon monitors the worker process through socket error events and process exit codes. As implemented in [`packages/coding-agent/test/daemon-client.test.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/test/daemon-client.test.ts) (line 710), socket errors emit immediately when the worker terminates unexpectedly, while the `workerExitHandler` captures non-zero exit statuses. These signals trigger the crash recovery protocol within milliseconds of failure.

### What happens to active sessions when a worker crashes?

Active sessions transition to a `status: "crash"` state in the session registry ([`packages/coding-agent/src/modes/daemon/daemon-session-list.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-session-list.ts)). The daemon removes these sessions from the active view to prevent new operations while preserving all session data, ledger entries, and snapshots on disk for subsequent restoration.

### How does the recovery journal prevent data loss?

The **Worker Recovery Journal** writes incrementally to disk in JSON-L format, recording every operation before execution completes. If a crash occurs, the journal contains a tombstone entry marking the exact failure point. The replacement worker reads this journal to identify which operations completed successfully and which require rollback, ensuring exactly-once semantics for all session modifications.

### Where is session state stored during a crash?

Session state persists in the **Resource Ledger** managed by [`packages/coding-agent/src/modes/daemon/rlm-ledger.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/rlm-ledger.ts), which flushes entries to disk after each operation (lines 704-789). Additionally, the session metadata file contains the full ledger history and latest snapshot ID. Both locations use atomic write operations to ensure consistency even if the worker terminates during I/O operations.