Worker Crash Recovery and Session Restoration in Prime Agent: Process Lifecycle Explained
Prime Agent implements a deterministic crash recovery protocol that captures worker state in a JSON-L recovery journal, marks crashed sessions with a "crash" status, and automatically spawns replacement workers to restore active sessions without data loss.
The Prime Agent system from the PrimeIntellect-ai/prime-agent repository uses a daemon-worker architecture to manage long-running LLM sessions. When the worker process crashes, the system follows a rigorous lifecycle to detect the failure, preserve session data, and restore execution continuity. This article examines the complete process lifecycle, referencing the actual source implementation to demonstrate how crash recovery and session restoration maintain data integrity.
Architecture Overview: Daemon-Worker Model
Prime Agent operates through a daemon process that spawns long-living worker processes to execute LLM-driven sessions. The daemon handles I/O multiplexing, manages concurrent sessions, and persists critical state, while the worker performs the actual computation.
When launched via DaemonWorker.start() in packages/coding-agent/src/modes/daemon/daemon-mode.ts, the worker receives optional initialization parameters. If restoring a previous session, the daemon passes restoreActiveSessionId in the worker launch options, instructing the worker to load existing state from the session ledger rather than creating a fresh session.
The Recovery Journal: Capturing Worker State
The cornerstone of crash recovery is the Worker Recovery Journal, implemented in packages/coding-agent/src/modes/daemon/worker-recovery-journal.ts. This persistent JSON-L append-only log records every critical operation during the worker's lifecycle.
Journal Format and Tombstone Entries
The journal captures:
- Session creation events
- Snapshot generation timestamps
- Ledger modifications
- A final tombstone entry with
status: "crash"when the process exits abnormally
This incremental write pattern ensures that even if the worker terminates unexpectedly, the daemon can determine exactly where execution stopped by scanning for the last consistent entry.
Safe Ledger Persistence with RLM
Supporting the journal is the Resource Ledger Manager (RLM), which flushes pending entries to disk after each successful operation. In packages/coding-agent/src/modes/daemon/rlm-ledger.ts (lines 704-789), the safeAppend method ensures atomic writes, preventing torn writes that could corrupt session state during a power loss or sudden termination.
Crash Detection and Analysis
The daemon monitors worker health through two primary mechanisms defined in the test suite and daemon implementation.
Detecting Socket Errors and Process Exits
Crash detection triggers through:
- Socket error emissions - When the worker socket emits an error event, as tested in
packages/coding-agent/test/daemon-client.test.ts(line 710) - Process exit handling - The daemon's
workerExitHandlercatches non-zero exit statuses, indicating abnormal termination
Upon detection, the daemon immediately halts I/O operations for that worker and initiates the recovery sequence.
Analyzing the Recovery Journal
The daemon instantiates a WorkerRecoveryJournal for the crashed worker's log file. The analysis method scans entries to determine consistency:
- If the final entry is a crash tombstone, the session requires restoration
- If the journal ends cleanly, the worker completed successfully
This logic in worker-recovery-journal.ts (lines 45-78) differentiates between recoverable crashes and clean shutdowns, preventing unnecessary restoration overhead.
Session State Transitions
Once a crash is confirmed, the system updates session metadata to reflect the failure state.
Marking Sessions as Crashed
The daemon updates the Session Registry in packages/coding-agent/src/modes/daemon/daemon-session-list.ts, explicitly setting the session status field to "crash". This atomic state change ensures that the session is no longer considered active for new operations, preventing data corruption from concurrent access.
Registry Filtering and Session Isolation
The registry implementation filters out status: "crash" entries from the active session view (lines 470-492 in daemon-session-list.ts). This isolation prevents crashed sessions from appearing in health checks or user listings while maintaining the session data for restoration purposes.
Worker Replacement and Session Restoration
With the crashed session isolated and analyzed, the daemon proceeds to restore execution continuity.
Spawning Replacement Workers with restoreActiveSessionId
The daemon spawns a fresh worker through the logic in packages/coding-agent/src/modes/daemon/daemon-mode.ts (lines 1088-1105). The spawn configuration includes:
{
restoreActiveSessionId: "session-abc123", // ID from crashed session
authenticationToken: "worker-token"
}
This parameter instructs the new worker to bypass initialization and load existing state.
Resuming from Snapshots
The replacement worker performs the following restoration sequence:
- Loads the session ledger from disk using the provided ID
- Reconstructs conversation state from the last valid snapshot
- Replays journal entries up to the crash point (excluding partial operations)
- Resumes the LLM conversation loop
The WorkerRecoveryJournal ensures the new worker does not replay operations that were partially written before the crash, maintaining exactly-once semantics for session modifications.
Implementation Code Examples
Launching a Daemon with Session Restoration
import { Daemon } from "prime-agent";
import { join } from "path";
import { tmpdir } from "os";
const socketPath = join(tmpdir(), "daemon.sock");
// Initialize daemon with crash recovery enabled
const daemon = new Daemon(socketPath, {
worker: {
authenticationToken: "secure-token",
// Restore session from previous crash
restoreActiveSessionId: "session-abc123",
},
});
await daemon.start();
Source: packages/coding-agent/test/daemon-launch.test.ts (lines 49-57)
Handling Worker Crash Events
daemon.on("workerCrash", async (workerId: string) => {
// Create recovery journal analyzer for crashed worker
const journal = daemon.createRecoveryJournal(workerId);
// Analyze crash state
const crashInfo = await journal.analyze();
// Returns: { status: "crash", latestSnapshot: "snap-001", sessionId: "..." }
// Update registry to mark session as crashed
daemon.sessionList.updateState(workerId, { status: "crash" });
// Spawn replacement worker attached to same session
await daemon.spawnWorker({
restoreActiveSessionId: crashInfo.sessionId,
});
});
Source: packages/coding-agent/src/modes/daemon/daemon-mode.ts (lines 1088-1105)
Session State File Structure
The persisted session state uses this JSON structure:
{
"sessionId": "session-abc123",
"status": "crash",
"ledger": {
"entries": [
{"op": "snapshot", "id": "snap-001", "timestamp": 1699123456},
{"op": "modify", "path": "/workspace/file.ts", "checksum": "abc..."}
]
},
"snapshotId": "snap-001"
}
When restoreActiveSessionId is provided, the daemon reads this file from packages/coding-agent/src/modes/daemon/session-state.ts and rebuilds the execution context.
Summary
- Worker Recovery Journal in
worker-recovery-journal.tsprovides append-only crash telemetry using JSON-L format with tombstone entries for abnormal exits - Crash Detection occurs through socket error monitoring and process exit handlers in the daemon, triggering immediate isolation of the failed worker
- Session State Management updates the registry in
daemon-session-list.tsto mark sessions as"crash", filtering them from active listings while preserving data - Worker Replacement uses
restoreActiveSessionIdto attach new processes to existing sessions, ensuring continuity without user intervention - RLM Safe Appends in
rlm-ledger.tsguarantee ledger integrity through atomic writes, preventing data corruption during unexpected termination
Frequently Asked Questions
How does Prime Agent detect a worker crash?
The daemon monitors the worker process through socket error events and process exit codes. As implemented in packages/coding-agent/test/daemon-client.test.ts (line 710), socket errors emit immediately when the worker terminates unexpectedly, while the workerExitHandler captures non-zero exit statuses. These signals trigger the crash recovery protocol within milliseconds of failure.
What happens to active sessions when a worker crashes?
Active sessions transition to a status: "crash" state in the session registry (packages/coding-agent/src/modes/daemon/daemon-session-list.ts). The daemon removes these sessions from the active view to prevent new operations while preserving all session data, ledger entries, and snapshots on disk for subsequent restoration.
How does the recovery journal prevent data loss?
The Worker Recovery Journal writes incrementally to disk in JSON-L format, recording every operation before execution completes. If a crash occurs, the journal contains a tombstone entry marking the exact failure point. The replacement worker reads this journal to identify which operations completed successfully and which require rollback, ensuring exactly-once semantics for all session modifications.
Where is session state stored during a crash?
Session state persists in the Resource Ledger managed by packages/coding-agent/src/modes/daemon/rlm-ledger.ts, which flushes entries to disk after each operation (lines 704-789). Additionally, the session metadata file contains the full ledger history and latest snapshot ID. Both locations use atomic write operations to ensure consistency even if the worker terminates during I/O operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →