# How the Prime Agent Daemon Protocol Handles Reconnection, Snapshots, and Crash Recovery

> Explore how the Prime Agent daemon protocol ensures seamless reconnection and crash recovery with JSON-L, chunked snapshots, and resume cursors. Learn more today.

- Repository: [Prime Intellect/prime-agent](https://github.com/PrimeIntellect-ai/prime-agent)
- Tags: internals
- Published: 2026-08-18

---

**The Prime Agent daemon protocol uses a JSON-L over local socket design with chunked snapshot streaming and resume cursors to ensure seamless reconnection and crash recovery without losing session state.**

The PrimeIntellect-ai/prime-agent repository implements a robust daemon protocol that separates commands from events to maintain session continuity across network interruptions and process restarts. This protocol leverages snapshot-based session recovery and transient caching to handle large transcripts efficiently while preserving state integrity during supervisor-managed restarts.

## JSON-L Protocol Architecture and Snapshot Lifecycle

The daemon operates over a **local socket** using **JSON-L (JSON Lines)** format with protocol version 7. It separates command submission from event streaming and advertises capabilities including `attach_snapshot` and `chunked_snapshot`.

When a client attaches or the daemon needs to synchronize state, it emits a sequence of snapshot events:

- **session_snapshot_begin**: Signals the start of a new snapshot stream with fields including `activeSessionId`, `snapshotId`, `messageCount`, and `purpose` (values: `"attach"`, `"replacement"`, or `"resync"`). This event is defined in [`packages/coding-agent/src/modes/daemon/daemon-protocol.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-protocol.ts).
- **session_snapshot_chunk**: Transmits transcript segments limited to `targetChunkBytes` (default approximately 512 KiB), containing `index` and `messages` arrays.
- **session_snapshot_end**: Marks completion with `chunkCount` and `lastEventSequence`.
- **session_snapshot_failed**: Indicates terminal errors with an `error` payload, rendering the snapshot unusable.

The `snapshotId` uniquely identifies a generation and regenerates whenever the daemon restarts, ensuring clients can distinguish between stale and current state.

## Reconnection Flow with Resume Cursors

When clients disconnect, the protocol supports precise state resumption through cursor-based replay rather than full state retransmission.

### The Attach Command Structure

Clients initiate recovery by sending an `attach` command containing an optional `resumeCursor`. The following example from [`daemon-protocol.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-protocol.ts) demonstrates the command structure:

```typescript
import { createDaemonCommandEnvelope } from "../daemon-protocol";

const attachCmd = {
  type: "attach",
  activeSessionId: "session-42",
  resumeCursor: { generation: "gen-1", sequence: 12345 },
};

const envelope = createDaemonCommandEnvelope(
  attachCmd,
  "cmd-001",
  "client-a",
);
socket.write(JSON.stringify(envelope) + "\n");

```

If the cursor's `generation` matches the daemon's current `supervisorGeneration`, the daemon replays events from the specified `sequence`. Otherwise, it initiates a fresh snapshot stream via the three-step chunked sequence.

### Chunked Delivery and Streaming

For large sessions, the daemon streams snapshots in chunks rather than loading entire transcripts into memory. If a reconnection occurs mid-stream, the daemon preserves partially-sent snapshots through the `SnapshotTranscriptCache`, allowing new clients to resume reading from the exact abort point without re-encoding the transcript.

## Crash Recovery Mechanisms

When the daemon process crashes and restarts, the supervisor coordinates state preservation without client data loss through generation management and cache continuity.

### Supervisor Generation Management

Upon restart, the supervisor increments `supervisorGeneration` and issues a new `daemon_hello` event. The old `SnapshotTranscriptCache` remains alive via reference counting (`retain()`/`dispose()`) until all readers release it, ensuring in-flight chunks remain readable even as a new worker initializes.

The protocol provides **DaemonReplayInfo** indicating replay availability: `"complete"`, `"partial"`, or `"unavailable"` based on cursor validity against the current `lastEventSequence`.

### Replacement Snapshots and Cache Continuity

The new worker may emit a **replacement snapshot** with `purpose: "replacement"` and a fresh `snapshotId`. Clients receive this while potentially finishing consumption of the previous generation's chunks—a behavior validated in test `ENG-4677 snapshot catch-up replacement` (see [`packages/coding-agent/test/suite/regressions/4677-snapshot-catchup-replacement.test.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/test/suite/regressions/4677-snapshot-catchup-replacement.test.ts)). The following handler pattern demonstrates client-side processing:

```typescript
socket.on("data", (line) => {
  const event = JSON.parse(line);
  switch (event.type) {
    case "session_snapshot_begin":
      if (event.purpose === "replacement") {
        // Start a new cache for the replacement snapshot
        currentCache = new SnapshotTranscriptCache({
          activeSessionId: event.activeSessionId,
          snapshotId: event.snapshotId,
          cacheRoot: "/tmp/daemon-snapshot",
        });
      }
      break;
    case "session_snapshot_chunk":
      currentCache.appendEncodedChunk(Buffer.from(line));
      break;
    case "session_snapshot_end":
      currentCache.markComplete();
      break;
  }
});

```

Once all readers complete the old snapshot (`session_snapshot_end` or `session_snapshot_failed`), the cache disposes safely and clients transition seamlessly to the new generation.

## SnapshotTranscriptCache Implementation Details

The `SnapshotTranscriptCache` class in [`packages/coding-agent/src/modes/daemon/snapshot-transcript-cache.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/snapshot-transcript-cache.ts) manages the storage layer for snapshot chunks with memory safety and concurrent access support.

### Memory Buffering and Disk Spilling

The cache maintains chunks in memory as `Buffer` objects until the total exceeds `memoryCacheBytes` (default 4 MiB). Upon exceeding this limit, it creates a temporary directory at `cacheRoot/<snapshotId>`, flushes existing buffers to files, and writes subsequent chunks directly to disk.

### Concurrent Reader Coordination

Multiple clients can share a single snapshot stream through `retain()` and `dispose()` methods that implement reference counting. The `waitForChunk(index)` method returns a `Promise` resolving when specific chunks become available, enabling efficient streaming without blocking. If errors occur, `markFailed(error)` propagates exceptions to all pending waiters.

The following implementation demonstrates consuming a cached snapshot stream:

```typescript
async function readSnapshot(clientId: string) {
  const cache = new SnapshotTranscriptCache({
    activeSessionId: "session-42",
    snapshotId: "snapshot-1",
    cacheRoot: "/tmp/daemon-snapshot",
  });

  // Wait for each chunk as they arrive
  let index = 0;
  while (true) {
    const chunk = await cache.waitForChunk(index);
    if (!chunk) break; // end of snapshot
    // Process messages in the chunk
    const events = JSON.parse(chunk.toString()).messages;
    handleMessages(events);
    index++;
  }
}

```

## Summary

- The Prime Agent daemon protocol uses JSON-L over local sockets with explicit snapshot events (`session_snapshot_begin`, `chunk`, `end`, `failed`) for state synchronization, defined in [`daemon-protocol.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-protocol.ts).
- Reconnection relies on `resumeCursor` values in the `attach` command to enable event replay or trigger fresh snapshots based on `supervisorGeneration` matching.
- Crash recovery preserves in-flight snapshot data through the `SnapshotTranscriptCache` reference counting system, allowing replacement snapshots to coexist with ongoing stream completion.
- The `SnapshotTranscriptCache` automatically spills to disk when memory limits (4 MiB default) are reached and supports concurrent readers without re-encoding transcripts.
- Test `ENG-4677 snapshot catch-up replacement` validates that partial snapshots complete correctly after daemon restarts.

## Frequently Asked Questions

### What protocol version does Prime Agent use for daemon communication?

The daemon protocol operates at version 7 using JSON-L (JSON Lines) over a local socket, separating command submission from event streaming to handle asynchronous session updates efficiently.

### How does the daemon handle large session transcripts during reconnection?

When the client advertises `chunked_snapshot` capability, the daemon splits the transcript into chunks limited by `targetChunkBytes` (approximately 512 KiB default) and streams them via `session_snapshot_chunk` events. The `SnapshotTranscriptCache` spills to disk if the total size exceeds 4 MiB, preventing memory exhaustion while maintaining stream integrity.

### What happens to active snapshots when the daemon crashes?

The supervisor preserves the existing `SnapshotTranscriptCache` through reference counting, allowing connected clients to finish receiving queued chunks. The new worker may start a replacement snapshot with a new `snapshotId`, and clients transition seamlessly once the old cache fully disposes after all readers call their release callbacks.

### How does the resume cursor mechanism determine whether to replay events or send a snapshot?

The daemon compares the `generation` field in the client's `resumeCursor` against its current `supervisorGeneration`. If they match, it replays events from the specified `sequence` number. If generations differ (indicating a daemon restart), it sends a full snapshot instead to ensure state consistency.