How the Prime Agent Daemon Protocol Handles Reconnection, Snapshots, and Crash Recovery

The Prime Agent daemon protocol uses a JSON-L over local socket design with chunked snapshot streaming and resume cursors to ensure seamless reconnection and crash recovery without losing session state.

The PrimeIntellect-ai/prime-agent repository implements a robust daemon protocol that separates commands from events to maintain session continuity across network interruptions and process restarts. This protocol leverages snapshot-based session recovery and transient caching to handle large transcripts efficiently while preserving state integrity during supervisor-managed restarts.

JSON-L Protocol Architecture and Snapshot Lifecycle

The daemon operates over a local socket using JSON-L (JSON Lines) format with protocol version 7. It separates command submission from event streaming and advertises capabilities including attach_snapshot and chunked_snapshot.

When a client attaches or the daemon needs to synchronize state, it emits a sequence of snapshot events:

  • session_snapshot_begin: Signals the start of a new snapshot stream with fields including activeSessionId, snapshotId, messageCount, and purpose (values: "attach", "replacement", or "resync"). This event is defined in packages/coding-agent/src/modes/daemon/daemon-protocol.ts.
  • session_snapshot_chunk: Transmits transcript segments limited to targetChunkBytes (default approximately 512 KiB), containing index and messages arrays.
  • session_snapshot_end: Marks completion with chunkCount and lastEventSequence.
  • session_snapshot_failed: Indicates terminal errors with an error payload, rendering the snapshot unusable.

The snapshotId uniquely identifies a generation and regenerates whenever the daemon restarts, ensuring clients can distinguish between stale and current state.

Reconnection Flow with Resume Cursors

When clients disconnect, the protocol supports precise state resumption through cursor-based replay rather than full state retransmission.

The Attach Command Structure

Clients initiate recovery by sending an attach command containing an optional resumeCursor. The following example from daemon-protocol.ts demonstrates the command structure:

import { createDaemonCommandEnvelope } from "../daemon-protocol";

const attachCmd = {
  type: "attach",
  activeSessionId: "session-42",
  resumeCursor: { generation: "gen-1", sequence: 12345 },
};

const envelope = createDaemonCommandEnvelope(
  attachCmd,
  "cmd-001",
  "client-a",
);
socket.write(JSON.stringify(envelope) + "\n");

If the cursor's generation matches the daemon's current supervisorGeneration, the daemon replays events from the specified sequence. Otherwise, it initiates a fresh snapshot stream via the three-step chunked sequence.

Chunked Delivery and Streaming

For large sessions, the daemon streams snapshots in chunks rather than loading entire transcripts into memory. If a reconnection occurs mid-stream, the daemon preserves partially-sent snapshots through the SnapshotTranscriptCache, allowing new clients to resume reading from the exact abort point without re-encoding the transcript.

Crash Recovery Mechanisms

When the daemon process crashes and restarts, the supervisor coordinates state preservation without client data loss through generation management and cache continuity.

Supervisor Generation Management

Upon restart, the supervisor increments supervisorGeneration and issues a new daemon_hello event. The old SnapshotTranscriptCache remains alive via reference counting (retain()/dispose()) until all readers release it, ensuring in-flight chunks remain readable even as a new worker initializes.

The protocol provides DaemonReplayInfo indicating replay availability: "complete", "partial", or "unavailable" based on cursor validity against the current lastEventSequence.

Replacement Snapshots and Cache Continuity

The new worker may emit a replacement snapshot with purpose: "replacement" and a fresh snapshotId. Clients receive this while potentially finishing consumption of the previous generation's chunks—a behavior validated in test ENG-4677 snapshot catch-up replacement (see packages/coding-agent/test/suite/regressions/4677-snapshot-catchup-replacement.test.ts). The following handler pattern demonstrates client-side processing:

socket.on("data", (line) => {
  const event = JSON.parse(line);
  switch (event.type) {
    case "session_snapshot_begin":
      if (event.purpose === "replacement") {
        // Start a new cache for the replacement snapshot
        currentCache = new SnapshotTranscriptCache({
          activeSessionId: event.activeSessionId,
          snapshotId: event.snapshotId,
          cacheRoot: "/tmp/daemon-snapshot",
        });
      }
      break;
    case "session_snapshot_chunk":
      currentCache.appendEncodedChunk(Buffer.from(line));
      break;
    case "session_snapshot_end":
      currentCache.markComplete();
      break;
  }
});

Once all readers complete the old snapshot (session_snapshot_end or session_snapshot_failed), the cache disposes safely and clients transition seamlessly to the new generation.

SnapshotTranscriptCache Implementation Details

The SnapshotTranscriptCache class in packages/coding-agent/src/modes/daemon/snapshot-transcript-cache.ts manages the storage layer for snapshot chunks with memory safety and concurrent access support.

Memory Buffering and Disk Spilling

The cache maintains chunks in memory as Buffer objects until the total exceeds memoryCacheBytes (default 4 MiB). Upon exceeding this limit, it creates a temporary directory at cacheRoot/<snapshotId>, flushes existing buffers to files, and writes subsequent chunks directly to disk.

Concurrent Reader Coordination

Multiple clients can share a single snapshot stream through retain() and dispose() methods that implement reference counting. The waitForChunk(index) method returns a Promise resolving when specific chunks become available, enabling efficient streaming without blocking. If errors occur, markFailed(error) propagates exceptions to all pending waiters.

The following implementation demonstrates consuming a cached snapshot stream:

async function readSnapshot(clientId: string) {
  const cache = new SnapshotTranscriptCache({
    activeSessionId: "session-42",
    snapshotId: "snapshot-1",
    cacheRoot: "/tmp/daemon-snapshot",
  });

  // Wait for each chunk as they arrive
  let index = 0;
  while (true) {
    const chunk = await cache.waitForChunk(index);
    if (!chunk) break; // end of snapshot
    // Process messages in the chunk
    const events = JSON.parse(chunk.toString()).messages;
    handleMessages(events);
    index++;
  }
}

Summary

  • The Prime Agent daemon protocol uses JSON-L over local sockets with explicit snapshot events (session_snapshot_begin, chunk, end, failed) for state synchronization, defined in daemon-protocol.ts.
  • Reconnection relies on resumeCursor values in the attach command to enable event replay or trigger fresh snapshots based on supervisorGeneration matching.
  • Crash recovery preserves in-flight snapshot data through the SnapshotTranscriptCache reference counting system, allowing replacement snapshots to coexist with ongoing stream completion.
  • The SnapshotTranscriptCache automatically spills to disk when memory limits (4 MiB default) are reached and supports concurrent readers without re-encoding transcripts.
  • Test ENG-4677 snapshot catch-up replacement validates that partial snapshots complete correctly after daemon restarts.

Frequently Asked Questions

What protocol version does Prime Agent use for daemon communication?

The daemon protocol operates at version 7 using JSON-L (JSON Lines) over a local socket, separating command submission from event streaming to handle asynchronous session updates efficiently.

How does the daemon handle large session transcripts during reconnection?

When the client advertises chunked_snapshot capability, the daemon splits the transcript into chunks limited by targetChunkBytes (approximately 512 KiB default) and streams them via session_snapshot_chunk events. The SnapshotTranscriptCache spills to disk if the total size exceeds 4 MiB, preventing memory exhaustion while maintaining stream integrity.

What happens to active snapshots when the daemon crashes?

The supervisor preserves the existing SnapshotTranscriptCache through reference counting, allowing connected clients to finish receiving queued chunks. The new worker may start a replacement snapshot with a new snapshotId, and clients transition seamlessly once the old cache fully disposes after all readers call their release callbacks.

How does the resume cursor mechanism determine whether to replay events or send a snapshot?

The daemon compares the generation field in the client's resumeCursor against its current supervisorGeneration. If they match, it replays events from the specified sequence number. If generations differ (indicating a daemon restart), it sends a full snapshot instead to ensure state consistency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →