How OpenClaude Manages Conversation Recovery Across Different Transports

OpenClaude manages conversation recovery by persisting conversation snapshots to JSON files keyed by transport-specific identifiers, then replaying those messages into new transport instances when connections are re-established.

The OpenClaude project implements a transport-agnostic conversation recovery system that allows sessions to survive interruptions regardless of whether users connect via REPL, HTTP API, or WebSocket. This article examines the source code architecture behind this feature, tracing how conversationRecovery.ts coordinates with transport implementations to restore conversational state.

Core Recovery Infrastructure in conversationRecovery.ts

The conversation recovery system centers on src/utils/conversationRecovery.ts, which provides deterministic snapshot management with atomic write semantics. This module exposes two primary functions that all transports invoke:

  • saveConversationSnapshot(transportId, messages) — Serializes the current message array to a JSON file at <runtime-dir>/recovery/<transportId>.json using atomic filesystem operations to prevent corruption during crashes.
  • loadConversationSnapshot(transportId) — Reads and parses the snapshot file, returning the message array or null if no recovery data exists.

The implementation handles edge cases including malformed JSON, permission errors, and concurrent access scenarios. Unit tests in src/utils/conversationRecovery.test.ts validate that snapshots survive process restarts and that corrupted files fail gracefully without crashing the recovery flow.

The BridgeTransport Interface Contract

All OpenClaude transports implement a common BridgeTransport interface defined in the bridge module. Two methods form the recovery contract:

Method Responsibility
getTransportId(): string Returns a stable, deterministic identifier unique to this transport instance (e.g., repl:<pid>, http:<request-id>, ws:<socket-id>).
onRecover(messages: Message[]): void Receives the restored message array and re-injects it into the active session, reconstructing conversation context including tool calls, system prompts, and user turns.

The REPL transport implementation in src/bridge/replBridgeTransport.ts demonstrates this contract in practice. It derives its transport ID from the process PID, ensuring that a restarted REPL process can locate its predecessor's snapshot. The HTTP and WebSocket transports follow identical patterns with identifiers tied to request or socket IDs respectively.

Recovery Flow in bridgeMain.ts

The central coordination point src/bridge/bridgeMain.ts orchestrates the recovery lifecycle across all transport types:

  1. Session initialization — bridgeMain queries the transport for its ID via getTransportId(), then attempts loadConversationSnapshot().
  2. State restoration — If a snapshot exists, bridgeMain invokes transport.onRecover(messages) before accepting new input.
  3. Continuous persistence — After each message is processed, bridgeMain calls saveConversationSnapshot() with the updated message array.

This design decouples recovery logic from transport-specific implementation details. The same bridgeMain code paths handle REPL crashes, HTTP request timeouts, and WebSocket disconnections without transport-specific branching.

Transport-Specific Recovery Behaviors

REPL Transport Recovery

The REPL implementation in replBridgeTransport.ts provides the most intuitive recovery experience. When a user restarts the openclaude repl command:

// From replBridgeTransport.ts — simplified recovery sequence
const transportId = `repl:${process.pid}`;
const priorMessages = await loadConversationSnapshot(transportId);

if (priorMessages && priorMessages.length > 0) {
  this.conversationHistory = priorMessages;
  this.renderPriorConversation();  // Re-display previous turns to user
  this.onRecover(priorMessages);   // Notify bridgeMain of restored state
}

The previous conversation appears immediately in the terminal, allowing users to continue as if uninterrupted.

HTTP API Recovery

HTTP transports use request-scoped identifiers. When a client reconnects with the same session token, bridgeMain locates the corresponding snapshot and restores the conversation state before processing the new request. This enables long-running conversational sessions over stateless HTTP connections.

WebSocket Recovery

WebSocket transports implement recovery through the same interface, using socket connection IDs as transport identifiers. On reconnection, clients receive their full conversation history replayed through the WebSocket, maintaining real-time conversational continuity.

Practical Implementation Pattern

Transport developers follow this pattern to integrate recovery:

import {
  saveConversationSnapshot,
  loadConversationSnapshot,
} from '@/utils/conversationRecovery';

class MyCustomTransport implements BridgeTransport {
  private messages: Message[] = [];
  private readonly transportId: string;

  constructor(id: string) {
    this.transportId = `custom:${id}`;
    this.attemptRecovery();
  }

  private async attemptRecovery(): Promise<void> {
    const recovered = await loadConversationSnapshot(this.transportId);
    if (recovered) {
      this.messages = recovered;
      this.emit('recovered', recovered);
    }
  }

  async processMessage(msg: Message): Promise<void> {
    // Process through OpenClaude engine...
    this.messages.push(msg);
    await saveConversationSnapshot(this.transportId, this.messages);
  }

  getTransportId(): string {
    return this.transportId;
  }

  onRecover(messages: Message[]): void {
    this.messages = messages;
    this.notifyClientOfRestoredState();
  }
}

This pattern appears consistently across replBridgeTransport.ts, the HTTP bridge implementation, and the WebSocket bridge implementation.

File-Level Architecture Reference

File Recovery Responsibility
src/utils/conversationRecovery.ts Core snapshot persistence and retrieval with atomic writes and error tolerance.
src/utils/conversationRecovery.test.ts Validates snapshot durability, corruption handling, and concurrent access.
src/bridge/bridgeMain.ts Orchestrates recovery flow across all transport types; invokes load/save operations.
src/bridge/replBridgeTransport.ts REPL-specific transport implementing BridgeTransport recovery contract.
src/bridge/sessionRunner.ts Manages session lifecycle and triggers recovery on transport reconnection.

Summary

  • Transport-agnostic design: All OpenClaude transports share the same recovery infrastructure through the BridgeTransport interface.
  • Atomic persistence: conversationRecovery.ts guarantees snapshot integrity using atomic filesystem operations.
  • Stable identifiers: Transport-specific IDs ensure snapshots are correctly matched to reconnecting clients.
  • Centralized coordination: bridgeMain.ts handles recovery uniformly, eliminating transport-specific recovery logic duplication.
  • Seamless user experience: Conversations restore automatically without user intervention across REPL, HTTP, and WebSocket transports.

Frequently Asked Questions

How does OpenClaude prevent snapshot corruption during crashes?

The saveConversationSnapshot() function in conversationRecovery.ts writes to a temporary file then performs an atomic rename operation. This ensures that the snapshot file is never in a partially written state, and loadConversationSnapshot() validates JSON parsing with try-catch blocks that return null for unreadable files.

Can conversation recovery work across different machine instances?

Recovery depends on access to the same filesystem where snapshots are stored (<runtime-dir>/recovery/). For distributed deployments, OpenClaude would require either shared storage or a pluggable backend for conversationRecovery.ts to store snapshots in a network-accessible location.

What happens if a transport ID collides between sessions?

Transport IDs are designed to be globally unique using compound identifiers (transport type plus UUID, PID, or request ID). The REPL transport uses repl:${process.pid}, HTTP uses http:${requestId}, and WebSocket uses ws:${socketId}, making collisions statistically improbable within the same runtime environment.

Is there a retention policy for old conversation snapshots?

The current implementation in conversationRecovery.ts does not automatically purge old snapshots. Transports or deployment operators must implement cleanup policies appropriate to their storage constraints, as snapshot files accumulate indefinitely until manually removed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →