Claude Context Snapshot Format and Indexing State Persistence Explained

Claude Context persists indexing state in a JSON snapshot file located at ~/.context/mcp-codebase-snapshot.json, managed by the SnapshotManager class which handles atomic writes, file locking, and automatic migration between legacy v1 and current v2 formats.

The Claude Context project from zilliztech/claude-context uses a sophisticated persistence mechanism to maintain codebase indexing state across process restarts. Understanding the snapshot format is essential for debugging indexing issues, migrating data between machines, or integrating with the Model Context Protocol (MCP) server internals. This article explores the JSON schema, version differences, and the atomic persistence strategy that prevents data corruption during concurrent access.

Snapshot File Location and Format

Claude Context stores all indexed codebase metadata in a single JSON file in the user's home directory.

File Location

The snapshot persists at:

~/.context/mcp-codebase-snapshot.json

This path is hardcoded in the configuration module and created automatically when the MCP server first initializes. The directory ~/.context/ also contains a .lock subdirectory used for file-locking during write operations.

JSON Structure Overview

The snapshot uses a discriminated union approach to versioning. The presence or absence of the formatVersion field determines which schema applies:

  • v1 (legacy): Missing formatVersion field, uses separate arrays for indexed and indexing codebases
  • v2 (current): Contains "formatVersion": "v2", uses a unified codebases record with rich metadata

Snapshot Versions and Data Models

The type definitions in packages/mcp/src/config.ts declare both supported formats.

Legacy v1 Format

The original snapshot structure used parallel arrays:

// From packages/mcp/src/config.ts
export interface CodebaseSnapshotV1 {
  indexedCodebases: string[];
  indexingCodebases: string[] | Record<string, number>;
  lastUpdated: string;
}

This format stores paths as simple strings, with indexingCodebases optionally containing progress percentages as a record mapping paths to completion percentages (0-100).

Current v2 Format

The modern format uses a path-keyed object with detailed status information:

// From packages/mcp/src/config.ts
export interface CodebaseSnapshotV2 {
  formatVersion: 'v2';
  codebases: Record<string, CodebaseInfo>;
  lastUpdated: string;
}

The codebases record maps absolute paths to CodebaseInfo objects, enabling storage of file counts, chunk totals, and error messages alongside the indexing status.

CodebaseInfo Union Types

CodebaseInfo is a discriminated union with three variants, distinguished by the status field:

Interface Status Value Key Fields
CodebaseInfoIndexing 'indexing' indexingPercentage: number (0-100)
CodebaseInfoIndexed 'indexed' indexedFiles: number, totalChunks: number, indexStatus: 'completed' | 'limit_reached'
CodebaseInfoIndexFailed 'indexfailed' errorMessage: string, lastAttemptedPercentage?: number

This design allows the snapshot to capture not just what is indexed, but how the indexing proceeded, including partial progress on failed attempts.

How Indexing State is Persisted

The SnapshotManager class in packages/mcp/src/snapshot.ts implements a crash-safe persistence layer using atomic file operations and directory locking.

Loading the Snapshot

When the MCP server starts, SnapshotManager.loadCodebaseSnapshot() executes the following sequence:

  1. Reads ~/.context/mcp-codebase-snapshot.json if it exists
  2. Detects the format version via isV2Format() helper
  3. Routes to loadV2Format() or loadV1Format() for deserialization
  4. Populates internal Map instances:
    • indexedCodebases: string[]
    • indexingCodebases: Map<string, number> (path → percentage)
    • codebaseInfoMap: Map<string, CodebaseInfo> (v2 rich data)
  5. Validates file-system existence, purging stale paths that no longer exist on disk

This validation prevents the index from referencing deleted directories after system changes.

Saving and Locking Mechanism

The saveCodebaseSnapshot() method (around line 380 in snapshot.ts) ensures data integrity through:

  1. Lock acquisition: Creates a .lock directory to prevent concurrent writes from multiple MCP instances
  2. Disk merging: Reads the current file state and merges any entries missing from memory (handles multi-process scenarios)
  3. Atomic writing: Constructs a fresh v2 JSON object from codebaseInfoMap and writes it atomically
  4. Cleanup: Clears the recentlyRemoved set and releases the lock

The write operation produces pretty-printed JSON for human readability while maintaining machine-parseable structure. All v1 data is automatically upgraded to v2 format on the next save operation.

Working with the SnapshotManager API

The SnapshotManager exposes a typed API for querying and updating indexing state used throughout the codebase.

Initializing and Loading

import { SnapshotManager } from '@zilliz/claude-context-mcp';

const snapshot = new SnapshotManager();
snapshot.loadCodebaseSnapshot();   // Reads ~/.context/mcp-codebase-snapshot.json

Querying Index Status

// Retrieve completed codebases
const indexed = snapshot.getIndexedCodebases();      // → ['/home/user/project-a']

// Retrieve in-progress codebases  
const indexing = snapshot.getIndexingCodebases();    // → ['/home/user/project-b']

// Check specific progress
const progress = snapshot.getIndexingProgress('/home/user/project-b');
console.log(`Indexing ${progress}% complete`);       // → Indexing 45% complete

Updating State During Indexing

State changes are persisted through typed setter methods:

// Mark as started
snapshot.setCodebaseIndexing('/path/to/repo', 0);

// Report incremental progress
snapshot.setCodebaseIndexing('/path/to/repo', 67);

// Mark as successfully completed
snapshot.setCodebaseIndexed('/path/to/repo', {
  indexedFiles: 150,
  totalChunks: 1200,
  status: 'completed',  // or 'limit_reached'
});

// Mark as failed with error context
snapshot.setCodebaseIndexFailed('/path/to/repo', {
  errorMessage: 'Embedding API rate limit exceeded',
  lastAttemptedPercentage: 67,
});

Each state transition in packages/mcp/src/handlers.ts (lines 130-150) invokes these methods, ensuring the JSON snapshot remains synchronized with the in-memory indexing pipeline.

Summary

  • Claude Context persists indexing state in ~/.context/mcp-codebase-snapshot.json using a JSON-based v2 format
  • The snapshot format migrated from simple arrays (v1) to a rich codebases record structure (v2) that tracks file counts, chunks, and error states
  • SnapshotManager in packages/mcp/src/snapshot.ts handles atomic writes using directory locking and disk merging to prevent corruption
  • Legacy v1 snapshots are automatically upgraded to v2 on the next save operation
  • The CodebaseInfo union type supports three states: indexing, indexed, and indexfailed, with specific metadata for each

Frequently Asked Questions

Where does Claude Context store its indexing state?

Claude Context stores indexing state in ~/.context/mcp-codebase-snapshot.json in the user's home directory. This JSON file is managed by the SnapshotManager class and updated atomically whenever indexing status changes, ensuring persistence across server restarts and crashes.

What is the difference between v1 and v2 snapshot formats?

The v1 format used separate indexedCodebases and indexingCodebases arrays with simple path strings, while v2 uses a unified codebases object keyed by path with rich metadata including file counts, chunk totals, and error messages. The v2 format is identified by the "formatVersion": "v2" field and is the current standard, though v1 is still supported for backward compatibility.

How does Claude Context prevent snapshot corruption during concurrent writes?

The SnapshotManager prevents corruption by acquiring a .lock directory before writing, merging any disk changes that occurred since the last read, and performing atomic file writes. This locking mechanism allows multiple MCP server instances to share the same snapshot file without losing state updates or corrupting the JSON structure.

What information is stored for a failed indexing attempt?

Failed codebases are stored with the indexfailed status and include an errorMessage string describing the failure reason and an optional lastAttemptedPercentage number indicating how far the indexing progressed before the error occurred. This allows the system to resume from the last known progress or report detailed failure diagnostics to the user.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →