# Claude Context Snapshot Format and Indexing State Persistence Explained

> Discover the Claude Context snapshot format and how the SnapshotManager ensures indexing state persistence. Learn about atomic writes and automatic migration between v1 and v2 formats.

- Repository: [Zilliz/claude-context](https://github.com/zilliztech/claude-context)
- Tags: internals
- Published: 2026-04-22

---

**Claude Context persists indexing state in a JSON snapshot file located at `~/.context/mcp-codebase-snapshot.json`, managed by the `SnapshotManager` class which handles atomic writes, file locking, and automatic migration between legacy v1 and current v2 formats.**

The Claude Context project from `zilliztech/claude-context` uses a sophisticated persistence mechanism to maintain codebase indexing state across process restarts. Understanding the snapshot format is essential for debugging indexing issues, migrating data between machines, or integrating with the Model Context Protocol (MCP) server internals. This article explores the JSON schema, version differences, and the atomic persistence strategy that prevents data corruption during concurrent access.

## Snapshot File Location and Format

Claude Context stores all indexed codebase metadata in a single JSON file in the user's home directory.

### File Location

The snapshot persists at:

```bash
~/.context/mcp-codebase-snapshot.json

```

This path is hardcoded in the configuration module and created automatically when the MCP server first initializes. The directory `~/.context/` also contains a `.lock` subdirectory used for file-locking during write operations.

### JSON Structure Overview

The snapshot uses a **discriminated union** approach to versioning. The presence or absence of the `formatVersion` field determines which schema applies:

- **v1 (legacy)**: Missing `formatVersion` field, uses separate arrays for indexed and indexing codebases
- **v2 (current)**: Contains `"formatVersion": "v2"`, uses a unified `codebases` record with rich metadata

## Snapshot Versions and Data Models

The type definitions in [`packages/mcp/src/config.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/mcp/src/config.ts) declare both supported formats.

### Legacy v1 Format

The original snapshot structure used parallel arrays:

```typescript
// From packages/mcp/src/config.ts
export interface CodebaseSnapshotV1 {
  indexedCodebases: string[];
  indexingCodebases: string[] | Record<string, number>;
  lastUpdated: string;
}

```

This format stores paths as simple strings, with `indexingCodebases` optionally containing progress percentages as a record mapping paths to completion percentages (0-100).

### Current v2 Format

The modern format uses a path-keyed object with detailed status information:

```typescript
// From packages/mcp/src/config.ts
export interface CodebaseSnapshotV2 {
  formatVersion: 'v2';
  codebases: Record<string, CodebaseInfo>;
  lastUpdated: string;
}

```

The `codebases` record maps absolute paths to `CodebaseInfo` objects, enabling storage of file counts, chunk totals, and error messages alongside the indexing status.

### CodebaseInfo Union Types

`CodebaseInfo` is a discriminated union with three variants, distinguished by the `status` field:

| Interface | Status Value | Key Fields |
|-----------|--------------|------------|
| `CodebaseInfoIndexing` | `'indexing'` | `indexingPercentage: number` (0-100) |
| `CodebaseInfoIndexed` | `'indexed'` | `indexedFiles: number`, `totalChunks: number`, `indexStatus: 'completed' \| 'limit_reached'` |
| `CodebaseInfoIndexFailed` | `'indexfailed'` | `errorMessage: string`, `lastAttemptedPercentage?: number` |

This design allows the snapshot to capture not just *what* is indexed, but *how* the indexing proceeded, including partial progress on failed attempts.

## How Indexing State is Persisted

The `SnapshotManager` class in [`packages/mcp/src/snapshot.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/mcp/src/snapshot.ts) implements a crash-safe persistence layer using atomic file operations and directory locking.

### Loading the Snapshot

When the MCP server starts, `SnapshotManager.loadCodebaseSnapshot()` executes the following sequence:

1. Reads `~/.context/mcp-codebase-snapshot.json` if it exists
2. Detects the format version via `isV2Format()` helper
3. Routes to `loadV2Format()` or `loadV1Format()` for deserialization
4. Populates internal `Map` instances:
   - `indexedCodebases: string[]`
   - `indexingCodebases: Map<string, number>` (path → percentage)
   - `codebaseInfoMap: Map<string, CodebaseInfo>` (v2 rich data)
5. Validates file-system existence, purging stale paths that no longer exist on disk

This validation prevents the index from referencing deleted directories after system changes.

### Saving and Locking Mechanism

The `saveCodebaseSnapshot()` method (around line 380 in [`snapshot.ts`](https://github.com/zilliztech/claude-context/blob/main/snapshot.ts)) ensures data integrity through:

1. **Lock acquisition**: Creates a `.lock` directory to prevent concurrent writes from multiple MCP instances
2. **Disk merging**: Reads the current file state and merges any entries missing from memory (handles multi-process scenarios)
3. **Atomic writing**: Constructs a fresh v2 JSON object from `codebaseInfoMap` and writes it atomically
4. **Cleanup**: Clears the `recentlyRemoved` set and releases the lock

The write operation produces pretty-printed JSON for human readability while maintaining machine-parseable structure. All v1 data is automatically upgraded to v2 format on the next save operation.

## Working with the SnapshotManager API

The `SnapshotManager` exposes a typed API for querying and updating indexing state used throughout the codebase.

### Initializing and Loading

```typescript
import { SnapshotManager } from '@zilliz/claude-context-mcp';

const snapshot = new SnapshotManager();
snapshot.loadCodebaseSnapshot();   // Reads ~/.context/mcp-codebase-snapshot.json

```

### Querying Index Status

```typescript
// Retrieve completed codebases
const indexed = snapshot.getIndexedCodebases();      // → ['/home/user/project-a']

// Retrieve in-progress codebases  
const indexing = snapshot.getIndexingCodebases();    // → ['/home/user/project-b']

// Check specific progress
const progress = snapshot.getIndexingProgress('/home/user/project-b');
console.log(`Indexing ${progress}% complete`);       // → Indexing 45% complete

```

### Updating State During Indexing

State changes are persisted through typed setter methods:

```typescript
// Mark as started
snapshot.setCodebaseIndexing('/path/to/repo', 0);

// Report incremental progress
snapshot.setCodebaseIndexing('/path/to/repo', 67);

// Mark as successfully completed
snapshot.setCodebaseIndexed('/path/to/repo', {
  indexedFiles: 150,
  totalChunks: 1200,
  status: 'completed',  // or 'limit_reached'
});

// Mark as failed with error context
snapshot.setCodebaseIndexFailed('/path/to/repo', {
  errorMessage: 'Embedding API rate limit exceeded',
  lastAttemptedPercentage: 67,
});

```

Each state transition in [`packages/mcp/src/handlers.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/mcp/src/handlers.ts) (lines 130-150) invokes these methods, ensuring the JSON snapshot remains synchronized with the in-memory indexing pipeline.

## Summary

- **Claude Context** persists indexing state in `~/.context/mcp-codebase-snapshot.json` using a JSON-based v2 format
- The **snapshot format** migrated from simple arrays (v1) to a rich `codebases` record structure (v2) that tracks file counts, chunks, and error states
- **SnapshotManager** in [`packages/mcp/src/snapshot.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/mcp/src/snapshot.ts) handles atomic writes using directory locking and disk merging to prevent corruption
- Legacy v1 snapshots are automatically upgraded to v2 on the next save operation
- The **CodebaseInfo** union type supports three states: `indexing`, `indexed`, and `indexfailed`, with specific metadata for each

## Frequently Asked Questions

### Where does Claude Context store its indexing state?

Claude Context stores indexing state in `~/.context/mcp-codebase-snapshot.json` in the user's home directory. This JSON file is managed by the `SnapshotManager` class and updated atomically whenever indexing status changes, ensuring persistence across server restarts and crashes.

### What is the difference between v1 and v2 snapshot formats?

The v1 format used separate `indexedCodebases` and `indexingCodebases` arrays with simple path strings, while v2 uses a unified `codebases` object keyed by path with rich metadata including file counts, chunk totals, and error messages. The v2 format is identified by the `"formatVersion": "v2"` field and is the current standard, though v1 is still supported for backward compatibility.

### How does Claude Context prevent snapshot corruption during concurrent writes?

The `SnapshotManager` prevents corruption by acquiring a `.lock` directory before writing, merging any disk changes that occurred since the last read, and performing atomic file writes. This locking mechanism allows multiple MCP server instances to share the same snapshot file without losing state updates or corrupting the JSON structure.

### What information is stored for a failed indexing attempt?

Failed codebases are stored with the `indexfailed` status and include an `errorMessage` string describing the failure reason and an optional `lastAttemptedPercentage` number indicating how far the indexing progressed before the error occurred. This allows the system to resume from the last known progress or report detailed failure diagnostics to the user.