Asynchronous Indexing in Claude Context: Complete Workflow and Progress Tracking Guide
Claude Context indexes codebases asynchronously through a background job system that reports progress every ~2 seconds to a persistent snapshot file, enabling real-time UI updates and crash recovery.
The Claude Context extension (from the zilliztech/claude-context repository) provides a robust asynchronous indexing pipeline that converts your codebase into searchable vector embeddings without blocking the editor. This article examines the complete workflow—from initial request to completion—and explains how progress is tracked through every stage.
Overview of the Asynchronous Indexing Workflow
The indexing process spans six distinct phases, each handled by specific modules in the codebase:
- Request handling —
ToolHandlers.handleIndexCodebasevalidates paths and triggers cloud synchronization - State initialization — Snapshot writes an
"indexing"entry with 0% progress - Background job launch —
startBackgroundIndexingspawns the async worker - File processing —
Context.indexCodebasescans, chunks, and embeds files - Progress reporting — Callbacks update the snapshot every ~2 seconds
- Finalization — State transitions to
"indexed"or"indexfailed"
The packages/mcp/src/handlers.ts file orchestrates this entire flow, delegating core indexing work to packages/core/src/context.ts.
Step-by-Step Workflow Deep Dive
Step 1: Request Validation and Cloud Sync
When you trigger indexing from VS Code, the index_codebase tool is invoked. In packages/mcp/src/handlers.ts (lines 15-30), handleIndexCodebase performs:
- Path argument parsing and validation
- Forced cloud synchronization via
forceCloudSync - Folder existence checks
- Detection of existing indexing operations to prevent conflicts
- Optional clearing of stale state with
force: true
// packages/mcp/src/handlers.ts
export class ToolHandlers {
async handleIndexCodebase(args: {
path: string;
force?: boolean;
splitter?: 'ast' | 'line';
customExtensions?: string[];
ignorePatterns?: string[];
}): Promise<ToolResult> {
// Validation and cloud sync occurs here (lines 15-30)
const absolutePath = path.resolve(args.path);
await this.forceCloudSync();
// ... validation logic
await this.startBackgroundIndexing(absolutePath, args);
}
}
Step 2: Snapshot State Initialization
Before background work begins, the system records an "indexing" state in packages/mcp/src/snapshot.ts (lines 42-56). The setCodebaseIndexing method:
- Adds the path to
indexingCodebasesmap with 0% progress - Creates a
CodebaseInfoIndexingrecord with timestamp - Triggers immediate snapshot persistence via
saveCodebaseSnapshot
This guarantees crash recovery—even if the MCP server restarts, the snapshot reveals incomplete work.
Step 3: Background Job Launch
The startBackgroundIndexing method (lines 9-23, 45-53 in handlers.ts) prepares the async environment:
// packages/mcp/src/handlers.ts
private async startBackgroundIndexing(
absolutePath: string,
args: IndexCodebaseArgs
): Promise<void> {
// Load ignore patterns for this codebase
const ignorePatterns = await this.context.getLoadedIgnorePatterns(
absolutePath,
args.ignorePatterns
);
// Create file synchronizer for change detection
const synchronizer = new FileSynchronizer(absolutePath, ignorePatterns);
// Prepare vector collection
await this.context.prepareCollection(absolutePath);
// Launch actual indexing with progress callback
await this.context.indexCodebase(
absolutePath,
{ /* options */ },
(progress) => this.onProgressUpdate(absolutePath, progress)
);
}
Step 4: File Processing and Embedding
The core indexing engine in packages/core/src/context.ts performs:
- File discovery —
getCodeFiles(lines 66-71) recursively walks the codebase, applying ignore patterns - Batch processing —
processFileList(lines 82-100) handles files in streaming batches - Chunk generation — AST or line-based splitting based on configuration
- Embedding creation — Batches sent to embedding model
- Vector insertion — Chunks stored in Milvus/Zilliz collection
// packages/core/src/context.ts
export class Context {
async indexCodebase(
codebasePath: string,
options: IndexOptions,
onProgress: (info: ProgressInfo) => void
): Promise<void> {
const files = await this.getCodeFiles(codebasePath, options.ignorePatterns);
await this.processFileList(files, {
onBatch: (chunks) => {
// Generate embeddings and insert
return this.processChunkBuffer(chunks);
},
onProgress // Forward progress updates
});
}
}
Step 5: Progress Reporting Mechanism
Progress flows through a callback chain culminating in snapshot updates. In handlers.ts (lines 53-57, 58-62):
// Inside startBackgroundIndexing
const onProgressUpdate = (progress: ProgressInfo) => {
// 1. Update in-memory map immediately
this.snapshotManager.setCodebaseIndexing(absolutePath, progress.percentage);
// 2. Throttle disk writes to ~2 second intervals
const now = Date.now();
if (now - lastSaveTime >= 2000) {
this.snapshotManager.saveCodebaseSnapshot();
lastSaveTime = now;
}
};
The SnapshotManager maintains indexingCodebases: Map<string, number> for fast lookups while saveCodebaseSnapshot() serializes the full CodebaseSnapshotV2 structure to ~/.context/mcp-codebase-snapshot.json.
Step 6: Completion and Error Handling
When indexCodebase resolves or rejects, handlers.ts (lines 66-71, 83-92) finalizes state:
try {
await this.context.indexCodebase(/* ... */);
// Success: mark as indexed with final stats
this.snapshotManager.setCodebaseIndexed(absolutePath, {
fileCount: progress.totalFiles,
chunkCount: progress.totalChunks,
durationMs: Date.now() - startTime
});
} catch (error) {
// Failure: preserve error and last known progress
this.snapshotManager.setCodebaseIndexFailed(absolutePath, {
error: error.message,
lastProgress: progress?.percentage || 0
});
} finally {
// Always persist final state
await this.snapshotManager.saveCodebaseSnapshot();
}
How Progress Is Tracked: Architecture Deep Dive
The Snapshot as Source of Truth
Claude Context uses a dual-layer persistence strategy for progress tracking:
| Layer | Data Structure | Persistence | Access Pattern |
|---|---|---|---|
| Runtime | Map<string, number> |
Memory only | Instant reads/writes |
| Persistent | CodebaseSnapshotV2 |
JSON file (~/.context/) | ~2s throttled writes |
The packages/mcp/src/snapshot.ts module encapsulates this logic. Key methods include:
setCodebaseIndexing(path, percent)— Updates runtime mapsaveCodebaseSnapshot()— Serializes to diskgetCodebaseStatus(path)— Returns current state for UI queries
Progress Callback Chain
Progress flows from the core engine through multiple layers:
Context.indexCodebase()
└── processFileList()
└── onProgress callback (per batch)
└── handlers.ts onProgressUpdate()
├── SnapshotManager.setCodebaseIndexing()
└── [every 2s] saveCodebaseSnapshot()
The ProgressInfo object contains:
percentage— Float from 0-100processedFiles/totalFilesprocessedChunks/totalChunkscurrentFile— Currently processing path
Querying Indexing Status
The get_indexing_status tool (implemented in handlers.ts lines 11-30) enables polling:
// Called by VS Code extension every few seconds
async handleGetIndexingStatus(args: { path: string }): Promise<ToolResult> {
const status = this.snapshotManager.getCodebaseStatus(args.path);
switch (status.state) {
case 'indexing':
return {
content: [{
text: `🔄 Codebase ${args.path} is currently being indexed. Progress: ${status.progress}%`
}]
};
case 'indexed':
return { content: [{ text: `✅ Indexed: ${status.stats.chunkCount} chunks` }] };
case 'indexfailed':
return { content: [{ text: `❌ Failed: ${status.error}` }] };
default:
return { content: [{ text: `⚠️ Not indexed` }] };
}
}
Practical Example: Triggering and Monitoring Indexing
// Trigger indexing from a Node.js MCP client
import { Client } from '@modelcontextprotocol/sdk/client/index.js';
const client = new Client({ name: 'example', version: '1.0.0' });
// 1. Start asynchronous indexing
await client.runTool('index_codebase', {
path: '/home/user/my-project',
force: false,
splitter: 'ast',
ignorePatterns: ['node_modules/**', '*.log', 'dist/**']
});
// 2. Poll for progress until complete
const pollInterval = setInterval(async () => {
const result = await client.runTool('get_indexing_status', {
path: '/home/user/my-project'
});
const message = result.content[0].text;
console.log(message);
// Check for completion or failure
if (message.includes('✅') || message.includes('❌')) {
clearInterval(pollInterval);
// Verify snapshot was updated
const finalStatus = await client.runTool('get_indexing_status', {
path: '/home/user/my-project'
});
console.log('Final:', finalStatus.content[0].text);
}
}, 3000);
Summary
-
The asynchronous indexing workflow in Claude Context spans six phases: request validation, snapshot initialization, background job launch, file processing, progress reporting, and final state commit.
-
Progress tracking relies on a dual-layer system: an in-memory
Mapfor instant updates and a throttled JSON snapshot (~/.context/mcp-codebase-snapshot.json) persisted every ~2 seconds. -
Key source files implementing this workflow include
packages/mcp/src/handlers.ts(orchestration),packages/mcp/src/snapshot.ts(persistence), andpackages/core/src/context.ts(core indexing engine). -
Status queries use the
get_indexing_statustool, which reads the snapshot and returns human-readable states: indexing, indexed, indexfailed, or not indexed.
Frequently Asked Questions
How does Claude Context recover indexing progress after a crash?
The mcp-codebase-snapshot.json file in ~/.context/ serves as the recovery source. On restart, SnapshotManager loads this file into indexingCodebases: Map<string, number>, restoring the exact percentage and state. The handleGetIndexingStatus tool then reports this recovered progress to the UI.
What happens if indexing fails midway through a large codebase?
The error handler in handlers.ts (lines 83-92) calls setCodebaseIndexFailed with the error message and last known progress percentage. This state is immediately persisted to the snapshot. Users can query get_indexing_status to see the failure reason and partial progress, then retry with force: true if needed.
How frequently does the UI receive progress updates?
The VS Code extension polls get_indexing_status every few seconds. The underlying snapshot updates every ~2 seconds (throttled in handlers.ts lines 58-62), with in-memory updates occurring immediately on each batch completion. This creates a maximum ~2-second lag between actual progress and UI display.
Can multiple codebases be indexed simultaneously?
Yes. The indexingCodebases Map stores progress for each path independently, and startBackgroundIndexing spawns separate async flows per invocation. The snapshot tracks all active jobs, and handleGetIndexingStatus accepts a path parameter to query specific codebases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →