Asynchronous Indexing in Claude Context: Complete Workflow and Progress Tracking Guide

Claude Context indexes codebases asynchronously through a background job system that reports progress every ~2 seconds to a persistent snapshot file, enabling real-time UI updates and crash recovery.

The Claude Context extension (from the zilliztech/claude-context repository) provides a robust asynchronous indexing pipeline that converts your codebase into searchable vector embeddings without blocking the editor. This article examines the complete workflow—from initial request to completion—and explains how progress is tracked through every stage.

Overview of the Asynchronous Indexing Workflow

The indexing process spans six distinct phases, each handled by specific modules in the codebase:

  1. Request handling — ToolHandlers.handleIndexCodebase validates paths and triggers cloud synchronization
  2. State initialization — Snapshot writes an "indexing" entry with 0% progress
  3. Background job launch — startBackgroundIndexing spawns the async worker
  4. File processing — Context.indexCodebase scans, chunks, and embeds files
  5. Progress reporting — Callbacks update the snapshot every ~2 seconds
  6. Finalization — State transitions to "indexed" or "indexfailed"

The packages/mcp/src/handlers.ts file orchestrates this entire flow, delegating core indexing work to packages/core/src/context.ts.

Step-by-Step Workflow Deep Dive

Step 1: Request Validation and Cloud Sync

When you trigger indexing from VS Code, the index_codebase tool is invoked. In packages/mcp/src/handlers.ts (lines 15-30), handleIndexCodebase performs:

  • Path argument parsing and validation
  • Forced cloud synchronization via forceCloudSync
  • Folder existence checks
  • Detection of existing indexing operations to prevent conflicts
  • Optional clearing of stale state with force: true
// packages/mcp/src/handlers.ts
export class ToolHandlers {
  async handleIndexCodebase(args: {
    path: string;
    force?: boolean;
    splitter?: 'ast' | 'line';
    customExtensions?: string[];
    ignorePatterns?: string[];
  }): Promise<ToolResult> {
    // Validation and cloud sync occurs here (lines 15-30)
    const absolutePath = path.resolve(args.path);
    await this.forceCloudSync();
    // ... validation logic
    await this.startBackgroundIndexing(absolutePath, args);
  }
}

Step 2: Snapshot State Initialization

Before background work begins, the system records an "indexing" state in packages/mcp/src/snapshot.ts (lines 42-56). The setCodebaseIndexing method:

  • Adds the path to indexingCodebases map with 0% progress
  • Creates a CodebaseInfoIndexing record with timestamp
  • Triggers immediate snapshot persistence via saveCodebaseSnapshot

This guarantees crash recovery—even if the MCP server restarts, the snapshot reveals incomplete work.

Step 3: Background Job Launch

The startBackgroundIndexing method (lines 9-23, 45-53 in handlers.ts) prepares the async environment:

// packages/mcp/src/handlers.ts
private async startBackgroundIndexing(
  absolutePath: string,
  args: IndexCodebaseArgs
): Promise<void> {
  // Load ignore patterns for this codebase
  const ignorePatterns = await this.context.getLoadedIgnorePatterns(
    absolutePath,
    args.ignorePatterns
  );
  
  // Create file synchronizer for change detection
  const synchronizer = new FileSynchronizer(absolutePath, ignorePatterns);
  
  // Prepare vector collection
  await this.context.prepareCollection(absolutePath);
  
  // Launch actual indexing with progress callback
  await this.context.indexCodebase(
    absolutePath,
    { /* options */ },
    (progress) => this.onProgressUpdate(absolutePath, progress)
  );
}

Step 4: File Processing and Embedding

The core indexing engine in packages/core/src/context.ts performs:

  1. File discovery — getCodeFiles (lines 66-71) recursively walks the codebase, applying ignore patterns
  2. Batch processing — processFileList (lines 82-100) handles files in streaming batches
  3. Chunk generation — AST or line-based splitting based on configuration
  4. Embedding creation — Batches sent to embedding model
  5. Vector insertion — Chunks stored in Milvus/Zilliz collection
// packages/core/src/context.ts
export class Context {
  async indexCodebase(
    codebasePath: string,
    options: IndexOptions,
    onProgress: (info: ProgressInfo) => void
  ): Promise<void> {
    const files = await this.getCodeFiles(codebasePath, options.ignorePatterns);
    
    await this.processFileList(files, {
      onBatch: (chunks) => {
        // Generate embeddings and insert
        return this.processChunkBuffer(chunks);
      },
      onProgress // Forward progress updates
    });
  }
}

Step 5: Progress Reporting Mechanism

Progress flows through a callback chain culminating in snapshot updates. In handlers.ts (lines 53-57, 58-62):

// Inside startBackgroundIndexing
const onProgressUpdate = (progress: ProgressInfo) => {
  // 1. Update in-memory map immediately
  this.snapshotManager.setCodebaseIndexing(absolutePath, progress.percentage);
  
  // 2. Throttle disk writes to ~2 second intervals
  const now = Date.now();
  if (now - lastSaveTime >= 2000) {
    this.snapshotManager.saveCodebaseSnapshot();
    lastSaveTime = now;
  }
};

The SnapshotManager maintains indexingCodebases: Map<string, number> for fast lookups while saveCodebaseSnapshot() serializes the full CodebaseSnapshotV2 structure to ~/.context/mcp-codebase-snapshot.json.

Step 6: Completion and Error Handling

When indexCodebase resolves or rejects, handlers.ts (lines 66-71, 83-92) finalizes state:

try {
  await this.context.indexCodebase(/* ... */);
  
  // Success: mark as indexed with final stats
  this.snapshotManager.setCodebaseIndexed(absolutePath, {
    fileCount: progress.totalFiles,
    chunkCount: progress.totalChunks,
    durationMs: Date.now() - startTime
  });
  
} catch (error) {
  // Failure: preserve error and last known progress
  this.snapshotManager.setCodebaseIndexFailed(absolutePath, {
    error: error.message,
    lastProgress: progress?.percentage || 0
  });
} finally {
  // Always persist final state
  await this.snapshotManager.saveCodebaseSnapshot();
}

How Progress Is Tracked: Architecture Deep Dive

The Snapshot as Source of Truth

Claude Context uses a dual-layer persistence strategy for progress tracking:

Layer Data Structure Persistence Access Pattern
Runtime Map<string, number> Memory only Instant reads/writes
Persistent CodebaseSnapshotV2 JSON file (~/.context/) ~2s throttled writes

The packages/mcp/src/snapshot.ts module encapsulates this logic. Key methods include:

  • setCodebaseIndexing(path, percent) — Updates runtime map
  • saveCodebaseSnapshot() — Serializes to disk
  • getCodebaseStatus(path) — Returns current state for UI queries

Progress Callback Chain

Progress flows from the core engine through multiple layers:


Context.indexCodebase()
  └── processFileList()
        └── onProgress callback (per batch)
              └── handlers.ts onProgressUpdate()
                    ├── SnapshotManager.setCodebaseIndexing()
                    └── [every 2s] saveCodebaseSnapshot()

The ProgressInfo object contains:

  • percentage — Float from 0-100
  • processedFiles / totalFiles
  • processedChunks / totalChunks
  • currentFile — Currently processing path

Querying Indexing Status

The get_indexing_status tool (implemented in handlers.ts lines 11-30) enables polling:

// Called by VS Code extension every few seconds
async handleGetIndexingStatus(args: { path: string }): Promise<ToolResult> {
  const status = this.snapshotManager.getCodebaseStatus(args.path);
  
  switch (status.state) {
    case 'indexing':
      return {
        content: [{ 
          text: `🔄 Codebase ${args.path} is currently being indexed. Progress: ${status.progress}%` 
        }]
      };
    case 'indexed':
      return { content: [{ text: `✅ Indexed: ${status.stats.chunkCount} chunks` }] };
    case 'indexfailed':
      return { content: [{ text: `❌ Failed: ${status.error}` }] };
    default:
      return { content: [{ text: `⚠️ Not indexed` }] };
  }
}

Practical Example: Triggering and Monitoring Indexing

// Trigger indexing from a Node.js MCP client
import { Client } from '@modelcontextprotocol/sdk/client/index.js';

const client = new Client({ name: 'example', version: '1.0.0' });

// 1. Start asynchronous indexing
await client.runTool('index_codebase', {
  path: '/home/user/my-project',
  force: false,
  splitter: 'ast',
  ignorePatterns: ['node_modules/**', '*.log', 'dist/**']
});

// 2. Poll for progress until complete
const pollInterval = setInterval(async () => {
  const result = await client.runTool('get_indexing_status', {
    path: '/home/user/my-project'
  });
  
  const message = result.content[0].text;
  console.log(message);
  
  // Check for completion or failure
  if (message.includes('✅') || message.includes('❌')) {
    clearInterval(pollInterval);
    
    // Verify snapshot was updated
    const finalStatus = await client.runTool('get_indexing_status', {
      path: '/home/user/my-project'
    });
    console.log('Final:', finalStatus.content[0].text);
  }
}, 3000);

Summary

  • The asynchronous indexing workflow in Claude Context spans six phases: request validation, snapshot initialization, background job launch, file processing, progress reporting, and final state commit.

  • Progress tracking relies on a dual-layer system: an in-memory Map for instant updates and a throttled JSON snapshot (~/.context/mcp-codebase-snapshot.json) persisted every ~2 seconds.

  • Key source files implementing this workflow include packages/mcp/src/handlers.ts (orchestration), packages/mcp/src/snapshot.ts (persistence), and packages/core/src/context.ts (core indexing engine).

  • Status queries use the get_indexing_status tool, which reads the snapshot and returns human-readable states: indexing, indexed, indexfailed, or not indexed.

Frequently Asked Questions

How does Claude Context recover indexing progress after a crash?

The mcp-codebase-snapshot.json file in ~/.context/ serves as the recovery source. On restart, SnapshotManager loads this file into indexingCodebases: Map<string, number>, restoring the exact percentage and state. The handleGetIndexingStatus tool then reports this recovered progress to the UI.

What happens if indexing fails midway through a large codebase?

The error handler in handlers.ts (lines 83-92) calls setCodebaseIndexFailed with the error message and last known progress percentage. This state is immediately persisted to the snapshot. Users can query get_indexing_status to see the failure reason and partial progress, then retry with force: true if needed.

How frequently does the UI receive progress updates?

The VS Code extension polls get_indexing_status every few seconds. The underlying snapshot updates every ~2 seconds (throttled in handlers.ts lines 58-62), with in-memory updates occurring immediately on each batch completion. This creates a maximum ~2-second lag between actual progress and UI display.

Can multiple codebases be indexed simultaneously?

Yes. The indexingCodebases Map stores progress for each path independently, and startBackgroundIndexing spawns separate async flows per invocation. The snapshot tracks all active jobs, and handleGetIndexingStatus accepts a path parameter to query specific codebases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →