# Multi-Pass Indexing Pipeline Architecture in Codebase-Memory-MCP: A 5-Stage Technical Breakdown

> Explore the five-stage multi-pass indexing pipeline architecture in Codebase-Memory-MCP. Learn how it transforms code into a searchable, language-aware knowledge base with ASTs, graphs, and embeddings.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: architecture
- Published: 2026-07-10

---

**The Codebase-Memory-MCP engine implements a sequential five-pass indexing pipeline that transforms raw repositories into a searchable, language-aware knowledge base by progressively enriching file metadata with ASTs, semantic graphs, and vector embeddings.**

The DeusData/codebase-memory-mcp repository provides a fast code intelligence engine that builds a persistent memory of your codebase. At its core lies a **multi-pass indexing pipeline architecture** designed to balance comprehensive semantic analysis with manageable memory consumption by staging data enrichment across five distinct phases.

## The Five-Pass Pipeline Architecture

The pipeline processes code sequentially, where each pass depends on the artifacts of the previous stage. This design allows the system to maintain a **cold index** of raw files while gradually adding computationally expensive semantic layers only when needed.

### Pass 1: Token and File Scan (Cold Index)

The initial pass creates a *cold* index of every file in the repository. According to the evaluation documentation in [`docs/EVALUATION_PLAN.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/EVALUATION_PLAN.md), this stage walks the repository tree, records file-type, size, and SHA-1 hash, and stores raw text blobs in a compact, deduplicated storage format. This lightweight snapshot enables incremental re-indexing; subsequent runs only trigger later passes for changed files.

### Pass 2: Language-Specific Front-End Parsing

The second pass dispatches each file to language-specific parsers as described in [`pkg/pypi/README.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pkg/pypi/README.md). For *LSP-hybrid* languages like Go, Rust, and TypeScript, the system utilizes Language Server Protocol (LSP) integrations and SSA (Static Single Assignment) forms. For simpler languages like Python, it employs native AST parsing. This stage emits syntax trees, symbol tables, and import/export edges that form the structural backbone of the index.

### Pass 3: Semantic Enrichment and Cross-File Resolution

During the third pass, the pipeline resolves cross-file references and builds project-wide graphs. As implemented in [`docs/ARCHITECTURE.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/ARCHITECTURE.md), this stage constructs **call-graphs** and **type-graphs** that span multiple files and languages. It also computes **semantic embeddings** for symbols using the Nomic model, generating dense vector representations that capture code meaning beyond surface syntax.

### Pass 4: Vector Index Construction (IVF-PQ)

The fourth pass transforms embeddings into a fast nearest-neighbor index. The [`scripts/extract_nomic_vectors.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/extract_nomic_vectors.py) script handles the serialization of **int8 vectors** from the Nomic model and inserts them into a disk-resident **IVF-PQ** (Inverted File + Product Quantization) index. This compression strategy reduces storage requirements while preserving millisecond-scale query performance for semantic similarity searches.

### Pass 5: Final Index Consolidation

The final pass combines all artifacts into a portable project bundle. The `index_repository` implementation in [`pkg/pypi/src/codebase_memory_mcp/_cli.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pkg/pypi/src/codebase_memory_mcp/_cli.py) orchestrates this stage, writing metadata (project name, version, timestamps) alongside the cold index, ASTs, semantic graphs, and vector index into a single `.cbm` file that the CLI can load instantly.

## How the Passes Interact

The pipeline implements several optimization strategies to minimize redundant computation and memory usage.

### Cold-Start vs. Warm-Start Optimization

The architecture distinguishes between **cold-start** and **warm-start** indexing. The first pass produces a cold snapshot that persists across runs. When re-indexing, the system compares SHA-1 hashes against this snapshot, triggering passes 2–5 only for modified files. This incremental approach keeps rebuild times under seconds even for large repositories.

### Language-Hybrid Strategy (Full vs. Fast Mode)

Codebase-Memory-MCP supports two indexing modes controlled by the `mode` parameter in `index_repository`. **Full mode** executes all five passes for LSP-hybrid languages (Go, Rust, TypeScript), enabling deep semantic analysis and cross-file call graph resolution. **Fast mode** skips the heavy semantic enrichment (passes 3–4) for simpler languages like Python and Bash, storing only ASTs and symbol tables to reduce memory footprint by approximately 60%.

### Just-In-Time Indexing

Rather than pre-computing every possible semantic relationship, the engine employs **just-in-time indexing**. When a query references a language or dependency that has not yet been fully processed, the pipeline launches the missing passes on-the-fly, materializing the required subgraphs without rebuilding the entire index.

### Cross-Repository Edge Handling

After Pass 3, the semantic graph may contain **cross-repo edges** (e.g., `CROSS_HTTP_CALLS` linking microservices). These edges are stored in the index but remain dormant until the dependent project is also indexed. This deferred materialization strategy prevents unbounded index growth while maintaining the ability to trace dependencies across repository boundaries.

## Implementation Details and Source Files

The multi-pass architecture is distributed across several key modules:

- **[`scripts/extract_nomic_vectors.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/extract_nomic_vectors.py)**: Generates Nomic token-level embeddings and writes the raw vector blob used in Pass 4.
- **[`pkg/pypi/src/codebase_memory_mcp/_cli.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pkg/pypi/src/codebase_memory_mcp/_cli.py)**: Contains the CLI entry point and `index_repository` function that orchestrates pass execution.
- **[`pkg/pypi/src/codebase_memory_mcp/__main__.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pkg/pypi/src/codebase_memory_mcp/__main__.py)**: Exposes the `codebase-memory-mcp` command used by agents to trigger indexing.
- **[`docs/ARCHITECTURE.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/ARCHITECTURE.md)**: Documents the language front-end abstraction and semantic enrichment layer.
- **[`docs/EVALUATION_PLAN.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/EVALUATION_PLAN.md)**: Details the evaluation methodology and explains the sequential design choice over parallel processing to manage memory constraints.
- **[`pkg/pypi/README.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pkg/pypi/README.md)**: Describes the fast versus full indexing modes and language support matrix.

## Practical Usage: Running the Pipeline

The following Python example demonstrates how to invoke the complete pipeline and query the resulting index:

```python
from codebase_memory_mcp import index_repository

# Index a repository in "full" mode (all five passes)

project = index_repository(
    path="/path/to/my/project",
    language="python",         # triggers language-specific front-end

    mode="full",              # runs all passes including semantic enrichment

    batch_size=128,
)

# Query the built index for a symbol's callers (leverages Pass 3 call-graph)

callers = project.callers_of("my_module.my_function")
print("Callers:", callers)

# Perform semantic similarity search (leverages Pass 4 IVF-PQ index)

similar = project.semantic_search(
    query="load configuration", top_k=5
)
for sym, score in similar:
    print(f"{sym} (score {score:.2f})")

```

## Summary

- **The multi-pass indexing pipeline** processes repositories in five sequential stages: file scanning, language parsing, semantic enrichment, vector indexing, and final consolidation.
- **Cold-start optimization** enables incremental updates by reusing Pass 1 artifacts and only reprocessing changed files.
- **Language-hybrid strategy** applies full semantic analysis to LSP-hybrid languages while using fast mode for simpler languages to conserve memory.
- **IVF-PQ vector storage** in Pass 4 provides efficient semantic search capabilities with minimal disk overhead.
- **Cross-repo edge handling** defers expensive inter-repository link resolution until dependencies are indexed.

## Frequently Asked Questions

### What triggers each pass in the multi-pass indexing pipeline?

The first pass triggers automatically when pointing the engine at a new repository. Subsequent passes execute based on the selected mode: **fast mode** stops after Pass 2, while **full mode** proceeds through Pass 5. In incremental updates, only files with changed SHA-1 hashes advance beyond Pass 1, with the system reusing cached artifacts for unchanged content.

### How does the architecture handle memory constraints during indexing?

The sequential design deliberately avoids parallel processing to keep memory usage bounded. As noted in [`docs/EVALUATION_PLAN.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/EVALUATION_PLAN.md), materializing the full semantic graph for large repositories can consume gigabytes of RAM; by staging computation across passes and using **int8 quantization** for vectors in Pass 4, the system maintains a working set size suitable for development laptops while still supporting repositories with millions of lines of code.

### What is the difference between "full" and "fast" indexing modes?

**Full mode** executes all five passes, enabling cross-file call graph resolution and semantic embedding search using the Nomic model. **Fast mode** truncates the pipeline after Pass 2, skipping the computationally expensive semantic enrichment and vector construction. Fast mode is recommended for CI/CD pipelines and resource-constrained environments, while full mode is required for "find similar code" and impact analysis queries.

### How are cross-repository dependencies handled in the semantic graph?

During Pass 3, the system identifies potential cross-repository references (such as `CROSS_HTTP_CALLS`) but stores these as **dangling edges** rather than resolving them immediately. These edges remain in the graph structure but are only materialized when the target repository is also indexed, preventing index bloat while preserving the ability to perform cross-repo analysis when both projects are loaded into memory.