Multi-Pass Indexing Pipeline Architecture in Codebase-Memory-MCP: A 5-Stage Technical Breakdown
The Codebase-Memory-MCP engine implements a sequential five-pass indexing pipeline that transforms raw repositories into a searchable, language-aware knowledge base by progressively enriching file metadata with ASTs, semantic graphs, and vector embeddings.
The DeusData/codebase-memory-mcp repository provides a fast code intelligence engine that builds a persistent memory of your codebase. At its core lies a multi-pass indexing pipeline architecture designed to balance comprehensive semantic analysis with manageable memory consumption by staging data enrichment across five distinct phases.
The Five-Pass Pipeline Architecture
The pipeline processes code sequentially, where each pass depends on the artifacts of the previous stage. This design allows the system to maintain a cold index of raw files while gradually adding computationally expensive semantic layers only when needed.
Pass 1: Token and File Scan (Cold Index)
The initial pass creates a cold index of every file in the repository. According to the evaluation documentation in docs/EVALUATION_PLAN.md, this stage walks the repository tree, records file-type, size, and SHA-1 hash, and stores raw text blobs in a compact, deduplicated storage format. This lightweight snapshot enables incremental re-indexing; subsequent runs only trigger later passes for changed files.
Pass 2: Language-Specific Front-End Parsing
The second pass dispatches each file to language-specific parsers as described in pkg/pypi/README.md. For LSP-hybrid languages like Go, Rust, and TypeScript, the system utilizes Language Server Protocol (LSP) integrations and SSA (Static Single Assignment) forms. For simpler languages like Python, it employs native AST parsing. This stage emits syntax trees, symbol tables, and import/export edges that form the structural backbone of the index.
Pass 3: Semantic Enrichment and Cross-File Resolution
During the third pass, the pipeline resolves cross-file references and builds project-wide graphs. As implemented in docs/ARCHITECTURE.md, this stage constructs call-graphs and type-graphs that span multiple files and languages. It also computes semantic embeddings for symbols using the Nomic model, generating dense vector representations that capture code meaning beyond surface syntax.
Pass 4: Vector Index Construction (IVF-PQ)
The fourth pass transforms embeddings into a fast nearest-neighbor index. The scripts/extract_nomic_vectors.py script handles the serialization of int8 vectors from the Nomic model and inserts them into a disk-resident IVF-PQ (Inverted File + Product Quantization) index. This compression strategy reduces storage requirements while preserving millisecond-scale query performance for semantic similarity searches.
Pass 5: Final Index Consolidation
The final pass combines all artifacts into a portable project bundle. The index_repository implementation in pkg/pypi/src/codebase_memory_mcp/_cli.py orchestrates this stage, writing metadata (project name, version, timestamps) alongside the cold index, ASTs, semantic graphs, and vector index into a single .cbm file that the CLI can load instantly.
How the Passes Interact
The pipeline implements several optimization strategies to minimize redundant computation and memory usage.
Cold-Start vs. Warm-Start Optimization
The architecture distinguishes between cold-start and warm-start indexing. The first pass produces a cold snapshot that persists across runs. When re-indexing, the system compares SHA-1 hashes against this snapshot, triggering passes 2–5 only for modified files. This incremental approach keeps rebuild times under seconds even for large repositories.
Language-Hybrid Strategy (Full vs. Fast Mode)
Codebase-Memory-MCP supports two indexing modes controlled by the mode parameter in index_repository. Full mode executes all five passes for LSP-hybrid languages (Go, Rust, TypeScript), enabling deep semantic analysis and cross-file call graph resolution. Fast mode skips the heavy semantic enrichment (passes 3–4) for simpler languages like Python and Bash, storing only ASTs and symbol tables to reduce memory footprint by approximately 60%.
Just-In-Time Indexing
Rather than pre-computing every possible semantic relationship, the engine employs just-in-time indexing. When a query references a language or dependency that has not yet been fully processed, the pipeline launches the missing passes on-the-fly, materializing the required subgraphs without rebuilding the entire index.
Cross-Repository Edge Handling
After Pass 3, the semantic graph may contain cross-repo edges (e.g., CROSS_HTTP_CALLS linking microservices). These edges are stored in the index but remain dormant until the dependent project is also indexed. This deferred materialization strategy prevents unbounded index growth while maintaining the ability to trace dependencies across repository boundaries.
Implementation Details and Source Files
The multi-pass architecture is distributed across several key modules:
scripts/extract_nomic_vectors.py: Generates Nomic token-level embeddings and writes the raw vector blob used in Pass 4.pkg/pypi/src/codebase_memory_mcp/_cli.py: Contains the CLI entry point andindex_repositoryfunction that orchestrates pass execution.pkg/pypi/src/codebase_memory_mcp/__main__.py: Exposes thecodebase-memory-mcpcommand used by agents to trigger indexing.docs/ARCHITECTURE.md: Documents the language front-end abstraction and semantic enrichment layer.docs/EVALUATION_PLAN.md: Details the evaluation methodology and explains the sequential design choice over parallel processing to manage memory constraints.pkg/pypi/README.md: Describes the fast versus full indexing modes and language support matrix.
Practical Usage: Running the Pipeline
The following Python example demonstrates how to invoke the complete pipeline and query the resulting index:
from codebase_memory_mcp import index_repository
# Index a repository in "full" mode (all five passes)
project = index_repository(
path="/path/to/my/project",
language="python", # triggers language-specific front-end
mode="full", # runs all passes including semantic enrichment
batch_size=128,
)
# Query the built index for a symbol's callers (leverages Pass 3 call-graph)
callers = project.callers_of("my_module.my_function")
print("Callers:", callers)
# Perform semantic similarity search (leverages Pass 4 IVF-PQ index)
similar = project.semantic_search(
query="load configuration", top_k=5
)
for sym, score in similar:
print(f"{sym} (score {score:.2f})")
Summary
- The multi-pass indexing pipeline processes repositories in five sequential stages: file scanning, language parsing, semantic enrichment, vector indexing, and final consolidation.
- Cold-start optimization enables incremental updates by reusing Pass 1 artifacts and only reprocessing changed files.
- Language-hybrid strategy applies full semantic analysis to LSP-hybrid languages while using fast mode for simpler languages to conserve memory.
- IVF-PQ vector storage in Pass 4 provides efficient semantic search capabilities with minimal disk overhead.
- Cross-repo edge handling defers expensive inter-repository link resolution until dependencies are indexed.
Frequently Asked Questions
What triggers each pass in the multi-pass indexing pipeline?
The first pass triggers automatically when pointing the engine at a new repository. Subsequent passes execute based on the selected mode: fast mode stops after Pass 2, while full mode proceeds through Pass 5. In incremental updates, only files with changed SHA-1 hashes advance beyond Pass 1, with the system reusing cached artifacts for unchanged content.
How does the architecture handle memory constraints during indexing?
The sequential design deliberately avoids parallel processing to keep memory usage bounded. As noted in docs/EVALUATION_PLAN.md, materializing the full semantic graph for large repositories can consume gigabytes of RAM; by staging computation across passes and using int8 quantization for vectors in Pass 4, the system maintains a working set size suitable for development laptops while still supporting repositories with millions of lines of code.
What is the difference between "full" and "fast" indexing modes?
Full mode executes all five passes, enabling cross-file call graph resolution and semantic embedding search using the Nomic model. Fast mode truncates the pipeline after Pass 2, skipping the computationally expensive semantic enrichment and vector construction. Fast mode is recommended for CI/CD pipelines and resource-constrained environments, while full mode is required for "find similar code" and impact analysis queries.
How are cross-repository dependencies handled in the semantic graph?
During Pass 3, the system identifies potential cross-repository references (such as CROSS_HTTP_CALLS) but stores these as dangling edges rather than resolving them immediately. These edges remain in the graph structure but are only materialized when the target repository is also indexed, preventing index bloat while preserving the ability to perform cross-repo analysis when both projects are loaded into memory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →