codebase-memory-mcp Indexing Performance Benchmarks for Large Codebases: Linux Kernel in 3 Minutes

codebase-memory-mcp indexes the 28-million-line Linux kernel in approximately 3 minutes using a RAM-first pipeline, LZ4 compression, and in-memory SQLite, delivering sub-millisecond query latencies on multi-million-node graphs.

The codebase-memory-mcp project from DeusData redefines what is possible for codebase indexing performance. Designed specifically for massive repositories, this tool transforms millions of lines of code into navigable graph structures faster than traditional static analysis tools. This article explores the architectural decisions and benchmark results that enable codebase-memory-mcp indexing performance benchmarks on large codebases to achieve linear scaling with repository size.

Architectural Optimizations Driving Indexing Speed

The indexer achieves its speed through a combination of low-level systems programming choices that minimize I/O bottlenecks and maximize CPU utilization. Every design decision prioritizes keeping data in RAM and avoiding disk access during the parsing phase.

RAM-First Pipeline with LZ4 Compression

The tool implements a RAM-first pipeline that reads files into LZ4-compressed blocks before processing. In src/main.c, the orchestration logic streams source files through LZ4 HC (high-compression) algorithms, storing them in memory until the final persistence step. This eliminates random disk reads during parsing, tokenization, and graph construction. Memory is released only after the final compressed dump is written to ~/.cache/codebase-memory-mcp/graph.db.zst.

Rather than writing intermediate results to disk, the system uses an in-memory SQLite instance with FTS5 full-text search capabilities. The src/store/sqlite.c file implements the graph storage layer, utilizing a custom cbm_camel_split tokenizer that understands camelCase and snake_case naming conventions. This provides O(1) symbol lookups without persisting temporary data during the indexing process.

Vendored Tree-Sitter and Hybrid LSP Layer

The src/pipeline/extract_usages.c file contains the core Tree-Sitter extraction routines, while src/pipeline/hybrid_lsp.c implements a lightweight C-based type resolver. With 158 vendored Tree-Sitter grammars compiled into the static binary, the tool parses files without dynamic loading overhead. The Hybrid LSP layer resolves imports, generics, inheritance, and standard-library types for 11 languages—including Python, TypeScript, Go, Rust, and C/C++—without spawning separate language-server processes.

Parallel Processing and Aho-Corasick Matching

Parallelism is controlled via the CBM_WORKERS environment variable, defaulting to the system's CPU core count. The src/pipeline/discover.c file handles file system walking with .gitignore and .cbmignore filtering. Additionally, an Aho-Corasick pattern matcher enables single-pass symbol scanning across the entire repository, making search operations 10× faster than naive grep implementations.

Real-World Benchmark Results

The project publishes specific performance metrics that demonstrate linear scaling with repository complexity. These figures represent single-machine execution on standard development hardware.

  • Linux kernel full index: 3 minutes (28 M LOC, 75 K files → 4.81 M nodes, 7.72 M edges)
  • Linux kernel fast index: 1 minute 12 seconds (1.88 M nodes, no hybrid LSP resolution)
  • Django full index: ~6 seconds (49 K nodes, 196 K edges)
  • Cypher relationship query: < 1 ms
  • Regex name search: < 10 ms (SQL-LIKE pre-filter)
  • Dead-code detection scan: ~150 ms (full-graph traversal)
  • Call-path trace (depth 5): < 10 ms (BFS traversal)

The benchmarks prove that while indexing time scales linearly with source tree size, query latency remains sub-millisecond because the graph resides entirely in memory or loads from a compressed SQLite snapshot on demand.

How the Indexing Pipeline Works

The source code in the src/ directory implements a multi-pass pipeline that reads each file exactly once:


src/
 ├─ main.c                → Entry point, CLI & MCP server
 ├─ pipeline/            → Multi-pass indexing:
 │   ├─ discover/        → File discovery, .gitignore & .cbmignore handling
 │   ├─ extract_usages.c → AST extraction via Tree-Sitter
 │   └─ hybrid_lsp/      → Type-aware resolution (Python, TS, …)
 ├─ store/               → SQLite graph storage, Louvain clustering
 └─ ui/                  → Optional 3D visualization server

The discover stage walks the filesystem once, applying ignore rules. The extract_usages stage parses each file with its vendored grammar, populating raw symbols and call sites. The hybrid LSP pass enriches those edges with type information, yielding a fully-linked call graph. All results are flushed into an in-memory SQLite DB, which is then persisted to ~/.cache/codebase-memory-mcp/graph.db.zst (a ZSTD-compressed SQLite snapshot).

Because each stage reads only once, the total I/O cost is bounded by the file size, not by the number of passes.

Reproducing the Benchmarks

You can validate these performance claims using the following commands against large repositories like the Linux kernel.

Index a Massive Repository


# Download the binary

curl -fsSL https://raw.githubusercontent.com/DeusData/codebase-memory-mcp/main/install.sh | bash

# Run the indexer

codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/linux"}' |& tee index.log

The JSON output contains timing data:

{
  "status":"indexed",
  "repo":"linux",
  "node_count":4812000,
  "edge_count":7720000,
  "duration_ms":180000
}

Extract the duration programmatically:

duration=$(codebase-memory-mcp cli index_repository '{"repo_path":"$HOME/linux"}' | jq -r '.duration_ms')
echo "Indexing took $((duration/1000)) seconds"

Measure Fast Index Mode


# Limit workers for reproducibility

export CBM_WORKERS=4
codebase-memory-mcp cli index_repository '{"repo_path":"$HOME/linux","fast":true}'

The fast flag stops after the first pass (no hybrid LSP resolution), typically completing in 72 seconds for the Linux kernel.

Execute Structural Queries


# Find all functions ending with "_init"

codebase-memory-mcp cli search_graph '{"label":"Function","name_pattern":"_init$"}' \
    | jq -r '.results[].qualified_name'

This query completes in under 10 milliseconds.

Trace Call Paths

codebase-memory-mcp cli trace_path '{"function_name":"schedule","direction":"outbound","depth":5}' \
    | jq '.results[].name'

This yields the call chain up to five hops in approximately 10 milliseconds.

Verify Persisted Graph Size

du -h ~/.cache/codebase-memory-mcp/graph.db.zst

The ZSTD-compressed snapshot typically achieves a 10-15× size reduction compared to raw SQLite files.

Repeatable Benchmark Script

#!/usr/bin/env bash
set -euo pipefail

REPO=$1
RUNS=5

for i in $(seq 1 $RUNS); do
    echo "Run $i..."
    time codebase-memory-mcp cli index_repository "{\"repo_path\":\"$REPO\"}" >/dev/null
done

Summary

  • codebase-memory-mcp achieves minutes-scale indexing on multi-million-line codebases by keeping the entire pipeline in RAM and using LZ4 compression.
  • The combination of vendored Tree-Sitter grammars, a C-based Hybrid LSP layer, and in-memory SQLite eliminates external dependencies and disk I/O bottlenecks.
  • Benchmarks demonstrate 3-minute full indexing of the Linux kernel and sub-millisecond query latencies, making it suitable for real-time AI-assisted coding agents.
  • The static binary deployment (no Docker, no dynamic libraries) ensures consistent performance across environments.

Frequently Asked Questions

How does codebase-memory-mcp achieve such fast indexing on large repositories?

The tool uses a RAM-first architecture where src/main.c orchestrates a pipeline that keeps all parsing, tokenization, and graph construction in memory. By combining LZ4 HC compression to reduce I/O volume with an in-memory SQLite database using the cbm_camel_split tokenizer, it eliminates disk access during processing. The Aho-Corasick pattern matcher and parallel workers controlled by CBM_WORKERS further accelerate the process.

What is the difference between full index and fast index modes?

Full index mode, implemented across src/pipeline/extract_usages.c and src/pipeline/hybrid_lsp.c, performs complete AST extraction followed by type-aware resolution to generate CALLS, IMPORTS, and IMPLEMENTED_BY edges. Fast index mode skips the hybrid LSP pass, running only the initial file discovery and Tree-Sitter extraction, which reduces the Linux kernel indexing time from 3 minutes to approximately 72 seconds but produces fewer semantic edges.

How much disk space does the indexed graph consume?

The final graph is stored at ~/.cache/codebase-memory-mcp/graph.db.zst using ZSTD level-3 compression. For the Linux kernel's 4.81 million nodes and 7.72 million edges, the compressed file occupies only a few hundred megabytes—a 10-15× reduction compared to uncompressed SQLite storage.

Can I limit CPU usage during indexing?

Yes. While the indexer auto-detects CPU cores, you can override the parallel worker count using the CBM_WORKERS environment variable. Setting CBM_WORKERS=4 restricts the process to four concurrent threads, allowing you to reserve CPU capacity for other development tasks during indexing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →