How Codebase-Memory-MCP Indexes a Repository: A Complete Technical Guide
Codebase-Memory-MCP indexes a repository by first discovering files with Git-aware ignore rules, then running a deterministic pipeline of analysis passes that parse symbols, resolve references, and compute semantic embeddings, finally persisting everything into a SQLite graph database via a supervised worker process that isolates crashes from the main MCP server.
Codebase-Memory-MCP (CBM) transforms your repository into a fully queryable, language-aware graph stored in SQLite. This indexing process enables semantic code search, cross-reference navigation, and clone detection without requiring external services. Understanding how Codebase-Memory-MCP indexes a repository reveals why it can handle large codebases robustly while providing detailed symbol relationships and vector-based similarity metrics.
Stage 1: Discovery and Preprocessing
The indexing process begins in src/discover/discover.c, where the discovery module builds the initial file inventory while respecting repository boundaries and ignore rules.
Walking the Filesystem
The function cbm_discover_repo recursively scans the repository root, following symbolic links but refusing to traverse outside the repository boundary. This ensures that symlinks to system directories or parent folders do not pollute the index.
Git-Aware and Custom Ignore Rules
CBM implements sophisticated filtering through cbm_gitignore_load in src/discover/gitignore.c. The system parses .gitignore files throughout the tree and applies pattern matching to filter the file list. Additionally, users can define a .cbmignore file in the repository root (documented in docs/cbmignore.md) which uses the same parsing logic but provides CBM-specific exclusions without modifying Git configuration.
Language Detection and File Filtering
Each discovered file undergoes classification via cbm_language_from_path in src/discover/language.c. This function maps file extensions to language enums, using file content heuristics for ambiguous cases. Files exceeding CBM_MAX_FILE_SIZE (defined in src/foundation/limits.h) are automatically skipped to prevent out-of-memory conditions during parsing.
The discovery stage outputs a list of source objects containing absolute paths, language classifications, and indexing flags.
Stage 2: The Pipeline Pass Architecture
The core indexing engine lives in src/pipeline/pipeline.c, which orchestrates a series of deterministic passes. Each pass processes all source objects in order, allowing downstream passes to depend on symbols created by earlier ones.
Core Parsing and Symbol Extraction
The pass_definitions.c module parses source files using Tree-Sitter grammars to create Definition nodes for functions, classes, and constants. Subsequently, pass_calls.c traverses the ASTs to emit CALLS edges linking call sites to their targets, building the backbone of the symbol graph.
For C/C++ projects, pass_compile_commands.c loads compile_commands.json when present to resolve include paths and improve header accuracy. The pass_pkgmap.c pass builds package-to-module mappings to resolve import statements correctly.
Semantic Analysis and Vector Embeddings
CBM generates semantic relationships without external APIs. The pass_semantic.c module feeds source tokens into the bundled Nomic-embed-code model, producing 768-dimensional vectors and creating SEMANTICALLY_RELATED edges. This enables natural language queries like "find code related to authentication" even when the exact keyword does not appear.
For code deduplication, pass_similarity.c computes SimHash signatures for each function, grouping near-duplicates and creating NEAR_DUPLICATE edges using the S2 metric.
Cross-Repository and Incremental Indexing
When enabled, pass_cross_repo.c matches symbols across distinct project databases, inserting CROSS_CALLS and CROSS_REFS edges to support monorepo analysis. For incremental updates, pass_gitdiff.c detects newly added or removed files, marking nodes for insertion or deletion without rebuilding the entire graph.
Additional passes include pass_envscan.c (detecting environment variable usage), pass_githistory.c (attaching Git blame information to nodes), pass_route_nodes.c (extracting HTTP route declarations), and pass_enrichment.c (computing cyclomatic complexity and line counts).
All passes write to the store layer (src/store/sqlite_writer.c), which translates nodes and edges into the nodes, edges, and metadata tables using bulk inserts for performance.
Stage 3: Supervision and Worker Isolation
To ensure the MCP server remains responsive during indexing, the heavy lifting runs in a separate supervised process managed by src/mcp/index_supervisor.c.
Process Isolation and Crash Protection
The cbm_index_spawn_worker function creates a child process via cbm_subprocess_run. If the worker crashes, hangs, or exhausts memory, the supervisor detects the failure and terminates the child without affecting the main server. The supervisor resolves the correct binary path using cbm_http_server_resolve_binary_path to ensure plugin compatibility.
Timeout Handling and Result Collection
The supervisor implements quiet-timeout detection through the worker_quiet_timeout_ms function. If a worker stops emitting log lines for 15 minutes (configurable via CBM_INDEX_WORKER_TIMEOUT_S), the supervisor kills the process and reports it as a hang.
The worker writes its JSON response to a temporary file (.worker-<pid>.response), which the supervisor reads via slurp_file after the child exits. Environment variables such as CBM_INDEX_SINGLE_THREAD, CBM_INDEX_MARKER_FILE, and CBM_INDEX_QUARANTINE_FILE temporarily control worker behavior during execution.
The worker binary is the same executable invoked with the --index-worker flag (handled in src/cli/cli.c).
Practical Usage Examples
Indexing from the Command Line
Index the current directory with default settings:
cbm index .
Force single-threaded execution for debugging or resource-constrained environments:
CBM_INDEX_SINGLE_THREAD=1 cbm index /path/to/project
Programmatic API Usage
Control indexing programmatically using the C API:
#include "mcp/mcp.h"
const char *opts = "{\"root\":\"/my/repo\",\"skipTests\":true}";
cbm_index_worker_result_t result;
if (cbm_index_spawn_worker(opts, false, NULL, NULL, &result) == 0) {
if (result.outcome == CBM_PROC_CLEAN) {
printf("Index succeeded, response: %s\n", result.response);
} else {
fprintf(stderr, "Worker failed (outcome=%d, signal=%d)\n",
result.outcome, result.term_signal);
}
cbm_index_worker_result_free(&result);
}
Querying the Generated Graph
Query call relationships directly via SQL:
-- Find all functions that call a function named `fetch_user`
SELECT src.name, dst.name
FROM edges e
JOIN nodes src ON e.src_id = src.id
JOIN nodes dst ON e.dst_id = dst.id
WHERE e.kind = 'CALLS' AND dst.name = 'fetch_user';
Perform semantic search using the built-in tool:
cbm search --semantic "publish"
Summary
- Codebase-Memory-MCP indexes a repository through three stages: filesystem discovery, multi-pass pipeline analysis, and supervised worker execution.
- The discovery phase (
src/discover/discover.c) applies Git-aware and custom ignore rules while classifying files by language. - Pipeline passes (
src/pipeline/pipeline.c) parse symbols, resolve calls, compute 768-dimensional semantic vectors, and detect near-duplicate code. - Worker isolation (
src/mcp/index_supervisor.c) prevents crashes and hangs from affecting the main MCP server through process supervision and timeout detection. - All data persists to a local SQLite database (
src/store/sqlite_writer.c) containing nodes, edges, and metadata tables for LSP and GraphQL queries.
Frequently Asked Questions
How does Codebase-Memory-MCP handle large repositories?
CBM handles large repositories through the pass_parallel.c module, which executes passes in parallel while preserving deterministic ordering. Additionally, files exceeding CBM_MAX_FILE_SIZE (defined in src/foundation/limits.h) are skipped during discovery to prevent memory exhaustion, and the supervised worker architecture ensures that resource constraints in the indexer do not crash the main MCP server.
What file types does the indexer support?
The indexer supports any language with a Tree-Sitter grammar available to the system. Language detection occurs in src/discover/language.c via cbm_language_from_path, which maps file extensions to language enums. For ambiguous files (such as those without extensions), the system examines a snippet of the file content to determine the appropriate parser.
How does the indexer prevent crashes from affecting the MCP server?
The src/mcp/index_supervisor.c module spawns indexing work in a separate child process via cbm_index_spawn_worker. If the worker crashes, hangs, or encounters an out-of-memory condition, the supervisor detects the failure through process monitoring and quiet-timeout detection (worker_quiet_timeout_ms), terminating the worker without impacting the main server process. Results are communicated through temporary files rather than shared memory.
Can I customize which files are included in the index?
Yes. In addition to standard .gitignore rules parsed by cbm_gitignore_load in src/discover/gitignore.c, you can create a .cbmignore file in the repository root (documented in docs/cbmignore.md). This file uses the same pattern syntax as Git ignore files but provides CBM-specific exclusions. You can also pass options via the C API or environment variables like CBM_INDEX_SINGLE_THREAD to control indexing behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →