# How Codebase-Memory-MCP Indexes a Repository: A Complete Technical Guide

> Learn how Codebase-Memory-MCP indexes a repository. Discover its Git-aware file discovery, deterministic analysis pipeline, and SQLite graph database persistence for efficient code understanding.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: deep-dive
- Published: 2026-07-04

---

**Codebase-Memory-MCP indexes a repository by first discovering files with Git-aware ignore rules, then running a deterministic pipeline of analysis passes that parse symbols, resolve references, and compute semantic embeddings, finally persisting everything into a SQLite graph database via a supervised worker process that isolates crashes from the main MCP server.**

Codebase-Memory-MCP (CBM) transforms your repository into a fully queryable, language-aware graph stored in SQLite. This indexing process enables semantic code search, cross-reference navigation, and clone detection without requiring external services. Understanding how Codebase-Memory-MCP indexes a repository reveals why it can handle large codebases robustly while providing detailed symbol relationships and vector-based similarity metrics.

## Stage 1: Discovery and Preprocessing

The indexing process begins in [`src/discover/discover.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/discover.c), where the **discovery module** builds the initial file inventory while respecting repository boundaries and ignore rules.

### Walking the Filesystem

The function `cbm_discover_repo` recursively scans the repository root, following symbolic links but refusing to traverse outside the repository boundary. This ensures that symlinks to system directories or parent folders do not pollute the index.

### Git-Aware and Custom Ignore Rules

CBM implements sophisticated filtering through `cbm_gitignore_load` in [`src/discover/gitignore.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/gitignore.c). The system parses `.gitignore` files throughout the tree and applies pattern matching to filter the file list. Additionally, users can define a `.cbmignore` file in the repository root (documented in [`docs/cbmignore.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/cbmignore.md)) which uses the same parsing logic but provides CBM-specific exclusions without modifying Git configuration.

### Language Detection and File Filtering

Each discovered file undergoes classification via `cbm_language_from_path` in [`src/discover/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/language.c). This function maps file extensions to language enums, using file content heuristics for ambiguous cases. Files exceeding `CBM_MAX_FILE_SIZE` (defined in [`src/foundation/limits.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/foundation/limits.h)) are automatically skipped to prevent out-of-memory conditions during parsing.

The discovery stage outputs a list of **source objects** containing absolute paths, language classifications, and indexing flags.

## Stage 2: The Pipeline Pass Architecture

The core indexing engine lives in [`src/pipeline/pipeline.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pipeline.c), which orchestrates a series of deterministic passes. Each pass processes all source objects in order, allowing downstream passes to depend on symbols created by earlier ones.

### Core Parsing and Symbol Extraction

The [`pass_definitions.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_definitions.c) module parses source files using Tree-Sitter grammars to create **Definition** nodes for functions, classes, and constants. Subsequently, [`pass_calls.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_calls.c) traverses the ASTs to emit **CALLS** edges linking call sites to their targets, building the backbone of the symbol graph.

For C/C++ projects, [`pass_compile_commands.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_compile_commands.c) loads [`compile_commands.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/compile_commands.json) when present to resolve include paths and improve header accuracy. The [`pass_pkgmap.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_pkgmap.c) pass builds package-to-module mappings to resolve import statements correctly.

### Semantic Analysis and Vector Embeddings

CBM generates semantic relationships without external APIs. The [`pass_semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_semantic.c) module feeds source tokens into the bundled Nomic-embed-code model, producing 768-dimensional vectors and creating **SEMANTICALLY_RELATED** edges. This enables natural language queries like "find code related to authentication" even when the exact keyword does not appear.

For code deduplication, [`pass_similarity.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_similarity.c) computes SimHash signatures for each function, grouping near-duplicates and creating **NEAR_DUPLICATE** edges using the S2 metric.

### Cross-Repository and Incremental Indexing

When enabled, [`pass_cross_repo.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_cross_repo.c) matches symbols across distinct project databases, inserting **CROSS_CALLS** and **CROSS_REFS** edges to support monorepo analysis. For incremental updates, [`pass_gitdiff.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_gitdiff.c) detects newly added or removed files, marking nodes for insertion or deletion without rebuilding the entire graph.

Additional passes include [`pass_envscan.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_envscan.c) (detecting environment variable usage), [`pass_githistory.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_githistory.c) (attaching Git blame information to nodes), [`pass_route_nodes.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_route_nodes.c) (extracting HTTP route declarations), and [`pass_enrichment.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_enrichment.c) (computing cyclomatic complexity and line counts).

All passes write to the **store layer** ([`src/store/sqlite_writer.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/sqlite_writer.c)), which translates nodes and edges into the `nodes`, `edges`, and `metadata` tables using bulk inserts for performance.

## Stage 3: Supervision and Worker Isolation

To ensure the MCP server remains responsive during indexing, the heavy lifting runs in a separate supervised process managed by [`src/mcp/index_supervisor.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/index_supervisor.c).

### Process Isolation and Crash Protection

The `cbm_index_spawn_worker` function creates a child process via `cbm_subprocess_run`. If the worker crashes, hangs, or exhausts memory, the supervisor detects the failure and terminates the child without affecting the main server. The supervisor resolves the correct binary path using `cbm_http_server_resolve_binary_path` to ensure plugin compatibility.

### Timeout Handling and Result Collection

The supervisor implements **quiet-timeout detection** through the `worker_quiet_timeout_ms` function. If a worker stops emitting log lines for 15 minutes (configurable via `CBM_INDEX_WORKER_TIMEOUT_S`), the supervisor kills the process and reports it as a hang.

The worker writes its JSON response to a temporary file (`.worker-<pid>.response`), which the supervisor reads via `slurp_file` after the child exits. Environment variables such as `CBM_INDEX_SINGLE_THREAD`, `CBM_INDEX_MARKER_FILE`, and `CBM_INDEX_QUARANTINE_FILE` temporarily control worker behavior during execution.

The worker binary is the same executable invoked with the `--index-worker` flag (handled in [`src/cli/cli.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/cli/cli.c)).

## Practical Usage Examples

### Indexing from the Command Line

Index the current directory with default settings:

```bash
cbm index .

```

Force single-threaded execution for debugging or resource-constrained environments:

```bash
CBM_INDEX_SINGLE_THREAD=1 cbm index /path/to/project

```

### Programmatic API Usage

Control indexing programmatically using the C API:

```c
#include "mcp/mcp.h"

const char *opts = "{\"root\":\"/my/repo\",\"skipTests\":true}";

cbm_index_worker_result_t result;
if (cbm_index_spawn_worker(opts, false, NULL, NULL, &result) == 0) {
    if (result.outcome == CBM_PROC_CLEAN) {
        printf("Index succeeded, response: %s\n", result.response);
    } else {
        fprintf(stderr, "Worker failed (outcome=%d, signal=%d)\n",
                result.outcome, result.term_signal);
    }
    cbm_index_worker_result_free(&result);
}

```

### Querying the Generated Graph

Query call relationships directly via SQL:

```sql
-- Find all functions that call a function named `fetch_user`
SELECT src.name, dst.name
FROM edges e
JOIN nodes src ON e.src_id = src.id
JOIN nodes dst ON e.dst_id = dst.id
WHERE e.kind = 'CALLS' AND dst.name = 'fetch_user';

```

Perform semantic search using the built-in tool:

```bash
cbm search --semantic "publish"

```

## Summary

- **Codebase-Memory-MCP indexes a repository** through three stages: filesystem discovery, multi-pass pipeline analysis, and supervised worker execution.
- The **discovery phase** ([`src/discover/discover.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/discover.c)) applies Git-aware and custom ignore rules while classifying files by language.
- **Pipeline passes** ([`src/pipeline/pipeline.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pipeline.c)) parse symbols, resolve calls, compute 768-dimensional semantic vectors, and detect near-duplicate code.
- **Worker isolation** ([`src/mcp/index_supervisor.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/index_supervisor.c)) prevents crashes and hangs from affecting the main MCP server through process supervision and timeout detection.
- All data persists to a local SQLite database ([`src/store/sqlite_writer.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/sqlite_writer.c)) containing nodes, edges, and metadata tables for LSP and GraphQL queries.

## Frequently Asked Questions

### How does Codebase-Memory-MCP handle large repositories?

CBM handles large repositories through the [`pass_parallel.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/pass_parallel.c) module, which executes passes in parallel while preserving deterministic ordering. Additionally, files exceeding `CBM_MAX_FILE_SIZE` (defined in [`src/foundation/limits.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/foundation/limits.h)) are skipped during discovery to prevent memory exhaustion, and the supervised worker architecture ensures that resource constraints in the indexer do not crash the main MCP server.

### What file types does the indexer support?

The indexer supports any language with a Tree-Sitter grammar available to the system. Language detection occurs in [`src/discover/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/language.c) via `cbm_language_from_path`, which maps file extensions to language enums. For ambiguous files (such as those without extensions), the system examines a snippet of the file content to determine the appropriate parser.

### How does the indexer prevent crashes from affecting the MCP server?

The [`src/mcp/index_supervisor.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/index_supervisor.c) module spawns indexing work in a separate child process via `cbm_index_spawn_worker`. If the worker crashes, hangs, or encounters an out-of-memory condition, the supervisor detects the failure through process monitoring and quiet-timeout detection (`worker_quiet_timeout_ms`), terminating the worker without impacting the main server process. Results are communicated through temporary files rather than shared memory.

### Can I customize which files are included in the index?

Yes. In addition to standard `.gitignore` rules parsed by `cbm_gitignore_load` in [`src/discover/gitignore.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/gitignore.c), you can create a `.cbmignore` file in the repository root (documented in [`docs/cbmignore.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/cbmignore.md)). This file uses the same pattern syntax as Git ignore files but provides CBM-specific exclusions. You can also pass options via the C API or environment variables like `CBM_INDEX_SINGLE_THREAD` to control indexing behavior.