# Semantic Search vs BM25 Full‑Text Search in codebase‑memory‑mcp: Architecture and Implementation

> Explore semantic search vs BM25 full-text search in codebase-memory-mcp. Learn how MCP integrates both for flexible and precise data retrieval.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: architecture
- Published: 2026-07-06

---

**codebase‑memory‑mcp implements two complementary search backends—dense vector semantic search for conceptual similarity and SQLite FTS5 BM25 for exact lexical matching—allowing MCP clients to choose the appropriate retrieval strategy based on query intent.**

The `codebase‑memory‑mcp` repository provides a graph‑centric memory layer for codebases that exposes both semantic and lexical search capabilities through a unified Model Context Protocol (MCP) interface. While both mechanisms serve the same high‑level goal of locating relevant code, they operate on fundamentally different data models: semantic search leverages high‑dimensional vector embeddings and Random Indexing, whereas BM25 relies on SQLite’s FTS5 inverted index with a custom camelCase‑aware tokenizer.

## Core Data Models: Embeddings vs. Inverted Indices

### Semantic Search Vectors

The semantic engine stores a dense vector for every distinct token and identifier in a `cbm_sem_corpus_t` structure. These vectors are constructed using a hybrid of **TF‑IDF**, **Random Indexing**, and pretrained embeddings (Nomic), augmented with code‑specific signals such as API signatures and type information. The corpus initialization and vector generation logic reside in [[`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c#L31-L50).

During the pipeline phase, the system emits `SEMANTICALLY_RELATED` edges by comparing these vectors using an 11‑signal weighted similarity model (`CBM_SEM_W_*` constants). This process is implemented in [[`src/pipeline/pass_semantic_edges.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pass_semantic_edges.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pass_semantic_edges.c#L2-L17).

### BM25 Full‑Text Index

In contrast, the BM25 implementation relies on SQLite’s **FTS5** virtual table named `nodes_fts`. The index is populated by a custom tokenizer called `cbm_camel_split`, which is registered in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.comDeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L1588-L1660). This tokenizer splits identifiers on camelCase and snake_case boundaries, ensuring that `openFile` and `open_file` both generate searchable tokens `open` and `file`. The underlying FTS5 engine is vendored in [[`vendored/sqlite3/sqlite3.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/vendored/sqlite3/sqlite3.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/vendored/sqlite3/sqlite3.c).

## Scoring and Ranking Mechanisms

### Vector Cosine Similarity

The semantic search computes similarity via **cosine distance** between a query vector and stored token vectors. The scoring incorporates a **Reflective Random Indexing (RRI)** pass that refines vector orientations based on co‑occurrence patterns, as seen in [[`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c#L40-L49). Results are filtered against a configurable threshold (`CBM_SEM_EDGE_THRESHOLD`, default 0.80) defined in [[`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c#L22-L31).

### BM25 Ranking Function

BM25 scores are calculated by SQLite’s `fts5Bm25Function`, which returns a **negative** numeric rank where lower values indicate higher relevance. To maintain performance on large codebases, the engine implements early termination via the `BM25_INNER_LIMIT` constant (default **2000**), limiting the inner candidate set before joining with the main `nodes` table. This logic is located in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L89-L96).

## Query Interfaces and MCP Integration

### Semantic Query Format

Semantic queries are submitted as JSON‑encoded arrays of keywords (e.g., `["open", "file"]`). The MCP layer converts these arrays into a dense query vector by calling `cbm_sem_random_index` for each term and aggregating results with `cbm_sem_vec_add_scaled`. This transformation happens in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L1880-L1888).

### BM25 Query Format

BM25 queries accept free‑form strings that are sanitized and tokenized by `bm25_build_match`. This function strips non‑alphanumerics, splits on whitespace, and joins tokens with the `OR` operator for the FTS5 `MATCH` clause, as shown in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L62-L70). The search executes a two‑stage SQL query—first fetching candidates from the FTS5 index, then joining with the `nodes` table—to keep complexity bounded, detailed in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L88-L92).

## Code Examples: Using Both Search Modes

The MCP CLI exposes both engines through the `search` endpoint. The following examples demonstrate how to invoke each mode.

```bash

# BM25 full‑text search for exact identifier matches

codebase-memory-mcp search \
  '{"project":"myrepo","query":"openFile","file_pattern":"src/**/*.c","limit":20}'

```

This command triggers `bm25_search()` in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L89-L96), which builds the FTS5 MATCH string and applies the file pattern filter.

```bash

# Semantic vector search for conceptual similarity

codebase-memory-mcp search \
  '{"project":"myrepo","semantic_query":["open","file"],"limit":20}'

```

When `semantic_query` is present, the CLI invokes `run_semantic_query()` in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L1880-L1888), constructing a query vector and traversing the `SEMANTICALLY_RELATED` edges generated by the pipeline.

Both commands return a JSON object with a `"results"` array. BM25 results include a `"rank"` field (negative float), while semantic results include a `"semantic_score"` (cosine similarity 0.0–1.0).

## Configuration Parameters and Tuning

### Semantic Search Tuning

Adjust the global similarity threshold at runtime using the environment variable `CBM_SEMANTIC_THRESHOLD`. Individual signal weights (e.g., `CBM_SEM_W_TFIDF`) can be modified in [[`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c#L22-L31) before corpus finalization.

### BM25 Tuning

The `BM25_INNER_LIMIT` constant controls the maximum number of candidates fetched from the FTS5 index before ranking. Lower values improve latency but may miss relevant matches in large result sets. This parameter is defined in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L92-L93).

## When to Use Semantic Search vs. BM25

- **Semantic Search**: Use for **conceptual** queries where you need fuzzy matching across naming conventions (e.g., finding functions that manipulate a `buffer` regardless of whether the variable is named `buf`, `buffer`, or `dataBuffer`). Ideal for large, polyglot codebases where exact token matches are too noisy.
- **BM25 Full‑Text Search**: Use for **exact** text retrieval, such as locating every occurrence of a specific function name, comment string, or error message. Best for precise, high‑speed lookups where the literal token is known.

## Summary

- **Semantic search** in `codebase‑memory‑mcp` uses dense vectors (TF‑IDF + Random Indexing) and cosine similarity to generate `SEMANTICALLY_RELATED` edges, configured via `CBM_SEM_EDGE_THRESHOLD` and signal weights.
- **BM25 search** leverages SQLite FTS5 with a custom `cbm_camel_split` tokenizer, ranking via the BM25 function and early‑terminating at `BM25_INNER_LIMIT` (default 2000).
- Both engines are exposed through the same MCP `search` endpoint: use `semantic_query` for vector search and `query` for lexical search.
- Semantic search excels at conceptual, cross‑language retrieval; BM25 excels at fast, exact token matching.

## Frequently Asked Questions

### What is the difference between semantic and BM25 search in codebase‑memory‑mcp?

Semantic search uses high‑dimensional vector embeddings and cosine similarity to find conceptually related code, generating `SEMANTICALLY_RELATED` graph edges. BM25 search uses SQLite’s FTS5 inverted index with a custom tokenizer to rank exact token matches lexically. The former handles fuzzy, conceptual queries, while the latter provides deterministic, fast literal lookups.

### How does the BM25 tokenizer handle camelCase and snake_case identifiers?

The `cbm_camel_split` tokenizer, defined in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L1588-L1660), splits identifiers on case transitions and underscores. For example, `openFileStream` tokenizes into `open`, `file`, and `stream`, while `open_file_stream` produces the same tokens. This ensures that partial matches against any sub‑token are possible.

### What is the default BM25_INNER_LIMIT and why does it matter?

The default `BM25_INNER_LIMIT` is **2000** candidates, as specified in [[`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L92-L93). This cap prevents the ranking function from processing excessive rows in large codebases, keeping query latency under 200 ms. Increasing this value improves recall at the cost of performance.

### Can I combine semantic and BM25 searches in a single query?

While the MCP protocol exposes them as separate fields (`query` for BM25 and `semantic_query` for semantic), the current implementation processes them independently. To combine results, issue both query types and merge the result sets client‑side, weighting the BM25 rank against the semantic score based on your application’s precision requirements.