Semantic Search vs BM25 Full‑Text Search in codebase‑memory‑mcp: Architecture and Implementation
codebase‑memory‑mcp implements two complementary search backends—dense vector semantic search for conceptual similarity and SQLite FTS5 BM25 for exact lexical matching—allowing MCP clients to choose the appropriate retrieval strategy based on query intent.
The codebase‑memory‑mcp repository provides a graph‑centric memory layer for codebases that exposes both semantic and lexical search capabilities through a unified Model Context Protocol (MCP) interface. While both mechanisms serve the same high‑level goal of locating relevant code, they operate on fundamentally different data models: semantic search leverages high‑dimensional vector embeddings and Random Indexing, whereas BM25 relies on SQLite’s FTS5 inverted index with a custom camelCase‑aware tokenizer.
Core Data Models: Embeddings vs. Inverted Indices
Semantic Search Vectors
The semantic engine stores a dense vector for every distinct token and identifier in a cbm_sem_corpus_t structure. These vectors are constructed using a hybrid of TF‑IDF, Random Indexing, and pretrained embeddings (Nomic), augmented with code‑specific signals such as API signatures and type information. The corpus initialization and vector generation logic reside in [src/semantic/semantic.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c#L31-L50).
During the pipeline phase, the system emits SEMANTICALLY_RELATED edges by comparing these vectors using an 11‑signal weighted similarity model (CBM_SEM_W_* constants). This process is implemented in [src/pipeline/pass_semantic_edges.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pass_semantic_edges.c#L2-L17).
BM25 Full‑Text Index
In contrast, the BM25 implementation relies on SQLite’s FTS5 virtual table named nodes_fts. The index is populated by a custom tokenizer called cbm_camel_split, which is registered in [src/mcp/mcp.c](https://github.comDeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L1588-L1660). This tokenizer splits identifiers on camelCase and snake_case boundaries, ensuring that openFile and open_file both generate searchable tokens open and file. The underlying FTS5 engine is vendored in [vendored/sqlite3/sqlite3.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/vendored/sqlite3/sqlite3.c).
Scoring and Ranking Mechanisms
Vector Cosine Similarity
The semantic search computes similarity via cosine distance between a query vector and stored token vectors. The scoring incorporates a Reflective Random Indexing (RRI) pass that refines vector orientations based on co‑occurrence patterns, as seen in [src/semantic/semantic.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c#L40-L49). Results are filtered against a configurable threshold (CBM_SEM_EDGE_THRESHOLD, default 0.80) defined in [src/semantic/semantic.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c#L22-L31).
BM25 Ranking Function
BM25 scores are calculated by SQLite’s fts5Bm25Function, which returns a negative numeric rank where lower values indicate higher relevance. To maintain performance on large codebases, the engine implements early termination via the BM25_INNER_LIMIT constant (default 2000), limiting the inner candidate set before joining with the main nodes table. This logic is located in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L89-L96).
Query Interfaces and MCP Integration
Semantic Query Format
Semantic queries are submitted as JSON‑encoded arrays of keywords (e.g., ["open", "file"]). The MCP layer converts these arrays into a dense query vector by calling cbm_sem_random_index for each term and aggregating results with cbm_sem_vec_add_scaled. This transformation happens in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L1880-L1888).
BM25 Query Format
BM25 queries accept free‑form strings that are sanitized and tokenized by bm25_build_match. This function strips non‑alphanumerics, splits on whitespace, and joins tokens with the OR operator for the FTS5 MATCH clause, as shown in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L62-L70). The search executes a two‑stage SQL query—first fetching candidates from the FTS5 index, then joining with the nodes table—to keep complexity bounded, detailed in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L88-L92).
Code Examples: Using Both Search Modes
The MCP CLI exposes both engines through the search endpoint. The following examples demonstrate how to invoke each mode.
# BM25 full‑text search for exact identifier matches
codebase-memory-mcp search \
'{"project":"myrepo","query":"openFile","file_pattern":"src/**/*.c","limit":20}'
This command triggers bm25_search() in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L89-L96), which builds the FTS5 MATCH string and applies the file pattern filter.
# Semantic vector search for conceptual similarity
codebase-memory-mcp search \
'{"project":"myrepo","semantic_query":["open","file"],"limit":20}'
When semantic_query is present, the CLI invokes run_semantic_query() in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L1880-L1888), constructing a query vector and traversing the SEMANTICALLY_RELATED edges generated by the pipeline.
Both commands return a JSON object with a "results" array. BM25 results include a "rank" field (negative float), while semantic results include a "semantic_score" (cosine similarity 0.0–1.0).
Configuration Parameters and Tuning
Semantic Search Tuning
Adjust the global similarity threshold at runtime using the environment variable CBM_SEMANTIC_THRESHOLD. Individual signal weights (e.g., CBM_SEM_W_TFIDF) can be modified in [src/semantic/semantic.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c#L22-L31) before corpus finalization.
BM25 Tuning
The BM25_INNER_LIMIT constant controls the maximum number of candidates fetched from the FTS5 index before ranking. Lower values improve latency but may miss relevant matches in large result sets. This parameter is defined in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L92-L93).
When to Use Semantic Search vs. BM25
- Semantic Search: Use for conceptual queries where you need fuzzy matching across naming conventions (e.g., finding functions that manipulate a
bufferregardless of whether the variable is namedbuf,buffer, ordataBuffer). Ideal for large, polyglot codebases where exact token matches are too noisy. - BM25 Full‑Text Search: Use for exact text retrieval, such as locating every occurrence of a specific function name, comment string, or error message. Best for precise, high‑speed lookups where the literal token is known.
Summary
- Semantic search in
codebase‑memory‑mcpuses dense vectors (TF‑IDF + Random Indexing) and cosine similarity to generateSEMANTICALLY_RELATEDedges, configured viaCBM_SEM_EDGE_THRESHOLDand signal weights. - BM25 search leverages SQLite FTS5 with a custom
cbm_camel_splittokenizer, ranking via the BM25 function and early‑terminating atBM25_INNER_LIMIT(default 2000). - Both engines are exposed through the same MCP
searchendpoint: usesemantic_queryfor vector search andqueryfor lexical search. - Semantic search excels at conceptual, cross‑language retrieval; BM25 excels at fast, exact token matching.
Frequently Asked Questions
What is the difference between semantic and BM25 search in codebase‑memory‑mcp?
Semantic search uses high‑dimensional vector embeddings and cosine similarity to find conceptually related code, generating SEMANTICALLY_RELATED graph edges. BM25 search uses SQLite’s FTS5 inverted index with a custom tokenizer to rank exact token matches lexically. The former handles fuzzy, conceptual queries, while the latter provides deterministic, fast literal lookups.
How does the BM25 tokenizer handle camelCase and snake_case identifiers?
The cbm_camel_split tokenizer, defined in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L1588-L1660), splits identifiers on case transitions and underscores. For example, openFileStream tokenizes into open, file, and stream, while open_file_stream produces the same tokens. This ensures that partial matches against any sub‑token are possible.
What is the default BM25_INNER_LIMIT and why does it matter?
The default BM25_INNER_LIMIT is 2000 candidates, as specified in [src/mcp/mcp.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c#L92-L93). This cap prevents the ranking function from processing excessive rows in large codebases, keeping query latency under 200 ms. Increasing this value improves recall at the cost of performance.
Can I combine semantic and BM25 searches in a single query?
While the MCP protocol exposes them as separate fields (query for BM25 and semantic_query for semantic), the current implementation processes them independently. To combine results, issue both query types and merge the result sets client‑side, weighting the BM25 rank against the semantic score based on your application’s precision requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →