How to Perform Semantic Search with Embeddings in codebase-memory-mcp

codebase-memory-mcp implements on-device semantic search using pre-computed nomic-embed-code token embeddings stored as a static lookup table, enabling fast cosine similarity queries without external services.

The codebase-memory-mcp repository provides a Model Context Protocol (MCP) server that performs semantic code search entirely offline. Unlike cloud-based embedding services, this project generates a static vector table from the open-source nomic-embed-code model and compiles it directly into the binary, allowing millisecond-latency semantic queries against your codebase.

Architecture of the Semantic Search Pipeline

The semantic search implementation follows a three-stage pipeline that separates heavy compute (embedding generation) from runtime performance (static lookup).

Stage 1: Embedding Extraction

The Python script scripts/extract_nomic_vectors.py handles the one-time generation of code token embeddings. This script loads the nomic-embed-code model, filters the tokenizer vocabulary to code-relevant identifiers using the is_code_relevant function, and discards BPE noise and punctuation. Valid tokens are normalized via clean_token (stripping BPE markers and underscores), then passed to the model as "search_query: <token>" strings.

The extraction process performs mean-pooling over non-padding tokens, applies L2 normalization, and optionally runs simulated_attention to blend each vector with its K-nearest neighbors. After mean-centering the entire corpus to reduce anisotropy, vectors are quantized to int8 in the range [-127, 127] and written to code_vectors.bin with the binary layout [int32 count][int32 dim] + count×dim int8.

Stage 2: Static Compilation

The generated artifacts (code_vectors.bin, code_vectors.h, code_vectors_blob.S, and code_tokens.h) are compiled into the executable via Makefile.cbm. The header vendored/nomic/code_vectors.h exposes PRETRAINED_VECTOR_BLOB and the token map PRETRAINED_TOKENS, embedding the entire embedding table directly into the binary's data section without requiring external file I/O at runtime.

Stage 3: Runtime Query Execution

At runtime, src/semantic/semantic.c handles query processing. When a client supplies a semantic query, the function cbm_sem_tokenize splits identifiers by delimiters and camel-case transitions, expands abbreviations, and retrieves token embeddings via pretrained_vec_at. The query vector is computed as the L2-normalized average of constituent token embeddings.

Cosine similarity is calculated between the query vector and every stored token vector using dot_product / (norms). The top-k results are blended with other signals (TF-IDF, Random Indexing, MinHash) using weights defined in semantic.c (e.g., CBM_SEM_W_TFIDF, CBM_SEM_W_RI), then returned as JSON via src/mcp/mcp.c.

Generating the Embedding Table

To build the static vector table from scratch, run the extraction script after installing the required dependencies.


# Install Python dependencies

pip install torch transformers sentence-transformers

# Generate embeddings (CPU mode shown; use --device cuda for GPU acceleration)

python3 scripts/extract_nomic_vectors.py \
    --output-dir vendored/nomic \
    --device cpu

The script outputs:

  • code_vectors.bin (13.2 MB): Quantized int8 vectors
  • code_vectors.h: C header with metadata
  • code_vectors_blob.S: Assembler wrapper for binary inclusion
  • code_tokens.h: Token string mappings

After generation, rebuild the MCP binary to embed the new vectors:

make -f Makefile.cbm clean
make -f Makefile.cbm

Performing Semantic Queries

Command-Line Interface

Query the codebase using natural language terms or code identifiers. The --semantic_query flag accepts multiple terms that get tokenized and averaged into a query vector.


# Search for files related to "send" and "publish" operations

codebase-mcp query \
    --semantic_query send publish \
    --path ./my-project \
    --limit 5

The JSON response includes ranked file matches under the semantic_query key:

{
  "semantic_query": [
    {"file":"src/net/transport.c","score":0.873},
    {"file":"src/pubsub/broker.c","score":0.862},
    {"file":"src/cli/cli.c","score":0.845}
  ]
}

C API Integration

For direct integration into other tools, use the C API exposed in src/semantic/semantic.c:

#include "nomic/code_vectors.h"
#include "semantic/semantic.h"

// Prepare buffer for query vector (768 dimensions)
float query_vec[PRETRAINED_DIM];
char tokens[MAX_TOKENS][MAX_TOKEN_LEN];

// Tokenize input and compute average embedding
int n = cbm_sem_tokenize("send_message_async", tokens, MAX_TOKENS);
cbm_sem_average_embedding(tokens, n, query_vec);

// Retrieve top-10 matches
int idxs[10];
float sims[10];
cbm_sem_topk(query_vec, PRETRAINED_DIM, idxs, sims, 10);

// Output results
for (int i = 0; i < 10; ++i) {
    printf("%s  (%.3f)\n", PRETRAINED_TOKENS[idxs[i]], sims[i]);
}

Summary

  • Offline Operation: All embeddings are pre-computed and compiled into the binary, eliminating network latency and external dependencies.
  • Nomic-Embed-Code Model: Uses the open-source code-specific embedding model with 768-dimensional vectors, filtered to code-relevant tokens and quantized to int8 for space efficiency.
  • Three-Stage Pipeline: Extraction (scripts/extract_nomic_vectors.py), compilation (Makefile.cbm), and runtime querying (src/semantic/semantic.c).
  • Cosine Similarity: Runtime queries compute the L2-normalized average of token embeddings and compare against the static table using cosine similarity.
  • Signal Blending: Semantic scores are combined with TF-IDF, Random Indexing, and MinHash signals using configurable weights in semantic.c.

Frequently Asked Questions

What embedding model does codebase-memory-mcp use?

The project uses nomic-embed-code, an open-source model specifically trained for code understanding. As implemented in scripts/extract_nomic_vectors.py, the model generates 768-dimensional vectors for code tokens, which are then quantized to int8 and stored in code_vectors.bin.

Can I customize the embedding table for my specific codebase?

Yes, though it requires regenerating the static vectors. Run scripts/extract_nomic_vectors.py with your custom token filtering logic or fine-tuned model weights, then recompile using Makefile.cbm. The is_code_relevant and clean_token functions in the extraction script control which tokens are included in the final table.

How does the runtime handle multi-token queries?

The runtime tokenizes queries using cbm_sem_tokenize in src/semantic/semantic.c, which splits camelCase and PascalCase identifiers and expands abbreviations. Each token's embedding is retrieved from PRETRAINED_TOKENS, summed together, and L2-normalized to create the final query vector for cosine similarity comparison.

Is GPU required for running semantic searches?

No. The embedding extraction phase can utilize GPU (--device cuda) for faster processing, but the runtime semantic search runs entirely on CPU using pre-computed vectors. The only heavy computation is the initial extraction; queries execute as fast in-memory table lookups with cosine distance calculations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →