# How to Perform Semantic Search with Embeddings in codebase-memory-mcp

> Learn to perform semantic search with embeddings in codebase-memory-mcp. Explore fast on-device queries using pre-computed nomic-embed-code token embeddings and cosine similarity.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: how-to-guide
- Published: 2026-07-09

---

**codebase-memory-mcp implements on-device semantic search using pre-computed nomic-embed-code token embeddings stored as a static lookup table, enabling fast cosine similarity queries without external services.**

The `codebase-memory-mcp` repository provides a Model Context Protocol (MCP) server that performs semantic code search entirely offline. Unlike cloud-based embedding services, this project generates a static vector table from the open-source **nomic-embed-code** model and compiles it directly into the binary, allowing millisecond-latency semantic queries against your codebase.

## Architecture of the Semantic Search Pipeline

The semantic search implementation follows a three-stage pipeline that separates heavy compute (embedding generation) from runtime performance (static lookup).

### Stage 1: Embedding Extraction

The Python script [`scripts/extract_nomic_vectors.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/extract_nomic_vectors.py) handles the one-time generation of code token embeddings. This script loads the `nomic-embed-code` model, filters the tokenizer vocabulary to code-relevant identifiers using the `is_code_relevant` function, and discards BPE noise and punctuation. Valid tokens are normalized via `clean_token` (stripping BPE markers and underscores), then passed to the model as `"search_query: <token>"` strings.

The extraction process performs mean-pooling over non-padding tokens, applies L2 normalization, and optionally runs `simulated_attention` to blend each vector with its K-nearest neighbors. After mean-centering the entire corpus to reduce anisotropy, vectors are quantized to `int8` in the range `[-127, 127]` and written to `code_vectors.bin` with the binary layout `[int32 count][int32 dim] + count×dim int8`.

### Stage 2: Static Compilation

The generated artifacts (`code_vectors.bin`, [`code_vectors.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/code_vectors.h), `code_vectors_blob.S`, and [`code_tokens.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/code_tokens.h)) are compiled into the executable via `Makefile.cbm`. The header [`vendored/nomic/code_vectors.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/vendored/nomic/code_vectors.h) exposes `PRETRAINED_VECTOR_BLOB` and the token map `PRETRAINED_TOKENS`, embedding the entire embedding table directly into the binary's data section without requiring external file I/O at runtime.

### Stage 3: Runtime Query Execution

At runtime, [`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c) handles query processing. When a client supplies a semantic query, the function `cbm_sem_tokenize` splits identifiers by delimiters and camel-case transitions, expands abbreviations, and retrieves token embeddings via `pretrained_vec_at`. The query vector is computed as the L2-normalized average of constituent token embeddings.

Cosine similarity is calculated between the query vector and every stored token vector using `dot_product / (norms)`. The top-k results are blended with other signals (TF-IDF, Random Indexing, MinHash) using weights defined in [`semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/semantic.c) (e.g., `CBM_SEM_W_TFIDF`, `CBM_SEM_W_RI`), then returned as JSON via [`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c).

## Generating the Embedding Table

To build the static vector table from scratch, run the extraction script after installing the required dependencies.

```bash

# Install Python dependencies

pip install torch transformers sentence-transformers

# Generate embeddings (CPU mode shown; use --device cuda for GPU acceleration)

python3 scripts/extract_nomic_vectors.py \
    --output-dir vendored/nomic \
    --device cpu

```

The script outputs:
- `code_vectors.bin` (13.2 MB): Quantized `int8` vectors
- [`code_vectors.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/code_vectors.h): C header with metadata
- `code_vectors_blob.S`: Assembler wrapper for binary inclusion
- [`code_tokens.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/code_tokens.h): Token string mappings

After generation, rebuild the MCP binary to embed the new vectors:

```bash
make -f Makefile.cbm clean
make -f Makefile.cbm

```

## Performing Semantic Queries

### Command-Line Interface

Query the codebase using natural language terms or code identifiers. The `--semantic_query` flag accepts multiple terms that get tokenized and averaged into a query vector.

```bash

# Search for files related to "send" and "publish" operations

codebase-mcp query \
    --semantic_query send publish \
    --path ./my-project \
    --limit 5

```

The JSON response includes ranked file matches under the `semantic_query` key:

```json
{
  "semantic_query": [
    {"file":"src/net/transport.c","score":0.873},
    {"file":"src/pubsub/broker.c","score":0.862},
    {"file":"src/cli/cli.c","score":0.845}
  ]
}

```

### C API Integration

For direct integration into other tools, use the C API exposed in [`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c):

```c
#include "nomic/code_vectors.h"
#include "semantic/semantic.h"

// Prepare buffer for query vector (768 dimensions)
float query_vec[PRETRAINED_DIM];
char tokens[MAX_TOKENS][MAX_TOKEN_LEN];

// Tokenize input and compute average embedding
int n = cbm_sem_tokenize("send_message_async", tokens, MAX_TOKENS);
cbm_sem_average_embedding(tokens, n, query_vec);

// Retrieve top-10 matches
int idxs[10];
float sims[10];
cbm_sem_topk(query_vec, PRETRAINED_DIM, idxs, sims, 10);

// Output results
for (int i = 0; i < 10; ++i) {
    printf("%s  (%.3f)\n", PRETRAINED_TOKENS[idxs[i]], sims[i]);
}

```

## Summary

- **Offline Operation**: All embeddings are pre-computed and compiled into the binary, eliminating network latency and external dependencies.
- **Nomic-Embed-Code Model**: Uses the open-source code-specific embedding model with 768-dimensional vectors, filtered to code-relevant tokens and quantized to `int8` for space efficiency.
- **Three-Stage Pipeline**: Extraction ([`scripts/extract_nomic_vectors.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/extract_nomic_vectors.py)), compilation (`Makefile.cbm`), and runtime querying ([`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c)).
- **Cosine Similarity**: Runtime queries compute the L2-normalized average of token embeddings and compare against the static table using cosine similarity.
- **Signal Blending**: Semantic scores are combined with TF-IDF, Random Indexing, and MinHash signals using configurable weights in [`semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/semantic.c).

## Frequently Asked Questions

### What embedding model does codebase-memory-mcp use?

The project uses **nomic-embed-code**, an open-source model specifically trained for code understanding. As implemented in [`scripts/extract_nomic_vectors.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/extract_nomic_vectors.py), the model generates 768-dimensional vectors for code tokens, which are then quantized to `int8` and stored in `code_vectors.bin`.

### Can I customize the embedding table for my specific codebase?

Yes, though it requires regenerating the static vectors. Run [`scripts/extract_nomic_vectors.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/extract_nomic_vectors.py) with your custom token filtering logic or fine-tuned model weights, then recompile using `Makefile.cbm`. The `is_code_relevant` and `clean_token` functions in the extraction script control which tokens are included in the final table.

### How does the runtime handle multi-token queries?

The runtime tokenizes queries using `cbm_sem_tokenize` in [`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c), which splits camelCase and PascalCase identifiers and expands abbreviations. Each token's embedding is retrieved from `PRETRAINED_TOKENS`, summed together, and L2-normalized to create the final query vector for cosine similarity comparison.

### Is GPU required for running semantic searches?

No. The embedding extraction phase can utilize GPU (`--device cuda`) for faster processing, but the runtime semantic search runs entirely on CPU using pre-computed vectors. The only heavy computation is the initial extraction; queries execute as fast in-memory table lookups with cosine distance calculations.