# How semantic_query Uses Vector Embeddings for Vocabulary-Mismatch Code Search

> Discover how semantic_query leverages vector embeddings for vocabulary-mismatch code search. Retrieve relevant code despite differing terminology with Nomic AI and dot-product similarity.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: how-to-guide
- Published: 2026-07-10

---

**The `semantic_query` feature converts query tokens into 768‑dimensional dense vectors using a pre‑trained Nomic Random‑Indexing model, then scores codebase functions by dot‑product similarity against pre‑computed semantic embeddings, enabling retrieval of relevant code even when terminology differs.**

The `codebase-memory-mcp` tool solves the vocabulary‑mismatch problem that plagues traditional code search through its `semantic_query` capability. Unlike exact string matching, this feature leverages continuous vector spaces to find semantically related functions even when they use different terminology than your query.

## The Vector Embedding Pipeline

When you invoke the `--semantic-query` flag, the CLI passes your keywords into the core MCP processor defined in [`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c). The `run_semantic_query` function orchestrates a multi‑stage pipeline that transforms text into comparable mathematical representations.

### Token Embedding with the Nomic Model

Each keyword supplied to `--semantic-query` is transformed into a dense vector using the **Nomic Random‑Indexing model** located in `vendored/nomic`. This pre‑trained model generates **768‑dimensional embeddings** for any input token, ensuring that unknown or out‑of‑vocabulary words still receive meaningful semantic representations. Because the model operates on learned semantic relationships rather than fixed vocabularies, it can associate related concepts like `send` and `publishMessage` within the same vector space.

### Vector Normalization for Efficient Similarity

Before comparison, the raw embedding vectors undergo **L2 normalization**. This critical step ensures that cosine similarity calculations reduce to simple dot‑products, significantly speeding up the search computation without sacrificing accuracy. Normalization eliminates magnitude differences between vectors so that only directional similarity—the actual semantic meaning—affects the final score.

## Indexing and Similarity Scoring

The search operates against a pre‑built index of semantic function representations generated during the codebase analysis phase.

### Per-Function Semantic Vectors

During the indexing pass, [`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c) generates **per‑function semantic vectors** that encode the aggregate meaning of identifiers, literals, and API calls within each function body. These vectors capture the "aboutness" of code—what concepts and operations the function performs—rather than its literal text. The vectors are stored in the codebase index and serve as the search targets during query execution.

### Similarity Computation and Ranking

For each indexed function, `run_semantic_query` computes the dot‑product between the function's semantic vector and each query keyword vector. These per‑keyword scores are then **summed or max‑pooled** to produce a final relevance score. Functions are sorted by descending score, and the top N results—controlled by the `--limit` parameter—are serialized into the JSON output under the `"semantic_query"` field.

## Practical Usage Examples

Invoke vocabulary‑agnostic search from the command line:

```bash
codebase-mcp --semantic-query send publish --limit 5

```

The tool returns ranked results in the JSON output:

```json
{
  "semantic_query": [
    {
      "path": "src/net/transport.c",
      "function": "send_message",
      "score": 0.92
    },
    {
      "path": "src/pub/sub.c",
      "function": "publish_event",
      "score": 0.89
    }
  ]
}

```

Integrate the functionality programmatically using the C API:

```c
/* Build a semantic query argument */
yyjson_mut_doc *doc = yyjson_mut_doc_new(pool);
yyjson_mut_val *root = yyjson_mut_doc_root(doc);
yyjson_mut_val *sq = yyjson_mut_arr_add(doc, root);
yyjson_mut_arr_add_str(doc, sq, "send");
yyjson_mut_arr_add_str(doc, sq, "publish");

/* Pass the doc to the MCP core */
run_semantic_query(doc, root, args, store, project, limit);

```

## Key Source Files

The implementation spans several components:

- [`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c) – Contains the `run_semantic_query` function that orchestrates the vector search and result ranking
- [`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c) – Generates per‑function semantic embeddings during the indexing phase
- [`src/semantic/semantic.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.h) – Declares embedding APIs and data structures for vector manipulation
- `vendored/nomic` – Houses the pre‑trained Random‑Indexing model and token‑to‑vector mappings

## Summary

- `semantic_query` uses **768‑dimensional dense vectors** from the Nomic Random‑Indexing model to represent both queries and code
- **L2 normalization** enables fast dot‑product similarity scoring equivalent to cosine similarity
- Per‑function embeddings are generated by [`src/semantic/semantic.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/semantic/semantic.c) during the semantic indexing pass
- The scoring algorithm uses **summed or max‑pooled dot‑products** to rank function relevance
- Results are returned as ranked JSON objects under the `"semantic_query"` field, with quantity controlled by `--limit`

## Frequently Asked Questions

### What makes semantic_query different from regular text search?

Regular text search relies on exact token matches or TF‑IDF weighting, which fails when terminology differs between query and implementation. `semantic_query` operates in a continuous vector space where semantically similar words have proximal embeddings, bridging vocabulary gaps between queries and source code.

### How does the Nomic model handle out-of-vocabulary terms?

The Random‑Indexing model in `vendored/nomic` is trained to produce meaningful 768‑dimensional vectors even for unseen tokens. This ensures that typos, abbreviations, or domain‑specific jargon still receive valid embeddings that approximate their true semantic meaning within the vector space.

### Why is L2 normalization important for the similarity calculation?

L2 normalization converts raw embedding vectors to unit length. This mathematical transformation reduces cosine similarity computation to a simple dot‑product, dramatically improving search performance while maintaining the angular relationship between vectors that determines semantic similarity.

### Can I adjust the number of results returned by semantic_query?

Yes. The `--limit` flag controls how many top‑scoring functions are returned. This parameter is passed directly to the ranking logic in `run_semantic_query` and determines the size of the `"semantic_query"` array in the final JSON output.