How semantic_query Uses Vector Embeddings for Vocabulary-Mismatch Code Search

The semantic_query feature converts query tokens into 768‑dimensional dense vectors using a pre‑trained Nomic Random‑Indexing model, then scores codebase functions by dot‑product similarity against pre‑computed semantic embeddings, enabling retrieval of relevant code even when terminology differs.

The codebase-memory-mcp tool solves the vocabulary‑mismatch problem that plagues traditional code search through its semantic_query capability. Unlike exact string matching, this feature leverages continuous vector spaces to find semantically related functions even when they use different terminology than your query.

The Vector Embedding Pipeline

When you invoke the --semantic-query flag, the CLI passes your keywords into the core MCP processor defined in src/mcp/mcp.c. The run_semantic_query function orchestrates a multi‑stage pipeline that transforms text into comparable mathematical representations.

Token Embedding with the Nomic Model

Each keyword supplied to --semantic-query is transformed into a dense vector using the Nomic Random‑Indexing model located in vendored/nomic. This pre‑trained model generates 768‑dimensional embeddings for any input token, ensuring that unknown or out‑of‑vocabulary words still receive meaningful semantic representations. Because the model operates on learned semantic relationships rather than fixed vocabularies, it can associate related concepts like send and publishMessage within the same vector space.

Vector Normalization for Efficient Similarity

Before comparison, the raw embedding vectors undergo L2 normalization. This critical step ensures that cosine similarity calculations reduce to simple dot‑products, significantly speeding up the search computation without sacrificing accuracy. Normalization eliminates magnitude differences between vectors so that only directional similarity—the actual semantic meaning—affects the final score.

Indexing and Similarity Scoring

The search operates against a pre‑built index of semantic function representations generated during the codebase analysis phase.

Per-Function Semantic Vectors

During the indexing pass, src/semantic/semantic.c generates per‑function semantic vectors that encode the aggregate meaning of identifiers, literals, and API calls within each function body. These vectors capture the "aboutness" of code—what concepts and operations the function performs—rather than its literal text. The vectors are stored in the codebase index and serve as the search targets during query execution.

Similarity Computation and Ranking

For each indexed function, run_semantic_query computes the dot‑product between the function's semantic vector and each query keyword vector. These per‑keyword scores are then summed or max‑pooled to produce a final relevance score. Functions are sorted by descending score, and the top N results—controlled by the --limit parameter—are serialized into the JSON output under the "semantic_query" field.

Practical Usage Examples

Invoke vocabulary‑agnostic search from the command line:

codebase-mcp --semantic-query send publish --limit 5

The tool returns ranked results in the JSON output:

{
  "semantic_query": [
    {
      "path": "src/net/transport.c",
      "function": "send_message",
      "score": 0.92
    },
    {
      "path": "src/pub/sub.c",
      "function": "publish_event",
      "score": 0.89
    }
  ]
}

Integrate the functionality programmatically using the C API:

/* Build a semantic query argument */
yyjson_mut_doc *doc = yyjson_mut_doc_new(pool);
yyjson_mut_val *root = yyjson_mut_doc_root(doc);
yyjson_mut_val *sq = yyjson_mut_arr_add(doc, root);
yyjson_mut_arr_add_str(doc, sq, "send");
yyjson_mut_arr_add_str(doc, sq, "publish");

/* Pass the doc to the MCP core */
run_semantic_query(doc, root, args, store, project, limit);

Key Source Files

The implementation spans several components:

  • src/mcp/mcp.c – Contains the run_semantic_query function that orchestrates the vector search and result ranking
  • src/semantic/semantic.c – Generates per‑function semantic embeddings during the indexing phase
  • src/semantic/semantic.h – Declares embedding APIs and data structures for vector manipulation
  • vendored/nomic – Houses the pre‑trained Random‑Indexing model and token‑to‑vector mappings

Summary

  • semantic_query uses 768‑dimensional dense vectors from the Nomic Random‑Indexing model to represent both queries and code
  • L2 normalization enables fast dot‑product similarity scoring equivalent to cosine similarity
  • Per‑function embeddings are generated by src/semantic/semantic.c during the semantic indexing pass
  • The scoring algorithm uses summed or max‑pooled dot‑products to rank function relevance
  • Results are returned as ranked JSON objects under the "semantic_query" field, with quantity controlled by --limit

Frequently Asked Questions

Regular text search relies on exact token matches or TF‑IDF weighting, which fails when terminology differs between query and implementation. semantic_query operates in a continuous vector space where semantically similar words have proximal embeddings, bridging vocabulary gaps between queries and source code.

How does the Nomic model handle out-of-vocabulary terms?

The Random‑Indexing model in vendored/nomic is trained to produce meaningful 768‑dimensional vectors even for unseen tokens. This ensures that typos, abbreviations, or domain‑specific jargon still receive valid embeddings that approximate their true semantic meaning within the vector space.

Why is L2 normalization important for the similarity calculation?

L2 normalization converts raw embedding vectors to unit length. This mathematical transformation reduces cosine similarity computation to a simple dot‑product, dramatically improving search performance while maintaining the angular relationship between vectors that determines semantic similarity.

Can I adjust the number of results returned by semantic_query?

Yes. The --limit flag controls how many top‑scoring functions are returned. This parameter is passed directly to the ranking logic in run_semantic_query and determines the size of the "semantic_query" array in the final JSON output.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →