How semantic_query Uses Vector Embeddings for Vocabulary-Mismatch Code Search
The semantic_query feature converts query tokens into 768‑dimensional dense vectors using a pre‑trained Nomic Random‑Indexing model, then scores codebase functions by dot‑product similarity against pre‑computed semantic embeddings, enabling retrieval of relevant code even when terminology differs.
The codebase-memory-mcp tool solves the vocabulary‑mismatch problem that plagues traditional code search through its semantic_query capability. Unlike exact string matching, this feature leverages continuous vector spaces to find semantically related functions even when they use different terminology than your query.
The Vector Embedding Pipeline
When you invoke the --semantic-query flag, the CLI passes your keywords into the core MCP processor defined in src/mcp/mcp.c. The run_semantic_query function orchestrates a multi‑stage pipeline that transforms text into comparable mathematical representations.
Token Embedding with the Nomic Model
Each keyword supplied to --semantic-query is transformed into a dense vector using the Nomic Random‑Indexing model located in vendored/nomic. This pre‑trained model generates 768‑dimensional embeddings for any input token, ensuring that unknown or out‑of‑vocabulary words still receive meaningful semantic representations. Because the model operates on learned semantic relationships rather than fixed vocabularies, it can associate related concepts like send and publishMessage within the same vector space.
Vector Normalization for Efficient Similarity
Before comparison, the raw embedding vectors undergo L2 normalization. This critical step ensures that cosine similarity calculations reduce to simple dot‑products, significantly speeding up the search computation without sacrificing accuracy. Normalization eliminates magnitude differences between vectors so that only directional similarity—the actual semantic meaning—affects the final score.
Indexing and Similarity Scoring
The search operates against a pre‑built index of semantic function representations generated during the codebase analysis phase.
Per-Function Semantic Vectors
During the indexing pass, src/semantic/semantic.c generates per‑function semantic vectors that encode the aggregate meaning of identifiers, literals, and API calls within each function body. These vectors capture the "aboutness" of code—what concepts and operations the function performs—rather than its literal text. The vectors are stored in the codebase index and serve as the search targets during query execution.
Similarity Computation and Ranking
For each indexed function, run_semantic_query computes the dot‑product between the function's semantic vector and each query keyword vector. These per‑keyword scores are then summed or max‑pooled to produce a final relevance score. Functions are sorted by descending score, and the top N results—controlled by the --limit parameter—are serialized into the JSON output under the "semantic_query" field.
Practical Usage Examples
Invoke vocabulary‑agnostic search from the command line:
codebase-mcp --semantic-query send publish --limit 5
The tool returns ranked results in the JSON output:
{
"semantic_query": [
{
"path": "src/net/transport.c",
"function": "send_message",
"score": 0.92
},
{
"path": "src/pub/sub.c",
"function": "publish_event",
"score": 0.89
}
]
}
Integrate the functionality programmatically using the C API:
/* Build a semantic query argument */
yyjson_mut_doc *doc = yyjson_mut_doc_new(pool);
yyjson_mut_val *root = yyjson_mut_doc_root(doc);
yyjson_mut_val *sq = yyjson_mut_arr_add(doc, root);
yyjson_mut_arr_add_str(doc, sq, "send");
yyjson_mut_arr_add_str(doc, sq, "publish");
/* Pass the doc to the MCP core */
run_semantic_query(doc, root, args, store, project, limit);
Key Source Files
The implementation spans several components:
src/mcp/mcp.c– Contains therun_semantic_queryfunction that orchestrates the vector search and result rankingsrc/semantic/semantic.c– Generates per‑function semantic embeddings during the indexing phasesrc/semantic/semantic.h– Declares embedding APIs and data structures for vector manipulationvendored/nomic– Houses the pre‑trained Random‑Indexing model and token‑to‑vector mappings
Summary
semantic_queryuses 768‑dimensional dense vectors from the Nomic Random‑Indexing model to represent both queries and code- L2 normalization enables fast dot‑product similarity scoring equivalent to cosine similarity
- Per‑function embeddings are generated by
src/semantic/semantic.cduring the semantic indexing pass - The scoring algorithm uses summed or max‑pooled dot‑products to rank function relevance
- Results are returned as ranked JSON objects under the
"semantic_query"field, with quantity controlled by--limit
Frequently Asked Questions
What makes semantic_query different from regular text search?
Regular text search relies on exact token matches or TF‑IDF weighting, which fails when terminology differs between query and implementation. semantic_query operates in a continuous vector space where semantically similar words have proximal embeddings, bridging vocabulary gaps between queries and source code.
How does the Nomic model handle out-of-vocabulary terms?
The Random‑Indexing model in vendored/nomic is trained to produce meaningful 768‑dimensional vectors even for unseen tokens. This ensures that typos, abbreviations, or domain‑specific jargon still receive valid embeddings that approximate their true semantic meaning within the vector space.
Why is L2 normalization important for the similarity calculation?
L2 normalization converts raw embedding vectors to unit length. This mathematical transformation reduces cosine similarity computation to a simple dot‑product, dramatically improving search performance while maintaining the angular relationship between vectors that determines semantic similarity.
Can I adjust the number of results returned by semantic_query?
Yes. The --limit flag controls how many top‑scoring functions are returned. This parameter is passed directly to the ranking logic in run_semantic_query and determines the size of the "semantic_query" array in the final JSON output.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →