What Is Semantic Search in code-review-graph and How Does It Work?
Semantic search in code-review-graph is an MCP tool that locates relevant functions and classes by combining full-text keyword matching with vector embedding similarity to rank code nodes by relevance.
The semantic_search_nodes tool is a core component of the tirth8205/code-review-graph repository, designed to bridge the gap between natural language queries and code navigation. Unlike traditional regex-based search, this semantic search implementation leverages both lexical and contextual understanding to identify relevant code entities. It enables developers and automated agents to discover functions, classes, and other graph nodes using intuitive descriptions rather than exact identifier matches.
What Is Semantic Search in code-review-graph?
Semantic search refers to the semantic_search_nodes MCP tool that indexes and queries code entities using dual search modalities. The system treats your codebase as a searchable knowledge graph where each node represents a function, class, or module with associated metadata and embeddings.
The tool generates structured results containing node identifiers, file paths, line numbers, and a provenance field that indicates whether each match originated from keyword matching, vector similarity, or both signals.
How Semantic Search Works: The Two-Strategy Engine
The implementation combines complementary search strategies to maximize recall and precision.
Full-Text Search (FTS)
The Full-Text Search strategy scans indexed source code and documentation for literal keyword matches. This component acts as a fast filter that retrieves nodes containing the query terms in their source text, docstrings, or comments. FTS provides high confidence for exact terminology but may miss conceptually related code using different vocabulary.
Vector Similarity Search
The Embedding (vector) similarity strategy encodes both the natural language query and pre-computed node embeddings into dense vectors. Using the embedding model defined in code_review_graph/embeddings.py (approximately line 1403 in the semantic_search function), the system calculates cosine similarity between the query vector and each node's stored vector. This captures semantic relationships beyond exact keyword matches, identifying functions that implement related concepts even when naming conventions differ.
Result Merging and Ranking
The engine constructs a candidate set from both FTS hits and embedding similarity scores, then merges and deduplicates the results. Each node receives a combined relevance score weighted to favor matches appearing in both search modes. The final ranked list includes provenance metadata indicating which strategies contributed to each result, allowing downstream agents to calibrate confidence levels appropriately.
Implementation Details and Source Code References
The semantic search architecture spans several key files in the repository:
- In
code_review_graph/tools/query.py(approximately line 699), thesemantic_search_nodesfunction implements the core orchestration logic, coordinating between text search and vector comparison. - The vector similarity computation resides in
code_review_graph/embeddings.py(around line 1403) within thesemantic_searchfunction, which handles embedding generation and similarity scoring. - The CLI wrapper exposing this functionality lives in
code_review_graph/main.py(lines 322-374), registering the tool assemantic_search_nodes_tool. - Tool registration for the MCP framework occurs in
code_review_graph/tools/__init__.py, making the function discoverable by the server. - User-facing documentation and usage guidance appear in
code_review_graph/skills.py(around line 1165).
How to Use Semantic Search in code-review-graph
You can invoke semantic search through three primary interfaces depending on your integration needs.
Command-Line Interface
Start the MCP server with the query tool enabled to access semantic search via CLI:
# Serve with the semantic search tool available
crg serve --tools query_graph_tool,semantic_search_nodes_tool
# Execute a search query
> semantic_search_nodes(query="user authentication")
This interface is defined in code_review_graph/main.py within the semantic_search_nodes_tool wrapper implementation.
Programmatic Python API
Import and call the function directly for custom scripts or agent implementations:
from code_review_graph.tools.query import semantic_search_nodes
# Execute semantic search with a natural language query
result = semantic_search_nodes("database connection pooling")
top_matches = result["nodes"][:5] # Access top 5 results
sources = result["provenance"] # Review match sources (FTS, embedding, or both)
This direct API is tested in tests/test_tools.py (lines 156-171) and provides full access to the provenance metadata without server overhead.
Tool Chaining
Combine semantic search with other graph analysis tools for complex workflows:
from code_review_graph.tools.query import semantic_search_nodes, query_graph
# First identify relevant function via semantic search
search_result = semantic_search_nodes("error handling middleware")
target = search_result["nodes"][0]
# Then retrieve its call relationships
call_graph = query_graph(node=target["id"])
This chaining pattern appears in integration tests at tests/test_main.py (lines 428-462), demonstrating how to bridge discovery and analysis operations.
Summary
- Semantic search in code-review-graph combines Full-Text Search and vector embedding similarity to locate code entities using natural language.
- The
semantic_search_nodestool returns structured results with provenance metadata indicating whether matches came from keyword matching, semantic similarity, or both. - Implementation spans
code_review_graph/tools/query.py(orchestration),code_review_graph/embeddings.py(vector operations), andcode_review_graph/main.py(CLI exposure). - Available via CLI, direct Python API, or MCP tool chaining for flexible integration into review workflows.
- The dual-strategy approach balances exact keyword recall with conceptual semantic understanding.
Frequently Asked Questions
What is the difference between FTS and embedding search in code-review-graph?
Full-Text Search (FTS) looks for literal keyword matches in source code and documentation, providing exact lexical matches. Embedding search converts queries and code into dense vectors to measure conceptual similarity, capturing related functionality even with different terminology. The provenance field in results indicates which method contributed to each match.
How does semantic search handle result ranking when both strategies find the same node?
When a node appears in both FTS and embedding results, the system assigns a higher combined relevance score, prioritizing these dual-signal matches in the final ranking. This weighting reflects increased confidence that the node is genuinely relevant to the query intent.
Can I use semantic search without running the MCP server?
Yes. While the CLI requires the server (crg serve), you can import semantic_search_nodes directly from code_review_graph.tools.query and execute searches programmatically without starting the server infrastructure, as demonstrated in the test suite.
What information does the provenance field provide?
The provenance field records whether each result originated from Full-Text Search, embedding similarity, or both sources. This metadata allows calling agents to filter results—accepting only embedding-confirmed matches or requiring both signals for high-stakes code review decisions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →