How code-review-graph Enables Semantic Search Using Embeddings for Natural Language Queries
code-review-graph implements a hybrid retrieval system that transforms code entities into dense vector embeddings using configurable providers, persists them in SQLite, and executes cosine-similarity search to answer natural-language queries alongside traditional full-text search.
The open-source repository tirth8205/code-review-graph provides a graph-based code analysis tool that extends beyond syntactic structure to support semantic search using embeddings. By converting qualified names, docstrings, and comments into high-dimensional vectors, the system enables developers to locate relevant code using plain English descriptions rather than exact keyword matches.
The Architecture of Semantic Search
The semantic search capability rests on a five-stage pipeline that bridges natural language and code representation. The architecture separates concerns between vector generation, persistent storage, approximate nearest neighbor (ANN) indexing, and result fusion.
Embedding Providers wrap external or local models to perform text-to-vector conversion. The codebase defines three concrete implementations: LocalEmbeddingProvider, OpenAIEmbeddingProvider, and MiniMaxEmbeddingProvider. Selection occurs via configuration keys _embedding_provider and _embedding_model, validated in tests/test_embeddings.py.
Vector Storage uses a dedicated SQLite table created within GraphStore._init_db(). The schema stores qualified_name as the primary key alongside a binary vector blob, plus metadata columns provider and model to track which model generated each embedding.
Index Construction happens through _build_embedding_index in code_review_graph/eval/runner.py. This function materializes an ANN index—backed by faiss or brute-force scanning—whenever new nodes are added or model configurations change.
Search Dispatch routes queries through search.search(query, mode=...) in the search module. The dispatcher invokes _embedding_search for vector similarity, _fts_search for keyword matching, or both when operating in hybrid mode.
Result Ranking merges outputs via _merge_results, combining cosine similarity scores from embeddings with relevance scores from the full-text search (FTS) engine to produce a unified ranking.
Core Implementation Components
Embedding Providers and Configuration
The system abstracts model interactions behind the provider interface defined in code_review_graph/embeddings.py. Concrete classes handle authentication, batching, and dimensionality normalization:
- LocalEmbeddingProvider: Runs inference locally using sentence-transformers or similar frameworks.
- OpenAIEmbeddingProvider: Calls the OpenAI API for
text-embedding-3-smallor comparable models. - MiniMaxEmbeddingProvider: Interfaces with the MiniMax API for alternative embedding models.
Configuration propagates through environment variables or CLI flags. The test suite in tests/test_embeddings.py verifies provider selection logic and error handling when invalid model names are supplied.
SQLite Schema and Persistence
When instantiating GraphStore, the _init_db() method executes DDL to create the embeddings table if absent:
CREATE TABLE IF NOT EXISTS embeddings (
qualified_name TEXT PRIMARY KEY,
vector BLOB NOT NULL,
provider TEXT,
model TEXT
);
As nodes are added via store.add_node(qualified_name, source), the system automatically invokes the configured provider to generate a vector, then inserts the serialized NumPy array into the vector column. Integration tests in tests/test_graph.py confirm that embeddings survive transactions and are correctly associated with their graph nodes.
ANN Index Management
The code_review_graph/eval/runner.py module exposes _build_embedding_index, which loads all vectors from the embeddings table into memory and constructs a searchable index. For small codebases, a brute-force cosine similarity scan suffices; larger repositories leverage faiss indexes for millisecond-scale nearest-neighbor lookups. The runner refreshes this index lazily or on-demand via the CLI.
Hybrid Search and Ranking
The public API entry point search.search() accepts a mode parameter that determines the retrieval strategy:
mode="semantic": Computes a query embedding and performs nearest-neighbor lookup against stored vectors using cosine similarity.mode="keyword": Executes FTS against thesearch_indexvirtual table.mode="hybrid": Runs both searches concurrently, then calls_merge_resultsto re-rank using a weighted combination of similarity and TF-IDF scores.
The merging logic resides in code_review_graph/search.py and is validated in tests/test_search.py, ensuring that results from different modalities are normalized and deduplicated before presentation.
End-to-End Workflow
The lifecycle of a semantic query follows five discrete stages:
-
Configuration: The user selects an embedding provider and model via
--embedding-providerand--embedding-modelflags or environment variables. -
Vectorisation: Upon node insertion,
EmbeddingProvider.get_embedding(text)returns a dense NumPy array representing the code entity's semantic meaning. -
Persistence: The vector, provider name, and model version are written to the
embeddingstable ingraph.py, ensuring durability across sessions. -
Index Refresh: The runner calls
_build_embedding_indexto load vectors into an ANN structure, enabling fast similarity queries. -
Query Execution: A natural language string is embedded using the same provider, then compared against the index. If hybrid mode is active, FTS results are retrieved in parallel and fused.
When code entities are removed, the forget logic (tested in tests/test_forget_parity.py) deletes orphaned vectors from the embeddings table to prevent stale data from influencing future searches.
Practical Usage Examples
Configure a local embedding provider via the command line:
code-review-graph --embedding-provider local --embedding-model eval-test
Create a graph and perform semantic search programmatically:
from code_review_graph import GraphStore, search
# Initialize an in-memory store
store = GraphStore(":memory:")
# Add a node—embedding generation happens automatically
store.add_node("my.module.function", source="def foo(): pass")
# Refresh the ANN index
store.refresh_embeddings()
# Execute a natural language query
results = search.search("how to compute a sum", mode="semantic", store=store)
for result in results:
print(result.qualified_name, result.score) # Output: my.module.function 0.92
Execute a hybrid search combining semantic and keyword signals:
# Ensure FTS index exists
store.build_fts()
# Query using both embeddings and full-text
results = search.search("read file", mode="hybrid", store=store)
Summary
- code-review-graph combines vector embeddings with full-text search to enable natural language code discovery.
- Three provider types (Local, OpenAI, MiniMax) generate embeddings via the abstraction layer in
code_review_graph/embeddings.py. - SQLite persistence stores vectors in the
embeddingstable managed byGraphStore._init_db()incode_review_graph/graph.py. - ANN indexing via
_build_embedding_indexincode_review_graph/eval/runner.pyaccelerates similarity queries. - Hybrid dispatch in the search module supports semantic, keyword, and hybrid retrieval modes, merging results with cosine similarity and relevance scoring.
- Lifecycle management ensures embeddings are removed when nodes are forgotten, maintaining index hygiene.
Frequently Asked Questions
What embedding models does code-review-graph support?
The system supports any model accessible through the three provider classes defined in code_review_graph/embeddings.py. You can use local sentence-transformers models via LocalEmbeddingProvider, OpenAI's text-embedding-3-small or text-embedding-ada-002 via OpenAIEmbeddingProvider, or MiniMax embeddings via MiniMaxEmbeddingProvider. Configuration is controlled by the _embedding_provider and _embedding_model settings.
How does the hybrid search mode combine semantic and keyword results?
Hybrid mode executes _embedding_search and _fts_search in parallel, then calls _merge_results to normalize scores. Cosine similarity from vector search and TF-IDF relevance from full-text search are weighted and fused into a single ranked list. This approach captures synonyms via embeddings while preserving exact keyword matches via FTS.
Where are the embedding vectors stored?
Vectors are serialized as BLOBs in the embeddings table within the SQLite database managed by GraphStore. The table schema—defined in GraphStore._init_db() in code_review_graph/graph.py—includes the qualified_name, vector, provider, and model columns to track provenance and enable retrieval.
How does the system handle outdated embeddings when code changes?
When a node is deleted or updated, the forget mechanism triggers cleanup logic that removes the corresponding row from the embeddings table. Integration tests in tests/test_forget_parity.py verify that orphaned vectors are eliminated, ensuring the ANN index only contains embeddings for extant code entities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →