# How code-review-graph Enables Semantic Search Using Embeddings for Natural Language Queries

> Discover how code-review-graph uses embeddings and semantic search to answer natural language queries on code. Learn about its hybrid retrieval system and vector embeddings.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: how-to-guide
- Published: 2026-08-15

---

**code-review-graph implements a hybrid retrieval system that transforms code entities into dense vector embeddings using configurable providers, persists them in SQLite, and executes cosine-similarity search to answer natural-language queries alongside traditional full-text search.**

The open-source repository `tirth8205/code-review-graph` provides a graph-based code analysis tool that extends beyond syntactic structure to support **semantic search using embeddings**. By converting qualified names, docstrings, and comments into high-dimensional vectors, the system enables developers to locate relevant code using plain English descriptions rather than exact keyword matches.

## The Architecture of Semantic Search

The semantic search capability rests on a five-stage pipeline that bridges natural language and code representation. The architecture separates concerns between vector generation, persistent storage, approximate nearest neighbor (ANN) indexing, and result fusion.

**Embedding Providers** wrap external or local models to perform text-to-vector conversion. The codebase defines three concrete implementations: `LocalEmbeddingProvider`, `OpenAIEmbeddingProvider`, and `MiniMaxEmbeddingProvider`. Selection occurs via configuration keys `_embedding_provider` and `_embedding_model`, validated in [`tests/test_embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_embeddings.py).

**Vector Storage** uses a dedicated SQLite table created within `GraphStore._init_db()`. The schema stores `qualified_name` as the primary key alongside a binary `vector` blob, plus metadata columns `provider` and `model` to track which model generated each embedding.

**Index Construction** happens through `_build_embedding_index` in [`code_review_graph/eval/runner.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/runner.py). This function materializes an ANN index—backed by `faiss` or brute-force scanning—whenever new nodes are added or model configurations change.

**Search Dispatch** routes queries through `search.search(query, mode=...)` in the search module. The dispatcher invokes `_embedding_search` for vector similarity, `_fts_search` for keyword matching, or both when operating in hybrid mode.

**Result Ranking** merges outputs via `_merge_results`, combining cosine similarity scores from embeddings with relevance scores from the full-text search (FTS) engine to produce a unified ranking.

## Core Implementation Components

### Embedding Providers and Configuration

The system abstracts model interactions behind the provider interface defined in [`code_review_graph/embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py). Concrete classes handle authentication, batching, and dimensionality normalization:

- **LocalEmbeddingProvider**: Runs inference locally using sentence-transformers or similar frameworks.
- **OpenAIEmbeddingProvider**: Calls the OpenAI API for `text-embedding-3-small` or comparable models.
- **MiniMaxEmbeddingProvider**: Interfaces with the MiniMax API for alternative embedding models.

Configuration propagates through environment variables or CLI flags. The test suite in [`tests/test_embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_embeddings.py) verifies provider selection logic and error handling when invalid model names are supplied.

### SQLite Schema and Persistence

When instantiating `GraphStore`, the `_init_db()` method executes DDL to create the `embeddings` table if absent:

```sql
CREATE TABLE IF NOT EXISTS embeddings (
    qualified_name TEXT PRIMARY KEY,
    vector BLOB NOT NULL,
    provider TEXT,
    model TEXT
);

```

As nodes are added via `store.add_node(qualified_name, source)`, the system automatically invokes the configured provider to generate a vector, then inserts the serialized NumPy array into the `vector` column. Integration tests in [`tests/test_graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_graph.py) confirm that embeddings survive transactions and are correctly associated with their graph nodes.

### ANN Index Management

The [`code_review_graph/eval/runner.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/runner.py) module exposes `_build_embedding_index`, which loads all vectors from the `embeddings` table into memory and constructs a searchable index. For small codebases, a brute-force cosine similarity scan suffices; larger repositories leverage `faiss` indexes for millisecond-scale nearest-neighbor lookups. The runner refreshes this index lazily or on-demand via the CLI.

### Hybrid Search and Ranking

The public API entry point `search.search()` accepts a `mode` parameter that determines the retrieval strategy:

- **`mode="semantic"`**: Computes a query embedding and performs nearest-neighbor lookup against stored vectors using cosine similarity.
- **`mode="keyword"`**: Executes FTS against the `search_index` virtual table.
- **`mode="hybrid"`**: Runs both searches concurrently, then calls `_merge_results` to re-rank using a weighted combination of similarity and TF-IDF scores.

The merging logic resides in [`code_review_graph/search.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py) and is validated in [`tests/test_search.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_search.py), ensuring that results from different modalities are normalized and deduplicated before presentation.

## End-to-End Workflow

The lifecycle of a semantic query follows five discrete stages:

1. **Configuration**: The user selects an embedding provider and model via `--embedding-provider` and `--embedding-model` flags or environment variables.

2. **Vectorisation**: Upon node insertion, `EmbeddingProvider.get_embedding(text)` returns a dense NumPy array representing the code entity's semantic meaning.

3. **Persistence**: The vector, provider name, and model version are written to the `embeddings` table in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py), ensuring durability across sessions.

4. **Index Refresh**: The runner calls `_build_embedding_index` to load vectors into an ANN structure, enabling fast similarity queries.

5. **Query Execution**: A natural language string is embedded using the same provider, then compared against the index. If hybrid mode is active, FTS results are retrieved in parallel and fused.

When code entities are removed, the `forget` logic (tested in [`tests/test_forget_parity.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_forget_parity.py)) deletes orphaned vectors from the `embeddings` table to prevent stale data from influencing future searches.

## Practical Usage Examples

Configure a local embedding provider via the command line:

```bash
code-review-graph --embedding-provider local --embedding-model eval-test

```

Create a graph and perform semantic search programmatically:

```python
from code_review_graph import GraphStore, search

# Initialize an in-memory store

store = GraphStore(":memory:")

# Add a node—embedding generation happens automatically

store.add_node("my.module.function", source="def foo(): pass")

# Refresh the ANN index

store.refresh_embeddings()

# Execute a natural language query

results = search.search("how to compute a sum", mode="semantic", store=store)

for result in results:
    print(result.qualified_name, result.score)  # Output: my.module.function 0.92

```

Execute a hybrid search combining semantic and keyword signals:

```python

# Ensure FTS index exists

store.build_fts()

# Query using both embeddings and full-text

results = search.search("read file", mode="hybrid", store=store)

```

## Summary

- **code-review-graph** combines vector embeddings with full-text search to enable natural language code discovery.
- **Three provider types** (Local, OpenAI, MiniMax) generate embeddings via the abstraction layer in [`code_review_graph/embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py).
- **SQLite persistence** stores vectors in the `embeddings` table managed by `GraphStore._init_db()` in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py).
- **ANN indexing** via `_build_embedding_index` in [`code_review_graph/eval/runner.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/runner.py) accelerates similarity queries.
- **Hybrid dispatch** in the search module supports semantic, keyword, and hybrid retrieval modes, merging results with cosine similarity and relevance scoring.
- **Lifecycle management** ensures embeddings are removed when nodes are forgotten, maintaining index hygiene.

## Frequently Asked Questions

### What embedding models does code-review-graph support?

The system supports any model accessible through the three provider classes defined in [`code_review_graph/embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py). You can use local sentence-transformers models via `LocalEmbeddingProvider`, OpenAI's `text-embedding-3-small` or `text-embedding-ada-002` via `OpenAIEmbeddingProvider`, or MiniMax embeddings via `MiniMaxEmbeddingProvider`. Configuration is controlled by the `_embedding_provider` and `_embedding_model` settings.

### How does the hybrid search mode combine semantic and keyword results?

Hybrid mode executes `_embedding_search` and `_fts_search` in parallel, then calls `_merge_results` to normalize scores. Cosine similarity from vector search and TF-IDF relevance from full-text search are weighted and fused into a single ranked list. This approach captures synonyms via embeddings while preserving exact keyword matches via FTS.

### Where are the embedding vectors stored?

Vectors are serialized as BLOBs in the `embeddings` table within the SQLite database managed by `GraphStore`. The table schema—defined in `GraphStore._init_db()` in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py)—includes the `qualified_name`, `vector`, `provider`, and `model` columns to track provenance and enable retrieval.

### How does the system handle outdated embeddings when code changes?

When a node is deleted or updated, the `forget` mechanism triggers cleanup logic that removes the corresponding row from the `embeddings` table. Integration tests in [`tests/test_forget_parity.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_forget_parity.py) verify that orphaned vectors are eliminated, ensuring the ANN index only contains embeddings for extant code entities.