# Which Embedding Models Are Supported by Code-Review-Graph for Semantic Search

> Explore supported embedding models for semantic search in code-review-graph. Discover options like Local, Google, MiniMax, OpenAI-compatible, and Voyage AI.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: api-reference
- Published: 2026-08-13

---

**The tirth8205/code-review-graph repository supports five distinct embedding providers—Local, Google, MiniMax, OpenAI-compatible, and Voyage AI—each configurable via the `EmbeddingProvider` interface and `get_provider()` factory in [`code_review_graph/embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py).**

Configuring semantic search for code reviews requires selecting the right embedding model for your infrastructure. The project implements a flexible provider architecture that supports both offline sentence-transformers and commercial API endpoints, ensuring you can run semantic search locally or at scale without modifying core logic.

## The EmbeddingProvider Architecture

All embedding capabilities are implemented through a pluggable interface defined in [`code_review_graph/embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py). The abstract base class specifies four critical methods: `embed()`, `embed_query()`, `dimension()`, and `name()`. This uniformity allows the `EmbeddingStore` (lines 557‑749) to treat every provider identically, regardless of whether it runs locally or calls a remote API.

The concrete provider classes—`LocalEmbeddingProvider`, `GoogleEmbeddingProvider`, `MiniMaxEmbeddingProvider`, `OpenAIEmbeddingProvider`, and `VoyageEmbeddingProvider`—occupy lines 13‑329 of the same file. A factory function `get_provider()` (lines 72‑124) validates environment variables and instantiates the appropriate class based on a user-supplied name.

## Supported Embedding Providers

Code-review-graph ships with five built-in providers, each optimized for different deployment scenarios.

### Local Provider (Sentence-Transformers)

The **LocalEmbeddingProvider** enables fully offline semantic search using the Hugging Face `sentence-transformers` library. By default, it loads `all-MiniLM-L6-v2`, though you can override this via the `CRG_EMBEDDING_MODEL` environment variable or the `model` parameter. No API keys are required.

- **Default model**: `all-MiniLM-L6-v2`
- **Optional env var**: `CRG_EMBEDDING_MODEL`

### Google Gemini Embeddings

The **GoogleEmbeddingProvider** interfaces with Google's Gemini API to generate cloud-based embeddings. It defaults to `gemini-embedding-001`, but accepts alternate model names via configuration.

- **Required env var**: `GOOGLE_API_KEY`
- **Default model**: `gemini-embedding-001`

### MiniMax Embeddings

The **MiniMaxEmbeddingProvider** offers high-quality 1536-dimensional vectors via MiniMax's "embo-01" model. Unlike other providers, this uses a fixed model and requires dedicated API credentials.

- **Required env var**: `MINIMAX_API_KEY`
- **Fixed model**: `embo-01`

### OpenAI-Compatible Provider

The **OpenAIEmbeddingProvider** is the most flexible option, supporting any service that adheres to the OpenAI `/v1/embeddings` schema. This includes OpenAI proper, Azure OpenAI, LiteLLM, vLLM, LocalAI, and Ollama.

- **Required env vars**: `CRG_OPENAI_API_KEY`, `CRG_OPENAI_BASE_URL` (optional for standard OpenAI)
- **Model configuration**: `CRG_OPENAI_MODEL` (e.g., `text-embedding-3-small`)

### Voyage AI Provider

The **VoyageEmbeddingProvider** specializes in code-retrieval tasks using Voyage AI's models. It defaults to `voyage-code-3` but supports model customization through the `model` argument.

- **Required env var**: `VOYAGE_API_KEY`
- **Default model**: `voyage-code-3`

## Provider Selection and Configuration

Switching between embedding models is handled by the `get_provider()` factory at lines 72‑124. The function maps case-insensitive strings (`"local"`, `"google"`, `"minimax"`, `"openai"`, `"voyage"`) to their respective classes and validates that required environment variables are present.

```python
from code_review_graph.embeddings import get_provider, EmbeddingStore

# 1️⃣ Choose a provider by name (case-insensitive)

provider = get_provider(provider="openai")   # picks OpenAI if env vars are set

# provider is an instance of OpenAIEmbeddingProvider

# 2️⃣ Use the provider directly

texts = ["def add(a, b): return a + b", "class User: pass"]
vectors = provider.embed(texts)               # → List[List[float]]

# 3️⃣ Store embeddings for graph nodes

store = EmbeddingStore(db_path="my_graph.db", provider="openai")
store.embed_nodes(nodes)                      # automatically uses the chosen provider

# 4️⃣ Perform a semantic search

results = store.search("how to add two numbers", limit=5)
print(results)  # [('my_module.add', 0.93), …]

```

Switching providers requires only changing the argument or environment variables and re-running the embedding step:

```python

# Use the local provider (no API key needed)

local = get_provider(provider="local", model="all-MiniLM-L6-v2")
store = EmbeddingStore(db_path="my_graph.db", provider="local")
store.embed_nodes(nodes)    # re-embeds with the local model

```

## Storing and Searching Embeddings

The **EmbeddingStore** class (lines 557‑749) manages SQLite persistence and semantic retrieval. It stores vectors keyed by the provider's unique `name` attribute, ensuring that changing providers triggers a fresh embedding generation rather than mixing incompatible vector spaces.

Because the store selects embeddings by the provider's unique `name` value, switching to a new model automatically invalidates old vectors and recomputes them for the new embedding space. This prevents dimensionality mismatches when migrating from, for example, the 384-dimensional MiniLM outputs to 1536-dimensional OpenAI vectors.

## Summary

- **Five providers** are supported: Local (sentence-transformers), Google Gemini, MiniMax, OpenAI-compatible, and Voyage AI.
- **Configuration** occurs via the `get_provider()` factory in [`code_review_graph/embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py) (lines 72‑124).
- **Environment variables** control API access: `GOOGLE_API_KEY`, `MINIMAX_API_KEY`, `CRG_OPENAI_API_KEY`, `VOYAGE_API_KEY`, and optionally `CRG_EMBEDDING_MODEL` for local overrides.
- The **`EmbeddingStore`** (lines 557‑749) isolates vector spaces by provider name, automatically handling provider switches and preventing vector contamination.

## Frequently Asked Questions

### Can I use code-review-graph without an internet connection?

Yes. Select the **LocalEmbeddingProvider** by setting `provider="local"` and ensuring the `sentence-transformers` package is installed. This downloads model weights (e.g., `all-MiniLM-L6-v2`) locally and requires no API keys or external network calls during embedding generation or search.

### How do I migrate from OpenAI to Azure OpenAI?

Use the **OpenAI-compatible provider** and export your Azure configuration:

```bash
export CRG_OPENAI_API_KEY="your-azure-key"
export CRG_OPENAI_BASE_URL="https://your-resource.openai.azure.com/openai/deployments/your-deployment"
export CRG_OPENAI_MODEL="your-deployment-name"

```

Then instantiate with `get_provider(provider="openai")`. The underlying client automatically routes to your Azure endpoint while maintaining the same interface as standard OpenAI.

### What happens if I change embedding models mid-project?

The `EmbeddingStore` stores embeddings keyed by the provider's `name` property. When you switch models or providers, the store detects a new unique name and generates fresh embeddings for your codebase. Previous vectors remain in the SQLite database but are ignored for search, preventing dimension mismatch errors or degraded retrieval quality from mixed embedding spaces.

### Which provider offers the best quality for code search?

**Voyage AI** (specifically `voyage-code-3`) provides embeddings optimized explicitly for code retrieval. For general-purpose use without external dependencies, **OpenAI** (`text-embedding-3-large`) offers high performance, while the **Local** provider with `all-MiniLM-L6-v2` delivers the best latency and zero cost for smaller repositories or development environments.