Which Embedding Models Are Supported by Code-Review-Graph for Semantic Search
The tirth8205/code-review-graph repository supports five distinct embedding providers—Local, Google, MiniMax, OpenAI-compatible, and Voyage AI—each configurable via the EmbeddingProvider interface and get_provider() factory in code_review_graph/embeddings.py.
Configuring semantic search for code reviews requires selecting the right embedding model for your infrastructure. The project implements a flexible provider architecture that supports both offline sentence-transformers and commercial API endpoints, ensuring you can run semantic search locally or at scale without modifying core logic.
The EmbeddingProvider Architecture
All embedding capabilities are implemented through a pluggable interface defined in code_review_graph/embeddings.py. The abstract base class specifies four critical methods: embed(), embed_query(), dimension(), and name(). This uniformity allows the EmbeddingStore (lines 557‑749) to treat every provider identically, regardless of whether it runs locally or calls a remote API.
The concrete provider classes—LocalEmbeddingProvider, GoogleEmbeddingProvider, MiniMaxEmbeddingProvider, OpenAIEmbeddingProvider, and VoyageEmbeddingProvider—occupy lines 13‑329 of the same file. A factory function get_provider() (lines 72‑124) validates environment variables and instantiates the appropriate class based on a user-supplied name.
Supported Embedding Providers
Code-review-graph ships with five built-in providers, each optimized for different deployment scenarios.
Local Provider (Sentence-Transformers)
The LocalEmbeddingProvider enables fully offline semantic search using the Hugging Face sentence-transformers library. By default, it loads all-MiniLM-L6-v2, though you can override this via the CRG_EMBEDDING_MODEL environment variable or the model parameter. No API keys are required.
- Default model:
all-MiniLM-L6-v2 - Optional env var:
CRG_EMBEDDING_MODEL
Google Gemini Embeddings
The GoogleEmbeddingProvider interfaces with Google's Gemini API to generate cloud-based embeddings. It defaults to gemini-embedding-001, but accepts alternate model names via configuration.
- Required env var:
GOOGLE_API_KEY - Default model:
gemini-embedding-001
MiniMax Embeddings
The MiniMaxEmbeddingProvider offers high-quality 1536-dimensional vectors via MiniMax's "embo-01" model. Unlike other providers, this uses a fixed model and requires dedicated API credentials.
- Required env var:
MINIMAX_API_KEY - Fixed model:
embo-01
OpenAI-Compatible Provider
The OpenAIEmbeddingProvider is the most flexible option, supporting any service that adheres to the OpenAI /v1/embeddings schema. This includes OpenAI proper, Azure OpenAI, LiteLLM, vLLM, LocalAI, and Ollama.
- Required env vars:
CRG_OPENAI_API_KEY,CRG_OPENAI_BASE_URL(optional for standard OpenAI) - Model configuration:
CRG_OPENAI_MODEL(e.g.,text-embedding-3-small)
Voyage AI Provider
The VoyageEmbeddingProvider specializes in code-retrieval tasks using Voyage AI's models. It defaults to voyage-code-3 but supports model customization through the model argument.
- Required env var:
VOYAGE_API_KEY - Default model:
voyage-code-3
Provider Selection and Configuration
Switching between embedding models is handled by the get_provider() factory at lines 72‑124. The function maps case-insensitive strings ("local", "google", "minimax", "openai", "voyage") to their respective classes and validates that required environment variables are present.
from code_review_graph.embeddings import get_provider, EmbeddingStore
# 1️⃣ Choose a provider by name (case-insensitive)
provider = get_provider(provider="openai") # picks OpenAI if env vars are set
# provider is an instance of OpenAIEmbeddingProvider
# 2️⃣ Use the provider directly
texts = ["def add(a, b): return a + b", "class User: pass"]
vectors = provider.embed(texts) # → List[List[float]]
# 3️⃣ Store embeddings for graph nodes
store = EmbeddingStore(db_path="my_graph.db", provider="openai")
store.embed_nodes(nodes) # automatically uses the chosen provider
# 4️⃣ Perform a semantic search
results = store.search("how to add two numbers", limit=5)
print(results) # [('my_module.add', 0.93), …]
Switching providers requires only changing the argument or environment variables and re-running the embedding step:
# Use the local provider (no API key needed)
local = get_provider(provider="local", model="all-MiniLM-L6-v2")
store = EmbeddingStore(db_path="my_graph.db", provider="local")
store.embed_nodes(nodes) # re-embeds with the local model
Storing and Searching Embeddings
The EmbeddingStore class (lines 557‑749) manages SQLite persistence and semantic retrieval. It stores vectors keyed by the provider's unique name attribute, ensuring that changing providers triggers a fresh embedding generation rather than mixing incompatible vector spaces.
Because the store selects embeddings by the provider's unique name value, switching to a new model automatically invalidates old vectors and recomputes them for the new embedding space. This prevents dimensionality mismatches when migrating from, for example, the 384-dimensional MiniLM outputs to 1536-dimensional OpenAI vectors.
Summary
- Five providers are supported: Local (sentence-transformers), Google Gemini, MiniMax, OpenAI-compatible, and Voyage AI.
- Configuration occurs via the
get_provider()factory incode_review_graph/embeddings.py(lines 72‑124). - Environment variables control API access:
GOOGLE_API_KEY,MINIMAX_API_KEY,CRG_OPENAI_API_KEY,VOYAGE_API_KEY, and optionallyCRG_EMBEDDING_MODELfor local overrides. - The
EmbeddingStore(lines 557‑749) isolates vector spaces by provider name, automatically handling provider switches and preventing vector contamination.
Frequently Asked Questions
Can I use code-review-graph without an internet connection?
Yes. Select the LocalEmbeddingProvider by setting provider="local" and ensuring the sentence-transformers package is installed. This downloads model weights (e.g., all-MiniLM-L6-v2) locally and requires no API keys or external network calls during embedding generation or search.
How do I migrate from OpenAI to Azure OpenAI?
Use the OpenAI-compatible provider and export your Azure configuration:
export CRG_OPENAI_API_KEY="your-azure-key"
export CRG_OPENAI_BASE_URL="https://your-resource.openai.azure.com/openai/deployments/your-deployment"
export CRG_OPENAI_MODEL="your-deployment-name"
Then instantiate with get_provider(provider="openai"). The underlying client automatically routes to your Azure endpoint while maintaining the same interface as standard OpenAI.
What happens if I change embedding models mid-project?
The EmbeddingStore stores embeddings keyed by the provider's name property. When you switch models or providers, the store detects a new unique name and generates fresh embeddings for your codebase. Previous vectors remain in the SQLite database but are ignored for search, preventing dimension mismatch errors or degraded retrieval quality from mixed embedding spaces.
Which provider offers the best quality for code search?
Voyage AI (specifically voyage-code-3) provides embeddings optimized explicitly for code retrieval. For general-purpose use without external dependencies, OpenAI (text-embedding-3-large) offers high performance, while the Local provider with all-MiniLM-L6-v2 delivers the best latency and zero cost for smaller repositories or development environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →