Embedding Providers in Local Deep Research: Configuration and Usage Guide
Local Deep Research supports three embedding providers—sentence-transformers, Ollama, and OpenAI—configured centrally via src/local_deep_research/embeddings/embeddings_config.py using settings keys like embeddings.provider and provider-specific model parameters.
The learningcircuit/local-deep-research repository centralizes all embedding-provider logic in a single configuration module. Understanding which embedding providers are supported and how to configure them enables you to switch seamlessly between local CPU/GPU inference, self-hosted Ollama servers, and cloud-based OpenAI APIs.
Supported Embedding Providers
Local Deep Research implements three distinct embedding providers, each defined in the VALID_EMBEDDING_PROVIDERS constant at lines 17-21 of src/local_deep_research/embeddings/embeddings_config.py:
- sentence_transformers: Local embeddings via the sentence-transformers library, running on CPU or GPU. Availability is checked via
is_sentence_transformers_available(), which returns true if the library is installed. - ollama: Local Ollama server integration (e.g.,
ollama run llama2). Theis_ollama_embeddings_available(settings_snapshot)function verifies that the configured Ollama endpoint is reachable. - openai: OpenAI embeddings API (e.g.,
text-embedding-ada-002). Theis_openai_embeddings_available(settings_snapshot)function validates that a valid OpenAI API key is present in your settings.
Each provider maps to a concrete implementation class: SentenceTransformersProvider, OllamaEmbeddingsProvider, or OpenAIEmbeddingsProvider.
Configuration Architecture
When get_embeddings() is called, the module resolves the provider from the embeddings.provider setting (defaulting to sentence_transformers). This default-provider lookup occurs at lines 42-48:
if provider is None:
provider = get_setting_from_snapshot(
"embeddings.provider",
default="sentence_transformers",
settings_snapshot=settings_snapshot,
)
The complete configuration schema uses dot-notation keys read transparently from YAML files, environment variables, or Python dictionaries:
| Setting key | Description | Example value |
|---|---|---|
embeddings.provider |
Active provider (sentence_transformers, ollama, openai) |
ollama |
embeddings.sentence_transformers.model |
Model name for local inference | all-MiniLM-L6-v2 |
embeddings.ollama.model |
Ollama model tag | llama2 |
embeddings.ollama.url |
HTTP endpoint for Ollama server | http://localhost:11434 |
embeddings.openai.model |
OpenAI embedding model ID | text-embedding-ada-002 |
embeddings.openai.api_key |
API authentication token | ${OPENAI_API_KEY} |
embeddings.openai.base_url |
Optional custom OpenAI-compatible endpoint | https://api.example.com/v1 |
How to Configure Embedding Providers
Via YAML Configuration File
Create or modify your settings.yaml to declare the active provider and its parameters. The library loads these via SettingsManager and passes them as a settings snapshot:
# settings.yaml
embeddings:
provider: ollama
ollama:
url: http://localhost:11434
model: llama2
openai:
api_key: ${OPENAI_API_KEY}
model: text-embedding-ada-002
sentence_transformers:
model: all-MiniLM-L6-v2
When the application initializes, get_embeddings() reads this configuration automatically if no explicit arguments override it.
Programmatic Configuration
Pass a settings snapshot directly to the factory functions for thread-safe, programmatic control. This approach is defined in src/local_deep_research/config/thread_settings.py:
from local_deep_research.embeddings.embeddings_config import get_embeddings
embeddings = get_embeddings(
provider="openai",
model="text-embedding-ada-002",
settings_snapshot=my_settings, # Thread-safe dict from SettingsManager
)
Runtime Overrides
You can override specific parameters without changing the underlying configuration file. For example, to keep the provider defined in settings but change the model:
from local_deep_research.embeddings.embeddings_config import get_embeddings
# Uses provider from settings, but overrides the model name
embeddings = get_embeddings(model="all-MiniLM-L6-v2")
Using Embeddings in Your Code
High-Level Helper Function
For quick integration, use get_embedding_function() to obtain a callable that accepts a list of strings and returns NumPy arrays:
from local_deep_research.embeddings.embeddings_config import get_embedding_function
embed = get_embedding_function(
provider="ollama",
model_name="llama2",
settings_snapshot=my_settings,
)
vectors = embed(["Hello world", "Local Deep Research"])
Direct LangChain Integration
To obtain a standard LangChain Embeddings object compatible with vector stores like FAISS or Chroma:
from local_deep_research.embeddings.embeddings_config import get_embeddings
from langchain_community.vectorstores import FAISS
embeddings = get_embeddings(
provider="sentence_transformers",
model="all-MiniLM-L6-v2",
)
vectorstore = FAISS.from_documents(docs, embeddings)
Verifying Provider Availability
Before initializing embeddings, check which providers are currently available in your environment using get_available_embedding_providers():
from local_deep_research.embeddings.embeddings_config import get_available_embedding_providers
available = get_available_embedding_providers(settings_snapshot=my_settings)
print(available)
# Output: {'sentence_transformers': 'Sentence Transformers (Local)', 'ollama': 'Ollama (Local)'}
This function internally calls the provider-specific availability checks (is_sentence_transformers_available(), is_ollama_embeddings_available(), is_openai_embeddings_available()) and returns only those providers whose dependencies and configuration requirements are satisfied.
Summary
- Three providers are supported:
sentence_transformers(local),ollama(local server), andopenai(cloud API), defined inVALID_EMBEDDING_PROVIDERSwithinembeddings_config.py. - Default provider:
sentence_transformersis used whenembeddings.provideris unset, as implemented in the fallback logic at lines 42-48. - Configuration keys: Use
embeddings.providerto select the backend, andembeddings.{provider}.modelto specify the model name. - Factory functions:
get_embeddings()returns a LangChainEmbeddingsobject;get_embedding_function()returns a simple callable for raw vector generation. - Availability checking: Call
get_available_embedding_providers()to verify which backends are operational in your current environment.
Frequently Asked Questions
Which embedding provider is the default in Local Deep Research?
sentence_transformers is the default provider. According to the source code in src/local_deep_research/embeddings/embeddings_config.py, the get_setting_from_snapshot() call at lines 42-48 explicitly defaults to "sentence_transformers" when no provider is specified in your settings.
How do I switch from sentence-transformers to Ollama without changing code?
Set embeddings.provider: ollama in your settings.yaml file and ensure embeddings.ollama.url points to your running server (default is http://localhost:11434). The get_embeddings() function reads this value automatically, so no code changes are required unless you need to override the default model.
Can I use a custom OpenAI-compatible endpoint for embeddings?
Yes. Configure embeddings.openai.base_url in your settings file or settings snapshot. The OpenAIEmbeddingsProvider implementation reads this key to redirect API calls to compatible servers such as LocalAI or custom OpenAI proxies, while still using standard OpenAI client libraries.
What file contains the list of valid embedding providers?
The canonical list resides in src/local_deep_research/embeddings/embeddings_config.py within the VALID_EMBEDDING_PROVIDERS constant (lines 17-21). This constant defines the supported string keys (sentence_transformers, ollama, openai) that you can pass to get_embeddings() or set in your configuration files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →