# Embedding Providers in Local Deep Research: Configuration and Usage Guide

> Explore embedding providers for Local Deep Research: sentence-transformers, Ollama, and OpenAI. Learn configuration and usage details from our comprehensive guide.

- Repository: [learningcircuit/local-deep-research](https://github.com/learningcircuit/local-deep-research)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Local Deep Research supports three embedding providers—sentence-transformers, Ollama, and OpenAI—configured centrally via [`src/local_deep_research/embeddings/embeddings_config.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/embeddings/embeddings_config.py) using settings keys like `embeddings.provider` and provider-specific model parameters.**

The `learningcircuit/local-deep-research` repository centralizes all embedding-provider logic in a single configuration module. Understanding which embedding providers are supported and how to configure them enables you to switch seamlessly between local CPU/GPU inference, self-hosted Ollama servers, and cloud-based OpenAI APIs.

## Supported Embedding Providers

Local Deep Research implements three distinct embedding providers, each defined in the `VALID_EMBEDDING_PROVIDERS` constant at [lines 17-21](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/embeddings/embeddings_config.py#L17-L21) of [`src/local_deep_research/embeddings/embeddings_config.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/embeddings/embeddings_config.py):

- **sentence_transformers**: Local embeddings via the *sentence-transformers* library, running on CPU or GPU. Availability is checked via `is_sentence_transformers_available()`, which returns true if the library is installed.
- **ollama**: Local Ollama server integration (e.g., `ollama run llama2`). The `is_ollama_embeddings_available(settings_snapshot)` function verifies that the configured Ollama endpoint is reachable.
- **openai**: OpenAI embeddings API (e.g., `text-embedding-ada-002`). The `is_openai_embeddings_available(settings_snapshot)` function validates that a valid OpenAI API key is present in your settings.

Each provider maps to a concrete implementation class: `SentenceTransformersProvider`, `OllamaEmbeddingsProvider`, or `OpenAIEmbeddingsProvider`.

## Configuration Architecture

When `get_embeddings()` is called, the module resolves the provider from the **`embeddings.provider`** setting (defaulting to `sentence_transformers`). This default-provider lookup occurs at [lines 42-48](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/embeddings/embeddings_config.py#L42-L48):

```python
if provider is None:
    provider = get_setting_from_snapshot(
        "embeddings.provider",
        default="sentence_transformers",
        settings_snapshot=settings_snapshot,
    )

```

The complete configuration schema uses dot-notation keys read transparently from YAML files, environment variables, or Python dictionaries:

| Setting key | Description | Example value |
|-------------|-------------|---------------|
| `embeddings.provider` | Active provider (`sentence_transformers`, `ollama`, `openai`) | `ollama` |
| `embeddings.sentence_transformers.model` | Model name for local inference | `all-MiniLM-L6-v2` |
| `embeddings.ollama.model` | Ollama model tag | `llama2` |
| `embeddings.ollama.url` | HTTP endpoint for Ollama server | `http://localhost:11434` |
| `embeddings.openai.model` | OpenAI embedding model ID | `text-embedding-ada-002` |
| `embeddings.openai.api_key` | API authentication token | `${OPENAI_API_KEY}` |
| `embeddings.openai.base_url` | Optional custom OpenAI-compatible endpoint | `https://api.example.com/v1` |

## How to Configure Embedding Providers

### Via YAML Configuration File

Create or modify your [`settings.yaml`](https://github.com/learningcircuit/local-deep-research/blob/main/settings.yaml) to declare the active provider and its parameters. The library loads these via `SettingsManager` and passes them as a settings snapshot:

```yaml

# settings.yaml

embeddings:
  provider: ollama
  ollama:
    url: http://localhost:11434
    model: llama2
  openai:
    api_key: ${OPENAI_API_KEY}
    model: text-embedding-ada-002
  sentence_transformers:
    model: all-MiniLM-L6-v2

```

When the application initializes, `get_embeddings()` reads this configuration automatically if no explicit arguments override it.

### Programmatic Configuration

Pass a settings snapshot directly to the factory functions for thread-safe, programmatic control. This approach is defined in [`src/local_deep_research/config/thread_settings.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/config/thread_settings.py):

```python
from local_deep_research.embeddings.embeddings_config import get_embeddings

embeddings = get_embeddings(
    provider="openai",
    model="text-embedding-ada-002",
    settings_snapshot=my_settings,  # Thread-safe dict from SettingsManager

)

```

### Runtime Overrides

You can override specific parameters without changing the underlying configuration file. For example, to keep the provider defined in settings but change the model:

```python
from local_deep_research.embeddings.embeddings_config import get_embeddings

# Uses provider from settings, but overrides the model name

embeddings = get_embeddings(model="all-MiniLM-L6-v2")

```

## Using Embeddings in Your Code

### High-Level Helper Function

For quick integration, use `get_embedding_function()` to obtain a callable that accepts a list of strings and returns NumPy arrays:

```python
from local_deep_research.embeddings.embeddings_config import get_embedding_function

embed = get_embedding_function(
    provider="ollama",
    model_name="llama2",
    settings_snapshot=my_settings,
)

vectors = embed(["Hello world", "Local Deep Research"])

```

### Direct LangChain Integration

To obtain a standard LangChain `Embeddings` object compatible with vector stores like FAISS or Chroma:

```python
from local_deep_research.embeddings.embeddings_config import get_embeddings
from langchain_community.vectorstores import FAISS

embeddings = get_embeddings(
    provider="sentence_transformers",
    model="all-MiniLM-L6-v2",
)

vectorstore = FAISS.from_documents(docs, embeddings)

```

## Verifying Provider Availability

Before initializing embeddings, check which providers are currently available in your environment using `get_available_embedding_providers()`:

```python
from local_deep_research.embeddings.embeddings_config import get_available_embedding_providers

available = get_available_embedding_providers(settings_snapshot=my_settings)
print(available)

# Output: {'sentence_transformers': 'Sentence Transformers (Local)', 'ollama': 'Ollama (Local)'}

```

This function internally calls the provider-specific availability checks (`is_sentence_transformers_available()`, `is_ollama_embeddings_available()`, `is_openai_embeddings_available()`) and returns only those providers whose dependencies and configuration requirements are satisfied.

## Summary

- **Three providers are supported**: `sentence_transformers` (local), `ollama` (local server), and `openai` (cloud API), defined in `VALID_EMBEDDING_PROVIDERS` within [`embeddings_config.py`](https://github.com/learningcircuit/local-deep-research/blob/main/embeddings_config.py).
- **Default provider**: `sentence_transformers` is used when `embeddings.provider` is unset, as implemented in the fallback logic at lines 42-48.
- **Configuration keys**: Use `embeddings.provider` to select the backend, and `embeddings.{provider}.model` to specify the model name.
- **Factory functions**: `get_embeddings()` returns a LangChain `Embeddings` object; `get_embedding_function()` returns a simple callable for raw vector generation.
- **Availability checking**: Call `get_available_embedding_providers()` to verify which backends are operational in your current environment.

## Frequently Asked Questions

### Which embedding provider is the default in Local Deep Research?

**`sentence_transformers`** is the default provider. According to the source code in [`src/local_deep_research/embeddings/embeddings_config.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/embeddings/embeddings_config.py), the `get_setting_from_snapshot()` call at lines 42-48 explicitly defaults to `"sentence_transformers"` when no provider is specified in your settings.

### How do I switch from sentence-transformers to Ollama without changing code?

Set `embeddings.provider: ollama` in your [`settings.yaml`](https://github.com/learningcircuit/local-deep-research/blob/main/settings.yaml) file and ensure `embeddings.ollama.url` points to your running server (default is `http://localhost:11434`). The `get_embeddings()` function reads this value automatically, so no code changes are required unless you need to override the default model.

### Can I use a custom OpenAI-compatible endpoint for embeddings?

Yes. Configure `embeddings.openai.base_url` in your settings file or settings snapshot. The `OpenAIEmbeddingsProvider` implementation reads this key to redirect API calls to compatible servers such as LocalAI or custom OpenAI proxies, while still using standard OpenAI client libraries.

### What file contains the list of valid embedding providers?

The canonical list resides in [`src/local_deep_research/embeddings/embeddings_config.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/embeddings/embeddings_config.py) within the `VALID_EMBEDDING_PROVIDERS` constant (lines 17-21). This constant defines the supported string keys (`sentence_transformers`, `ollama`, `openai`) that you can pass to `get_embeddings()` or set in your configuration files.