# Configuring Embedding Models in Private-GPT: HuggingFace, OpenAI, Ollama, and Mistral Options

> Explore embedding model options for Private-GPT. Learn to configure HuggingFace, OpenAI, Ollama, and Mistral backends using settings.yaml for seamless integration.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: configuration
- Published: 2026-03-06

---

**Private-GPT uses a single `embedding.mode` setting to switch between embedding backends including HuggingFace, OpenAI, Ollama, and Mistral, with each provider configured through dedicated YAML blocks in [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml).**

Private-GPT is designed with a flexible embedding architecture that decouples the retrieval logic from the underlying vectorization service. By modifying the **embedding model configuration** in the central settings file, you can swap between local transformer models, commercial APIs, and self-hosted solutions without changing application code.

## Understanding Embedding Configuration Architecture

The configuration system centers on the **`EmbeddingSettings`** model defined in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) (lines 998–1008). This Pydantic model validates the `embedding.mode` field against a strict enum of supported backends and maps each mode to its required credential fields.

When the application initializes, the **`EmbeddingComponent`** factory class located in [`private_gpt/components/embedding/embedding_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/embedding/embedding_component.py) (lines 18–67) reads this setting and instantiates the appropriate LlamaIndex embedding class using a Python `match` statement.

## Available Embedding Modes

Private-GPT supports eight distinct embedding providers via the `mode` field. The primary options requested—HuggingFace, OpenAI, Ollama, and Mistral—are configured as follows:

- **`huggingface`** – Local inference using Hugging Face transformer models
- **`openai`** – OpenAI API or compatible self-hosted endpoints
- **`ollama`** – Local Ollama server for embedding generation
- **`mistralai`** – Mistral AI's embedding endpoint (OpenAI-compatible)
- **`azopenai`** – Azure OpenAI Service (enterprise deployments)
- **`sagemaker`** – Amazon SageMaker hosted endpoints
- **`gemini`** – Google Gemini API
- **`mock`** – Fixed-dimension random vectors for testing

## Provider-Specific Configuration Options

Each embedding backend requires distinct authentication and model parameters defined under separate YAML keys.

### HuggingFace (Local Models)

For **local embedding execution**, set `embedding.mode` to `huggingface` and specify the model identifier. The system downloads weights to `models_cache_path` and runs inference locally.

```yaml
embedding:
  mode: huggingface

huggingface:
  embedding_hf_model_name: sentence-transformers/all-MiniLM-L6-v2
  trust_remote_code: false
  access_token: ${HUGGINGFACE_TOKEN}  # Optional, for gated models

```

Key parameters include `embedding_hf_model_name` (the HuggingFace Hub model ID) and `trust_remote_code` (boolean for custom model architectures).

### OpenAI (Cloud API)

To use **OpenAI embeddings**, configure the API credentials and model name. The component constructs requests to the `/embeddings` endpoint.

```yaml
embedding:
  mode: openai

openai:
  api_key: ${OPENAI_API_KEY}
  embedding_model: text-embedding-3-small
  embedding_api_base: https://api.openai.com/v1  # Optional override

```

The `embedding_api_base` field supports proxy servers or OpenAI-compatible alternatives like LocalAI.

### Ollama (Self-Hosted)

The **Ollama integration** connects to a locally running Ollama instance (default `http://localhost:11434`). Enable `autopull_models` to automatically fetch the specified model tag if not present.

```yaml
embedding:
  mode: ollama

ollama:
  embedding_api_base: http://localhost:11434
  embedding_model: nomic-embed-text
  autopull_models: true

```

This mode utilizes the `OllamaEmbedding` class from LlamaIndex and validates server reachability before ingestion begins.

### Mistral AI (OpenAI-Compatible)

**Mistral AI** leverages the same underlying client as OpenAI but targets Mistral's endpoint. Configure it using the standard OpenAI credential fields while setting `mode: mistralai`.

```yaml
embedding:
  mode: mistralai

openai:
  api_key: ${MISTRAL_API_KEY}
  embedding_model: mistral-embed
  embedding_api_base: https://api.mistral.ai/v1

```

As implemented in [`embedding_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/embedding_component.py), this mode instantiates `MistralAIEmbedding` but pulls configuration from the shared `openai` settings block.

## Runtime Component Instantiation

The selection logic in [`private_gpt/components/embedding/embedding_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/embedding/embedding_component.py) uses structural pattern matching to bind configuration to implementation:

```python

# Simplified excerpt from embedding_component.py

match settings.embedding.mode:
    case "huggingface":
        from llama_index.embeddings.huggingface import HuggingFaceEmbedding
        embedding_model = HuggingFaceEmbedding(
            model_name=settings.huggingface.embedding_hf_model_name
        )
    case "openai":
        from llama_index.embeddings.openai import OpenAIEmbedding
        embedding_model = OpenAIEmbedding(
            api_key=settings.openai.api_key,
            model=settings.openai.embedding_model
        )
    case "ollama":
        from llama_index.embeddings.ollama import OllamaEmbedding
        embedding_model = OllamaEmbedding(
            base_url=settings.ollama.embedding_api_base,
            model_name=settings.ollama.embedding_model
        )
    case "mistralai":
        from llama_index.embeddings.mistralai import MistralAIEmbedding
        embedding_model = MistralAIEmbedding(
            api_key=settings.openai.api_key,
            model=settings.openai.embedding_model
        )

```

This factory approach ensures that changing the embedding provider requires only a configuration edit, with no code modifications.

## Practical Implementation Example

Load settings and generate embeddings programmatically using the component:

```python
from private_gpt.components.embedding.embedding_component import EmbeddingComponent
from private_gpt.settings.settings_loader import SettingsLoader

# Load configuration from private_gpt/settings.yaml

settings = SettingsLoader().load()

# Initialize the embedding component

component = EmbeddingComponent(settings)

# Generate embeddings

texts = ["Private-GPT supports multiple embedding providers"]
vectors = component.embedding_model.get_text_embedding_batch(texts)

print(f"Generated {len(vectors)} embeddings of dimension {len(vectors[0])}")

```

## Summary

- **Centralized Configuration**: The `embedding.mode` field in [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) controls the backend selection.
- **File Locations**: Settings are defined in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) (lines 998–1008) and instantiated in [`private_gpt/components/embedding/embedding_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/embedding/embedding_component.py) (lines 18–67).
- **Local Options**: HuggingFace and Ollama provide fully local embedding without external API dependencies.
- **Cloud Options**: OpenAI, Mistral, Azure OpenAI, and Gemini require API keys and support high-throughput managed inference.
- **Testing**: The `mock` mode generates deterministic random vectors for CI/CD pipelines.

## Frequently Asked Questions

### How do I switch from OpenAI to a local HuggingFace model?

Change `embedding.mode` from `openai` to `huggingface` in [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml), then populate the `huggingface` block with your chosen `embedding_hf_model_name`. Ensure sufficient disk space for model weights (typically 100MB–1GB) and adequate RAM for inference.

### Can I use a custom OpenAI-compatible endpoint for embeddings?

Yes. Set `embedding.mode` to `openai` and specify your custom base URL in `openai.embedding_api_base`. This works with LocalAI, text-embeddings-inference, or any service implementing the OpenAI embeddings specification.

### What is the difference between Ollama and HuggingFace modes?

**HuggingFace** downloads transformer weights directly via the HuggingFace Hub and runs them in-process using PyTorch. **Ollama** sends HTTP requests to an external Ollama server process, which manages its own model cache and runtime environment independently of Private-GPT.

### Does Mistral AI require separate API credentials from OpenAI?

No. Mistral AI configuration reuses the `openai.api_key` and `openai.embedding_model` fields in [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml). However, you must set `embedding.mode: mistralai` to ensure the component instantiates `MistralAIEmbedding` rather than the standard OpenAI client.