Configuring Embedding Models in Private-GPT: HuggingFace, OpenAI, Ollama, and Mistral Options

Private-GPT uses a single embedding.mode setting to switch between embedding backends including HuggingFace, OpenAI, Ollama, and Mistral, with each provider configured through dedicated YAML blocks in settings.yaml.

Private-GPT is designed with a flexible embedding architecture that decouples the retrieval logic from the underlying vectorization service. By modifying the embedding model configuration in the central settings file, you can swap between local transformer models, commercial APIs, and self-hosted solutions without changing application code.

Understanding Embedding Configuration Architecture

The configuration system centers on the EmbeddingSettings model defined in private_gpt/settings/settings.py (lines 998–1008). This Pydantic model validates the embedding.mode field against a strict enum of supported backends and maps each mode to its required credential fields.

When the application initializes, the EmbeddingComponent factory class located in private_gpt/components/embedding/embedding_component.py (lines 18–67) reads this setting and instantiates the appropriate LlamaIndex embedding class using a Python match statement.

Available Embedding Modes

Private-GPT supports eight distinct embedding providers via the mode field. The primary options requested—HuggingFace, OpenAI, Ollama, and Mistral—are configured as follows:

  • huggingface – Local inference using Hugging Face transformer models
  • openai – OpenAI API or compatible self-hosted endpoints
  • ollama – Local Ollama server for embedding generation
  • mistralai – Mistral AI's embedding endpoint (OpenAI-compatible)
  • azopenai – Azure OpenAI Service (enterprise deployments)
  • sagemaker – Amazon SageMaker hosted endpoints
  • gemini – Google Gemini API
  • mock – Fixed-dimension random vectors for testing

Provider-Specific Configuration Options

Each embedding backend requires distinct authentication and model parameters defined under separate YAML keys.

HuggingFace (Local Models)

For local embedding execution, set embedding.mode to huggingface and specify the model identifier. The system downloads weights to models_cache_path and runs inference locally.

embedding:
  mode: huggingface

huggingface:
  embedding_hf_model_name: sentence-transformers/all-MiniLM-L6-v2
  trust_remote_code: false
  access_token: ${HUGGINGFACE_TOKEN}  # Optional, for gated models

Key parameters include embedding_hf_model_name (the HuggingFace Hub model ID) and trust_remote_code (boolean for custom model architectures).

OpenAI (Cloud API)

To use OpenAI embeddings, configure the API credentials and model name. The component constructs requests to the /embeddings endpoint.

embedding:
  mode: openai

openai:
  api_key: ${OPENAI_API_KEY}
  embedding_model: text-embedding-3-small
  embedding_api_base: https://api.openai.com/v1  # Optional override

The embedding_api_base field supports proxy servers or OpenAI-compatible alternatives like LocalAI.

Ollama (Self-Hosted)

The Ollama integration connects to a locally running Ollama instance (default http://localhost:11434). Enable autopull_models to automatically fetch the specified model tag if not present.

embedding:
  mode: ollama

ollama:
  embedding_api_base: http://localhost:11434
  embedding_model: nomic-embed-text
  autopull_models: true

This mode utilizes the OllamaEmbedding class from LlamaIndex and validates server reachability before ingestion begins.

Mistral AI (OpenAI-Compatible)

Mistral AI leverages the same underlying client as OpenAI but targets Mistral's endpoint. Configure it using the standard OpenAI credential fields while setting mode: mistralai.

embedding:
  mode: mistralai

openai:
  api_key: ${MISTRAL_API_KEY}
  embedding_model: mistral-embed
  embedding_api_base: https://api.mistral.ai/v1

As implemented in embedding_component.py, this mode instantiates MistralAIEmbedding but pulls configuration from the shared openai settings block.

Runtime Component Instantiation

The selection logic in private_gpt/components/embedding/embedding_component.py uses structural pattern matching to bind configuration to implementation:


# Simplified excerpt from embedding_component.py

match settings.embedding.mode:
    case "huggingface":
        from llama_index.embeddings.huggingface import HuggingFaceEmbedding
        embedding_model = HuggingFaceEmbedding(
            model_name=settings.huggingface.embedding_hf_model_name
        )
    case "openai":
        from llama_index.embeddings.openai import OpenAIEmbedding
        embedding_model = OpenAIEmbedding(
            api_key=settings.openai.api_key,
            model=settings.openai.embedding_model
        )
    case "ollama":
        from llama_index.embeddings.ollama import OllamaEmbedding
        embedding_model = OllamaEmbedding(
            base_url=settings.ollama.embedding_api_base,
            model_name=settings.ollama.embedding_model
        )
    case "mistralai":
        from llama_index.embeddings.mistralai import MistralAIEmbedding
        embedding_model = MistralAIEmbedding(
            api_key=settings.openai.api_key,
            model=settings.openai.embedding_model
        )

This factory approach ensures that changing the embedding provider requires only a configuration edit, with no code modifications.

Practical Implementation Example

Load settings and generate embeddings programmatically using the component:

from private_gpt.components.embedding.embedding_component import EmbeddingComponent
from private_gpt.settings.settings_loader import SettingsLoader

# Load configuration from private_gpt/settings.yaml

settings = SettingsLoader().load()

# Initialize the embedding component

component = EmbeddingComponent(settings)

# Generate embeddings

texts = ["Private-GPT supports multiple embedding providers"]
vectors = component.embedding_model.get_text_embedding_batch(texts)

print(f"Generated {len(vectors)} embeddings of dimension {len(vectors[0])}")

Summary

  • Centralized Configuration: The embedding.mode field in settings.yaml controls the backend selection.
  • File Locations: Settings are defined in private_gpt/settings/settings.py (lines 998–1008) and instantiated in private_gpt/components/embedding/embedding_component.py (lines 18–67).
  • Local Options: HuggingFace and Ollama provide fully local embedding without external API dependencies.
  • Cloud Options: OpenAI, Mistral, Azure OpenAI, and Gemini require API keys and support high-throughput managed inference.
  • Testing: The mock mode generates deterministic random vectors for CI/CD pipelines.

Frequently Asked Questions

How do I switch from OpenAI to a local HuggingFace model?

Change embedding.mode from openai to huggingface in settings.yaml, then populate the huggingface block with your chosen embedding_hf_model_name. Ensure sufficient disk space for model weights (typically 100MB–1GB) and adequate RAM for inference.

Can I use a custom OpenAI-compatible endpoint for embeddings?

Yes. Set embedding.mode to openai and specify your custom base URL in openai.embedding_api_base. This works with LocalAI, text-embeddings-inference, or any service implementing the OpenAI embeddings specification.

What is the difference between Ollama and HuggingFace modes?

HuggingFace downloads transformer weights directly via the HuggingFace Hub and runs them in-process using PyTorch. Ollama sends HTTP requests to an external Ollama server process, which manages its own model cache and runtime environment independently of Private-GPT.

Does Mistral AI require separate API credentials from OpenAI?

No. Mistral AI configuration reuses the openai.api_key and openai.embedding_model fields in settings.yaml. However, you must set embedding.mode: mistralai to ensure the component instantiates MistralAIEmbedding rather than the standard OpenAI client.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →