How to Use the PrivateGPT Low-Level Embeddings API for Custom RAG Pipelines

PrivateGPT exposes an OpenAI-compatible POST /v1/embeddings endpoint backed by the EmbeddingComponent and EmbeddingsService classes, allowing you to generate dense vectors for any text and integrate them into bespoke Retrieval-Augmented Generation workflows.

The zylon-ai/private-gpt repository decouples its embeddings stack from the chat interface, providing a thin, low-level API that you can call from external services or internal Python modules. This architecture enables you to embed documents using any supported provider—HuggingFace, OpenAI, Ollama, Azure, Gemini, MistralAI, or SageMaker—and store the resulting vectors in your preferred vector database.

Architectural Overview of the Embeddings Stack

Configuration and Settings

All embedding-related configuration lives in the central Settings model defined in private_gpt/settings/settings.py. This includes the provider mode, model name, API keys, and dimensionality settings. The application inspects settings.embedding.mode at startup to determine which concrete implementation to instantiate.

The EmbeddingComponent Class

The EmbeddingComponent class in private_gpt/components/embedding/embedding_component.py acts as the core abstraction. During application initialization, the dependency injector creates a singleton instance of this component. It reads the embedding mode from settings and constructs a concrete BaseEmbedding implementation from the LlamaIndex library. The selected model is stored as self.embedding_model, exposing methods like get_text_embedding_batch() for vector generation.

EmbeddingsService and the HTTP Endpoint

The FastAPI route handler in private_gpt/server/embeddings/embeddings_router.py exposes the POST /v1/embeddings endpoint. This route injects the EmbeddingsService singleton, defined in private_gpt/server/embeddings/embeddings_service.py. The service's texts_embeddings() method receives a list of input strings, calls self.embedding_model.get_text_embedding_batch(texts), and wraps each resulting vector in a Pydantic Embedding model. The router then returns an OpenAI-compatible JSON payload containing object: "list", model: "private-gpt", and the data array.

Calling the Low-Level Embeddings API

HTTP Client Example (OpenAI-Compatible)

Because the endpoint follows the OpenAI specification, you can use standard HTTP clients or the official OpenAI SDK. Below is a raw Python example using requests:

import json
import requests

url = "http://localhost:8000/v1/embeddings"
payload = {
    "model": "text-embedding-ada-002",  # any model name supported by your Settings

    "input": [
        "Explain the difference between LLMs and embeddings.",
        "What is a vector store?"
    ]
}
headers = {"Content-Type": "application/json"}

resp = requests.post(url, json=payload, headers=headers)
resp.raise_for_status()
embeddings = resp.json()["data"]  # list of {"object":"embedding","embedding":[...],"index":...}

print(json.dumps(embeddings, indent=2))

Direct Python Component Usage

For in-process pipelines where you want to avoid HTTP overhead, instantiate the EmbeddingComponent directly:

from private_gpt.components.embedding.embedding_component import EmbeddingComponent
from private_gpt.settings.settings import Settings

# Load settings (e.g., from a YAML file)

settings = Settings.from_yaml("settings.yaml")  # see private_gpt/settings/yaml.py

# Initialise the component (injector does this automatically in the app)

embed_comp = EmbeddingComponent(settings)

# Batch-embed a list of texts

texts = ["What is Retrieval‑Augmented Generation?", "How does Chroma store vectors?"]
vectors = embed_comp.embedding_model.get_text_embedding_batch(texts)

for i, vec in enumerate(vectors):
    print(f"Doc {i} → dim {len(vec)}")

Building a Custom RAG Pipeline

The low-level embeddings API enables you to construct bespoke RAG workflows that are completely independent of the built-in chat UI. The following skeleton demonstrates how to combine the EmbeddingComponent with the vector store to ingest documents and perform similarity search:

from private_gpt.components.embedding.embedding_component import EmbeddingComponent
from private_gpt.components.vector_store.batched_chroma import BatchedChroma
from private_gpt.settings.settings import Settings

# 1️⃣ Load configuration

settings = Settings.from_yaml("settings.yaml")

# 2️⃣ Initialise components

embed = EmbeddingComponent(settings)
vector_store = BatchedChroma(settings)  # uses settings.embedding.embed_dim etc.

# 3️⃣ Ingest documents

documents = ["Doc A text …", "Doc B text …", "Doc C text …"]
doc_vectors = embed.embedding_model.get_text_embedding_batch(documents)

# Store vectors + original texts

vector_store.upsert(
    ids=[f"doc-{i}" for i in range(len(documents))],
    embeddings=doc_vectors,
    metadatas=[{"text": txt} for txt in documents],
)

# 4️⃣ Query

query = "How can I retrieve similar docs?"
query_vec = embed.embedding_model.get_text_embedding_batch([query])[0]

# 5️⃣ Retrieve top‑k neighbours

results = vector_store.query(
    query_embeddings=[query_vec],
    n_results=3,
    include=["metadatas"]
)

print("Retrieved docs:")
for meta in results["metadatas"][0]:
    print("-", meta["text"])

The same logic applies if you prefer to call the HTTP /v1/embeddings endpoint instead of the Python component—simply replace the embedding generation steps with a requests.post call to the endpoint and pass the returned vectors into vector_store.query.

Key Source Files Reference

Understanding the following files is essential for extending or debugging the embeddings stack:

Summary

  • PrivateGPT exposes a thin, OpenAI-compatible POST /v1/embeddings endpoint that is decoupled from the chat interface, making it ideal for custom RAG pipelines.
  • The stack consists of three layers: Settings for configuration, EmbeddingComponent for model instantiation, and EmbeddingsService for request handling.
  • You can interact with the system via HTTP (compatible with the OpenAI SDK) or directly via Python by importing EmbeddingComponent and calling get_text_embedding_batch().
  • Generated vectors can be stored in any compatible vector database (Chroma, PostgreSQL, etc.) and retrieved using standard similarity search to feed context into LLM prompts.

Frequently Asked Questions

What is the difference between the embeddings API and the chat API in PrivateGPT?

The embeddings API (POST /v1/embeddings) is a stateless endpoint that converts text into dense vectors and returns them immediately. It does not perform retrieval or generation. The chat API (POST /v1/chat/completions) orchestrates the full RAG flow: it embeds the query, retrieves documents from the vector store, injects them into a prompt template, and streams a response from the LLM. Using the low-level embeddings API allows you to bypass the built-in chat logic and implement your own retrieval strategies.

Can I use external embedding providers like OpenAI or Ollama with this API?

Yes. The EmbeddingComponent in private_gpt/components/embedding/embedding_component.py supports multiple providers including HuggingFace, OpenAI, Ollama, Azure OpenAI, Gemini, MistralAI, and SageMaker. You configure the desired provider in your settings.yaml file under the embedding section (e.g., mode: openai or mode: ollama). The component automatically instantiates the correct BaseEmbedding implementation from LlamaIndex, and the /v1/embeddings endpoint will use that model for all requests.

How do I configure the embedding model dimensions for my vector store?

Embedding dimensions are determined by the specific model you select in settings.yaml (for example, text-embedding-ada-002 produces 1536 dimensions, while local HuggingFace models may produce 384 or 768). The vector store components, such as BatchedChroma in private_gpt/components/vector_store/batched_chroma.py, read settings.embedding.embed_dim to initialize their collections with the correct dimensionality. Ensure that the value in your settings file matches the output dimension of your chosen embedding model to avoid dimension mismatch errors during ingestion or query time.

Is the embeddings endpoint compatible with the OpenAI Python SDK?

Yes. The POST /v1/embeddings endpoint in private_gpt/server/embeddings/embeddings_router.py implements the OpenAI API specification, returning a JSON payload with object: "list", model: "private-gpt", and a data array containing embedding objects with object, embedding, and index fields. You can point the official openai Python library (or any other OpenAI-compatible client) at your PrivateGPT base URL (e.g., http://localhost:8000) and call client.embeddings.create() to generate vectors, making integration with existing toolchains seamless.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →