# How to Use the PrivateGPT Low-Level Embeddings API for Custom RAG Pipelines

> Learn to use PrivateGPT's low-level embeddings API to create custom RAG pipelines. Integrate dense vectors into your bespoke generation workflows via the OpenAI-compatible endpoint.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: how-to-guide
- Published: 2026-03-06

---

**PrivateGPT exposes an OpenAI-compatible `POST /v1/embeddings` endpoint backed by the `EmbeddingComponent` and `EmbeddingsService` classes, allowing you to generate dense vectors for any text and integrate them into bespoke Retrieval-Augmented Generation workflows.**

The `zylon-ai/private-gpt` repository decouples its embeddings stack from the chat interface, providing a thin, low-level API that you can call from external services or internal Python modules. This architecture enables you to embed documents using any supported provider—HuggingFace, OpenAI, Ollama, Azure, Gemini, MistralAI, or SageMaker—and store the resulting vectors in your preferred vector database.

## Architectural Overview of the Embeddings Stack

### Configuration and Settings

All embedding-related configuration lives in the central `Settings` model defined in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py). This includes the provider mode, model name, API keys, and dimensionality settings. The application inspects `settings.embedding.mode` at startup to determine which concrete implementation to instantiate.

### The EmbeddingComponent Class

The `EmbeddingComponent` class in [`private_gpt/components/embedding/embedding_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/embedding/embedding_component.py) acts as the core abstraction. During application initialization, the dependency injector creates a singleton instance of this component. It reads the embedding mode from settings and constructs a concrete `BaseEmbedding` implementation from the LlamaIndex library. The selected model is stored as `self.embedding_model`, exposing methods like `get_text_embedding_batch()` for vector generation.

### EmbeddingsService and the HTTP Endpoint

The FastAPI route handler in [`private_gpt/server/embeddings/embeddings_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/embeddings/embeddings_router.py) exposes the `POST /v1/embeddings` endpoint. This route injects the `EmbeddingsService` singleton, defined in [`private_gpt/server/embeddings/embeddings_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/embeddings/embeddings_service.py). The service's `texts_embeddings()` method receives a list of input strings, calls `self.embedding_model.get_text_embedding_batch(texts)`, and wraps each resulting vector in a Pydantic `Embedding` model. The router then returns an OpenAI-compatible JSON payload containing `object: "list"`, `model: "private-gpt"`, and the `data` array.

## Calling the Low-Level Embeddings API

### HTTP Client Example (OpenAI-Compatible)

Because the endpoint follows the OpenAI specification, you can use standard HTTP clients or the official OpenAI SDK. Below is a raw Python example using `requests`:

```python
import json
import requests

url = "http://localhost:8000/v1/embeddings"
payload = {
    "model": "text-embedding-ada-002",  # any model name supported by your Settings

    "input": [
        "Explain the difference between LLMs and embeddings.",
        "What is a vector store?"
    ]
}
headers = {"Content-Type": "application/json"}

resp = requests.post(url, json=payload, headers=headers)
resp.raise_for_status()
embeddings = resp.json()["data"]  # list of {"object":"embedding","embedding":[...],"index":...}

print(json.dumps(embeddings, indent=2))

```

### Direct Python Component Usage

For in-process pipelines where you want to avoid HTTP overhead, instantiate the `EmbeddingComponent` directly:

```python
from private_gpt.components.embedding.embedding_component import EmbeddingComponent
from private_gpt.settings.settings import Settings

# Load settings (e.g., from a YAML file)

settings = Settings.from_yaml("settings.yaml")  # see private_gpt/settings/yaml.py

# Initialise the component (injector does this automatically in the app)

embed_comp = EmbeddingComponent(settings)

# Batch-embed a list of texts

texts = ["What is Retrieval‑Augmented Generation?", "How does Chroma store vectors?"]
vectors = embed_comp.embedding_model.get_text_embedding_batch(texts)

for i, vec in enumerate(vectors):
    print(f"Doc {i} → dim {len(vec)}")

```

## Building a Custom RAG Pipeline

The low-level embeddings API enables you to construct bespoke RAG workflows that are completely independent of the built-in chat UI. The following skeleton demonstrates how to combine the `EmbeddingComponent` with the vector store to ingest documents and perform similarity search:

```python
from private_gpt.components.embedding.embedding_component import EmbeddingComponent
from private_gpt.components.vector_store.batched_chroma import BatchedChroma
from private_gpt.settings.settings import Settings

# 1️⃣ Load configuration

settings = Settings.from_yaml("settings.yaml")

# 2️⃣ Initialise components

embed = EmbeddingComponent(settings)
vector_store = BatchedChroma(settings)  # uses settings.embedding.embed_dim etc.

# 3️⃣ Ingest documents

documents = ["Doc A text …", "Doc B text …", "Doc C text …"]
doc_vectors = embed.embedding_model.get_text_embedding_batch(documents)

# Store vectors + original texts

vector_store.upsert(
    ids=[f"doc-{i}" for i in range(len(documents))],
    embeddings=doc_vectors,
    metadatas=[{"text": txt} for txt in documents],
)

# 4️⃣ Query

query = "How can I retrieve similar docs?"
query_vec = embed.embedding_model.get_text_embedding_batch([query])[0]

# 5️⃣ Retrieve top‑k neighbours

results = vector_store.query(
    query_embeddings=[query_vec],
    n_results=3,
    include=["metadatas"]
)

print("Retrieved docs:")
for meta in results["metadatas"][0]:
    print("-", meta["text"])

```

The same logic applies if you prefer to call the HTTP `/v1/embeddings` endpoint instead of the Python component—simply replace the embedding generation steps with a `requests.post` call to the endpoint and pass the returned vectors into `vector_store.query`.

## Key Source Files Reference

Understanding the following files is essential for extending or debugging the embeddings stack:

- **[`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py)** – Central Pydantic settings model; holds provider-specific fields such as `openai.embedding_model`, `ollama.embedding_model`, and dimensionality configuration.
- **[`private_gpt/components/embedding/embedding_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/embedding/embedding_component.py)** – Instantiates the concrete `BaseEmbedding` implementation based on `settings.embedding.mode` and exposes `self.embedding_model`.
- **[`private_gpt/server/embeddings/embeddings_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/embeddings/embeddings_service.py)** – Thin service layer that wraps `get_text_embedding_batch()` results into Pydantic `Embedding` models.
- **[`private_gpt/server/embeddings/embeddings_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/embeddings/embeddings_router.py)** – FastAPI route handler exposing `POST /v1/embeddings` with OpenAI-compatible request/response schemas.
- **[`private_gpt/components/vector_store/batched_chroma.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/vector_store/batched_chroma.py)** – Example vector-store implementation for ChromaDB that consumes embeddings generated by the API.
- **[`private_gpt/components/vector_store/vector_store_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/vector_store/vector_store_component.py)** – Abstract interface for any vector database (Postgres, Chroma, etc.).
- **[`private_gpt/server/ingest/ingest_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/ingest/ingest_service.py)** and **[`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py)** – Demonstrate how higher-level services reuse the same `EmbeddingComponent` for document ingestion and retrieval.

## Summary

- **PrivateGPT** exposes a thin, OpenAI-compatible `POST /v1/embeddings` endpoint that is decoupled from the chat interface, making it ideal for custom RAG pipelines.
- The stack consists of three layers: `Settings` for configuration, `EmbeddingComponent` for model instantiation, and `EmbeddingsService` for request handling.
- You can interact with the system via HTTP (compatible with the OpenAI SDK) or directly via Python by importing `EmbeddingComponent` and calling `get_text_embedding_batch()`.
- Generated vectors can be stored in any compatible vector database (Chroma, PostgreSQL, etc.) and retrieved using standard similarity search to feed context into LLM prompts.

## Frequently Asked Questions

### What is the difference between the embeddings API and the chat API in PrivateGPT?

The **embeddings API** (`POST /v1/embeddings`) is a stateless endpoint that converts text into dense vectors and returns them immediately. It does not perform retrieval or generation. The **chat API** (`POST /v1/chat/completions`) orchestrates the full RAG flow: it embeds the query, retrieves documents from the vector store, injects them into a prompt template, and streams a response from the LLM. Using the low-level embeddings API allows you to bypass the built-in chat logic and implement your own retrieval strategies.

### Can I use external embedding providers like OpenAI or Ollama with this API?

Yes. The `EmbeddingComponent` in [`private_gpt/components/embedding/embedding_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/embedding/embedding_component.py) supports multiple providers including HuggingFace, OpenAI, Ollama, Azure OpenAI, Gemini, MistralAI, and SageMaker. You configure the desired provider in your [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) file under the `embedding` section (e.g., `mode: openai` or `mode: ollama`). The component automatically instantiates the correct `BaseEmbedding` implementation from LlamaIndex, and the `/v1/embeddings` endpoint will use that model for all requests.

### How do I configure the embedding model dimensions for my vector store?

Embedding dimensions are determined by the specific model you select in [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) (for example, `text-embedding-ada-002` produces 1536 dimensions, while local HuggingFace models may produce 384 or 768). The vector store components, such as `BatchedChroma` in [`private_gpt/components/vector_store/batched_chroma.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/vector_store/batched_chroma.py), read `settings.embedding.embed_dim` to initialize their collections with the correct dimensionality. Ensure that the value in your settings file matches the output dimension of your chosen embedding model to avoid dimension mismatch errors during ingestion or query time.

### Is the embeddings endpoint compatible with the OpenAI Python SDK?

Yes. The `POST /v1/embeddings` endpoint in [`private_gpt/server/embeddings/embeddings_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/embeddings/embeddings_router.py) implements the OpenAI API specification, returning a JSON payload with `object: "list"`, `model: "private-gpt"`, and a `data` array containing embedding objects with `object`, `embedding`, and `index` fields. You can point the official `openai` Python library (or any other OpenAI-compatible client) at your PrivateGPT base URL (e.g., `http://localhost:8000`) and call `client.embeddings.create()` to generate vectors, making integration with existing toolchains seamless.