How to Use the PrivateGPT Low-Level Embeddings API for Custom RAG Pipelines
PrivateGPT exposes an OpenAI-compatible POST /v1/embeddings endpoint backed by the EmbeddingComponent and EmbeddingsService classes, allowing you to generate dense vectors for any text and integrate them into bespoke Retrieval-Augmented Generation workflows.
The zylon-ai/private-gpt repository decouples its embeddings stack from the chat interface, providing a thin, low-level API that you can call from external services or internal Python modules. This architecture enables you to embed documents using any supported provider—HuggingFace, OpenAI, Ollama, Azure, Gemini, MistralAI, or SageMaker—and store the resulting vectors in your preferred vector database.
Architectural Overview of the Embeddings Stack
Configuration and Settings
All embedding-related configuration lives in the central Settings model defined in private_gpt/settings/settings.py. This includes the provider mode, model name, API keys, and dimensionality settings. The application inspects settings.embedding.mode at startup to determine which concrete implementation to instantiate.
The EmbeddingComponent Class
The EmbeddingComponent class in private_gpt/components/embedding/embedding_component.py acts as the core abstraction. During application initialization, the dependency injector creates a singleton instance of this component. It reads the embedding mode from settings and constructs a concrete BaseEmbedding implementation from the LlamaIndex library. The selected model is stored as self.embedding_model, exposing methods like get_text_embedding_batch() for vector generation.
EmbeddingsService and the HTTP Endpoint
The FastAPI route handler in private_gpt/server/embeddings/embeddings_router.py exposes the POST /v1/embeddings endpoint. This route injects the EmbeddingsService singleton, defined in private_gpt/server/embeddings/embeddings_service.py. The service's texts_embeddings() method receives a list of input strings, calls self.embedding_model.get_text_embedding_batch(texts), and wraps each resulting vector in a Pydantic Embedding model. The router then returns an OpenAI-compatible JSON payload containing object: "list", model: "private-gpt", and the data array.
Calling the Low-Level Embeddings API
HTTP Client Example (OpenAI-Compatible)
Because the endpoint follows the OpenAI specification, you can use standard HTTP clients or the official OpenAI SDK. Below is a raw Python example using requests:
import json
import requests
url = "http://localhost:8000/v1/embeddings"
payload = {
"model": "text-embedding-ada-002", # any model name supported by your Settings
"input": [
"Explain the difference between LLMs and embeddings.",
"What is a vector store?"
]
}
headers = {"Content-Type": "application/json"}
resp = requests.post(url, json=payload, headers=headers)
resp.raise_for_status()
embeddings = resp.json()["data"] # list of {"object":"embedding","embedding":[...],"index":...}
print(json.dumps(embeddings, indent=2))
Direct Python Component Usage
For in-process pipelines where you want to avoid HTTP overhead, instantiate the EmbeddingComponent directly:
from private_gpt.components.embedding.embedding_component import EmbeddingComponent
from private_gpt.settings.settings import Settings
# Load settings (e.g., from a YAML file)
settings = Settings.from_yaml("settings.yaml") # see private_gpt/settings/yaml.py
# Initialise the component (injector does this automatically in the app)
embed_comp = EmbeddingComponent(settings)
# Batch-embed a list of texts
texts = ["What is Retrieval‑Augmented Generation?", "How does Chroma store vectors?"]
vectors = embed_comp.embedding_model.get_text_embedding_batch(texts)
for i, vec in enumerate(vectors):
print(f"Doc {i} → dim {len(vec)}")
Building a Custom RAG Pipeline
The low-level embeddings API enables you to construct bespoke RAG workflows that are completely independent of the built-in chat UI. The following skeleton demonstrates how to combine the EmbeddingComponent with the vector store to ingest documents and perform similarity search:
from private_gpt.components.embedding.embedding_component import EmbeddingComponent
from private_gpt.components.vector_store.batched_chroma import BatchedChroma
from private_gpt.settings.settings import Settings
# 1️⃣ Load configuration
settings = Settings.from_yaml("settings.yaml")
# 2️⃣ Initialise components
embed = EmbeddingComponent(settings)
vector_store = BatchedChroma(settings) # uses settings.embedding.embed_dim etc.
# 3️⃣ Ingest documents
documents = ["Doc A text …", "Doc B text …", "Doc C text …"]
doc_vectors = embed.embedding_model.get_text_embedding_batch(documents)
# Store vectors + original texts
vector_store.upsert(
ids=[f"doc-{i}" for i in range(len(documents))],
embeddings=doc_vectors,
metadatas=[{"text": txt} for txt in documents],
)
# 4️⃣ Query
query = "How can I retrieve similar docs?"
query_vec = embed.embedding_model.get_text_embedding_batch([query])[0]
# 5️⃣ Retrieve top‑k neighbours
results = vector_store.query(
query_embeddings=[query_vec],
n_results=3,
include=["metadatas"]
)
print("Retrieved docs:")
for meta in results["metadatas"][0]:
print("-", meta["text"])
The same logic applies if you prefer to call the HTTP /v1/embeddings endpoint instead of the Python component—simply replace the embedding generation steps with a requests.post call to the endpoint and pass the returned vectors into vector_store.query.
Key Source Files Reference
Understanding the following files is essential for extending or debugging the embeddings stack:
private_gpt/settings/settings.py– Central Pydantic settings model; holds provider-specific fields such asopenai.embedding_model,ollama.embedding_model, and dimensionality configuration.private_gpt/components/embedding/embedding_component.py– Instantiates the concreteBaseEmbeddingimplementation based onsettings.embedding.modeand exposesself.embedding_model.private_gpt/server/embeddings/embeddings_service.py– Thin service layer that wrapsget_text_embedding_batch()results into PydanticEmbeddingmodels.private_gpt/server/embeddings/embeddings_router.py– FastAPI route handler exposingPOST /v1/embeddingswith OpenAI-compatible request/response schemas.private_gpt/components/vector_store/batched_chroma.py– Example vector-store implementation for ChromaDB that consumes embeddings generated by the API.private_gpt/components/vector_store/vector_store_component.py– Abstract interface for any vector database (Postgres, Chroma, etc.).private_gpt/server/ingest/ingest_service.pyandprivate_gpt/server/chat/chat_service.py– Demonstrate how higher-level services reuse the sameEmbeddingComponentfor document ingestion and retrieval.
Summary
- PrivateGPT exposes a thin, OpenAI-compatible
POST /v1/embeddingsendpoint that is decoupled from the chat interface, making it ideal for custom RAG pipelines. - The stack consists of three layers:
Settingsfor configuration,EmbeddingComponentfor model instantiation, andEmbeddingsServicefor request handling. - You can interact with the system via HTTP (compatible with the OpenAI SDK) or directly via Python by importing
EmbeddingComponentand callingget_text_embedding_batch(). - Generated vectors can be stored in any compatible vector database (Chroma, PostgreSQL, etc.) and retrieved using standard similarity search to feed context into LLM prompts.
Frequently Asked Questions
What is the difference between the embeddings API and the chat API in PrivateGPT?
The embeddings API (POST /v1/embeddings) is a stateless endpoint that converts text into dense vectors and returns them immediately. It does not perform retrieval or generation. The chat API (POST /v1/chat/completions) orchestrates the full RAG flow: it embeds the query, retrieves documents from the vector store, injects them into a prompt template, and streams a response from the LLM. Using the low-level embeddings API allows you to bypass the built-in chat logic and implement your own retrieval strategies.
Can I use external embedding providers like OpenAI or Ollama with this API?
Yes. The EmbeddingComponent in private_gpt/components/embedding/embedding_component.py supports multiple providers including HuggingFace, OpenAI, Ollama, Azure OpenAI, Gemini, MistralAI, and SageMaker. You configure the desired provider in your settings.yaml file under the embedding section (e.g., mode: openai or mode: ollama). The component automatically instantiates the correct BaseEmbedding implementation from LlamaIndex, and the /v1/embeddings endpoint will use that model for all requests.
How do I configure the embedding model dimensions for my vector store?
Embedding dimensions are determined by the specific model you select in settings.yaml (for example, text-embedding-ada-002 produces 1536 dimensions, while local HuggingFace models may produce 384 or 768). The vector store components, such as BatchedChroma in private_gpt/components/vector_store/batched_chroma.py, read settings.embedding.embed_dim to initialize their collections with the correct dimensionality. Ensure that the value in your settings file matches the output dimension of your chosen embedding model to avoid dimension mismatch errors during ingestion or query time.
Is the embeddings endpoint compatible with the OpenAI Python SDK?
Yes. The POST /v1/embeddings endpoint in private_gpt/server/embeddings/embeddings_router.py implements the OpenAI API specification, returning a JSON payload with object: "list", model: "private-gpt", and a data array containing embedding objects with object, embedding, and index fields. You can point the official openai Python library (or any other OpenAI-compatible client) at your PrivateGPT base URL (e.g., http://localhost:8000) and call client.embeddings.create() to generate vectors, making integration with existing toolchains seamless.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →