How to Use the High-Level RAG Chat API in PrivateGPT for Conversational Context
The high-level RAG chat API in PrivateGPT provides a ChatService class that abstracts the entire Retrieval-Augmented Generation pipeline, enabling conversational context from ingested documents through the chat() and stream_chat() methods when use_context=True is configured.
PrivateGPT's zylon-ai/private-gpt repository ships with a production-ready high-level RAG chat API that eliminates the need to manually orchestrate vector stores, retrievers, and LLM prompts. This service seamlessly connects document ingestion pipelines with conversational interfaces, allowing developers to build context-aware chatbots with minimal boilerplate.
Understanding the High-Level RAG Chat Architecture
The RAG chat functionality centers on the ChatService class located in private_gpt/server/chat/chat_service.py. This service acts as a unified interface that coordinates multiple underlying components to provide contextual responses.
Core Components and Data Flow
When you invoke the high-level RAG chat API with use_context=True, the service executes the following pipeline:
-
Retrieval Configuration: The service obtains a retriever from
VectorStoreComponent.get_retrieverand configures post-processors including similarity thresholds and optional reranking (source lines 15-31 ofchat_service.py). -
Engine Instantiation: It constructs a
ContextChatEngineusing the retriever and node post-processors (source lines 37-44 ofchat_service.py). -
Context Injection: The engine automatically queries the vector store, retrieves relevant chunks, and injects them into the LLM prompt alongside the conversation history.
-
Response Generation: The configured LLM (from
LLMComponent) generates the final response, with support for both synchronous and streaming outputs.
Configuration via RagSettings
All retrieval behavior is governed by the RagSettings model defined in private_gpt/settings/settings.py (lines 99-108). Key parameters include:
similarity_top_k: Number of chunks to retrievesimilarity_value: Minimum similarity score thresholdrerank.enabled: Whether to apply a reranking model to retrieved results
Implementing Synchronous RAG Chat
The chat() method provides a blocking call that returns the complete response and source documents. This is ideal for batch processing or simple scripts.
from private_gpt.server.chat.chat_service import ChatService
from private_gpt.server.chat.chat_service import Completion
from llama_index.core.schema import ChatMessage, MessageRole
# Resolve the singleton via the DI container (or instantiate directly in tests)
chat_service: ChatService = ChatService() # DI provides the required components
# Build a conversation history – system prompt + user question
messages = [
ChatMessage(
role=MessageRole.SYSTEM,
content="You are a helpful assistant that answers based on the provided documents."
),
ChatMessage(role=MessageRole.USER, content="What does the privacy policy say about data retention?")
]
# Enable RAG context
result: Completion = chat_service.chat(messages, use_context=True)
print("Answer:", result.response)
print("\nSources:")
for src in result.sources or []:
print(f"- {src.metadata['doc_id']} (score: {src.score:.2f})")
The call to chat_service.chat(..., use_context=True) triggers the ContextChatEngine path in chat_service.py. The returned Completion object contains both the generated text and a list of source Chunk objects with metadata and relevance scores.
Implementing Streaming RAG Chat
For interactive user interfaces, the stream_chat() method returns an async generator that yields tokens as they are generated, while still providing access to source documents upon completion.
from private_gpt.server.chat.chat_service import ChatService
from llama_index.core.schema import ChatMessage, MessageRole
chat = ChatService()
messages = [
ChatMessage(role=MessageRole.USER, content="Summarize the key points from the last meeting notes.")
]
# Stream the response
gen = chat.stream_chat(messages, use_context=True)
# The generator yields tokens; we also have source chunks at the end
for token in gen.response:
print(token, end="", flush=True)
print("\n\nSources:")
for src in gen.sources or []:
print(f"- {src.metadata['doc_id']}")
The stream_chat method builds the same retrieval engine but invokes engine.stream_chat, which returns a TokenGen generator. Source nodes are collected once the stream completes, as implemented in lines 75-84 of chat_service.py.
Filtering Context with ContextFilter
The ContextFilter class enables scoped retrieval, allowing you to restrict searches to specific documents. This is essential for multi-tenant applications or when users need answers from a subset of ingested files.
from private_gpt.open_ai.extensions.context_filter import ContextFilter
filter = ContextFilter(docs_ids=["c202d5e6-7b69-4869-81cc-dd574ee8ee11"])
answer = chat.chat(messages, use_context=True, context_filter=filter)
print(answer.response)
The ContextFilter model is defined in private_gpt/open_ai/extensions/context_filter.py (lines 4-8). When provided, the retriever only considers nodes whose doc_id matches the whitelist, effectively filtering the vector search results before they reach the LLM.
Ingesting Documents for RAG
Before utilizing the high-level RAG chat API, documents must be ingested into the vector store. The ingestion pipeline is handled by components in private_gpt/components/ingest/ingest_component.py.
from private_gpt.components.ingest.ingest_helper import IngestionHelper
from private_gpt.components.ingest.ingest_component import get_ingestion_component
from private_gpt.server.utils.storage import get_storage_context
from private_gpt.settings import settings
from pathlib import Path
# Build the storage context (vector store + docstore)
storage_context = get_storage_context()
# Obtain the component matching the embedding configuration
ingest = get_ingestion_component(
storage_context=storage_context,
embed_model=settings().embedding,
transformations=IngestionHelper.default_transformations(),
settings=settings(),
)
# Ingest every PDF in ./data/
folder = Path("data")
for file_path in folder.rglob("*.pdf"):
ingest.ingest(file_path.name, file_path)
The get_ingestion_component factory (lines 90-115 of ingest_component.py) selects the appropriate implementation—SimpleIngestComponent, BatchIngestComponent, or others—based on settings.embedding.ingest_mode. Documents are parsed, chunked, embedded, and stored in the vector database, making them available for the ChatService retrieval pipeline.
Summary
- The high-level RAG chat API centers on
ChatServiceinprivate_gpt/server/chat/chat_service.py, which orchestrates retrieval and generation. - Enable RAG context by setting
use_context=Trueinchat()orstream_chat()calls; this instantiates aContextChatEnginethat queries the vector store. - Configure retrieval behavior via
RagSettingsinsettings.py, controllingsimilarity_top_k, similarity thresholds, and reranking options. - Scope searches using
ContextFilterto restrict retrieval to specific document IDs, supporting multi-tenant or filtered search scenarios. - Stream responses using
stream_chat()for real-time token generation while still receiving source attribution upon completion.
Frequently Asked Questions
How do I enable RAG context in PrivateGPT chat requests?
Set the use_context parameter to True when calling ChatService.chat() or ChatService.stream_chat(). This triggers the service to instantiate a ContextChatEngine that retrieves relevant document chunks from the vector store and injects them into the LLM prompt alongside your conversation history.
What is the difference between chat() and stream_chat() in the high-level API?
The chat() method returns a Completion object synchronously containing the full response text and source chunks, making it ideal for batch processing. The stream_chat() method returns a CompletionGen generator that yields tokens in real-time as the LLM generates them, which is better suited for interactive user interfaces, while still providing source attribution after the stream completes.
How can I restrict the chat to only search specific ingested documents?
Pass a ContextFilter object to the context_filter parameter of the chat methods. Create the filter with a list of specific document IDs: ContextFilter(docs_ids=["your-doc-id"]). This restricts the vector store retriever to only consider chunks from the specified documents, enabling multi-tenant scenarios or focused searches within a document subset.
Where is the RAG retrieval behavior configured in PrivateGPT?
Retrieval parameters are controlled by the RagSettings model in private_gpt/settings/settings.py (lines 99-108). Key settings include similarity_top_k (number of chunks to retrieve), similarity_value (minimum similarity threshold), and rerank.enabled (whether to apply a reranking model to improve result quality). These settings automatically configure the ContextChatEngine when use_context=True is invoked.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →