# How to Use the High-Level RAG Chat API in PrivateGPT for Conversational Context

> Learn to use PrivateGPT's high-level RAG chat API with ChatService for conversational context from ingested documents. Explore chat() and stream_chat() with use_context=True.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: tutorial
- Published: 2026-03-06

---

**The high-level RAG chat API in PrivateGPT provides a `ChatService` class that abstracts the entire Retrieval-Augmented Generation pipeline, enabling conversational context from ingested documents through the `chat()` and `stream_chat()` methods when `use_context=True` is configured.**

PrivateGPT's `zylon-ai/private-gpt` repository ships with a production-ready high-level RAG chat API that eliminates the need to manually orchestrate vector stores, retrievers, and LLM prompts. This service seamlessly connects document ingestion pipelines with conversational interfaces, allowing developers to build context-aware chatbots with minimal boilerplate.

## Understanding the High-Level RAG Chat Architecture

The RAG chat functionality centers on the `ChatService` class located in [`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py). This service acts as a unified interface that coordinates multiple underlying components to provide contextual responses.

### Core Components and Data Flow

When you invoke the high-level RAG chat API with `use_context=True`, the service executes the following pipeline:

1. **Retrieval Configuration**: The service obtains a retriever from `VectorStoreComponent.get_retriever` and configures post-processors including similarity thresholds and optional reranking (source lines 15-31 of [`chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/chat_service.py)).

2. **Engine Instantiation**: It constructs a `ContextChatEngine` using the retriever and node post-processors (source lines 37-44 of [`chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/chat_service.py)).

3. **Context Injection**: The engine automatically queries the vector store, retrieves relevant chunks, and injects them into the LLM prompt alongside the conversation history.

4. **Response Generation**: The configured LLM (from `LLMComponent`) generates the final response, with support for both synchronous and streaming outputs.

### Configuration via RagSettings

All retrieval behavior is governed by the `RagSettings` model defined in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) (lines 99-108). Key parameters include:

- `similarity_top_k`: Number of chunks to retrieve
- `similarity_value`: Minimum similarity score threshold
- `rerank.enabled`: Whether to apply a reranking model to retrieved results

## Implementing Synchronous RAG Chat

The `chat()` method provides a blocking call that returns the complete response and source documents. This is ideal for batch processing or simple scripts.

```python
from private_gpt.server.chat.chat_service import ChatService
from private_gpt.server.chat.chat_service import Completion
from llama_index.core.schema import ChatMessage, MessageRole

# Resolve the singleton via the DI container (or instantiate directly in tests)

chat_service: ChatService = ChatService()   # DI provides the required components

# Build a conversation history – system prompt + user question

messages = [
    ChatMessage(
        role=MessageRole.SYSTEM,
        content="You are a helpful assistant that answers based on the provided documents."
    ),
    ChatMessage(role=MessageRole.USER, content="What does the privacy policy say about data retention?")
]

# Enable RAG context

result: Completion = chat_service.chat(messages, use_context=True)

print("Answer:", result.response)
print("\nSources:")
for src in result.sources or []:
    print(f"- {src.metadata['doc_id']} (score: {src.score:.2f})")

```

The call to `chat_service.chat(..., use_context=True)` triggers the `ContextChatEngine` path in [`chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/chat_service.py). The returned `Completion` object contains both the generated text and a list of source `Chunk` objects with metadata and relevance scores.

## Implementing Streaming RAG Chat

For interactive user interfaces, the `stream_chat()` method returns an async generator that yields tokens as they are generated, while still providing access to source documents upon completion.

```python
from private_gpt.server.chat.chat_service import ChatService
from llama_index.core.schema import ChatMessage, MessageRole

chat = ChatService()
messages = [
    ChatMessage(role=MessageRole.USER, content="Summarize the key points from the last meeting notes.")
]

# Stream the response

gen = chat.stream_chat(messages, use_context=True)

# The generator yields tokens; we also have source chunks at the end

for token in gen.response:
    print(token, end="", flush=True)

print("\n\nSources:")
for src in gen.sources or []:
    print(f"- {src.metadata['doc_id']}")

```

The `stream_chat` method builds the same retrieval engine but invokes `engine.stream_chat`, which returns a `TokenGen` generator. Source nodes are collected once the stream completes, as implemented in lines 75-84 of [`chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/chat_service.py).

## Filtering Context with ContextFilter

The `ContextFilter` class enables scoped retrieval, allowing you to restrict searches to specific documents. This is essential for multi-tenant applications or when users need answers from a subset of ingested files.

```python
from private_gpt.open_ai.extensions.context_filter import ContextFilter

filter = ContextFilter(docs_ids=["c202d5e6-7b69-4869-81cc-dd574ee8ee11"])

answer = chat.chat(messages, use_context=True, context_filter=filter)
print(answer.response)

```

The `ContextFilter` model is defined in [`private_gpt/open_ai/extensions/context_filter.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/open_ai/extensions/context_filter.py) (lines 4-8). When provided, the retriever only considers nodes whose `doc_id` matches the whitelist, effectively filtering the vector search results before they reach the LLM.

## Ingesting Documents for RAG

Before utilizing the high-level RAG chat API, documents must be ingested into the vector store. The ingestion pipeline is handled by components in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py).

```python
from private_gpt.components.ingest.ingest_helper import IngestionHelper
from private_gpt.components.ingest.ingest_component import get_ingestion_component
from private_gpt.server.utils.storage import get_storage_context
from private_gpt.settings import settings
from pathlib import Path

# Build the storage context (vector store + docstore)

storage_context = get_storage_context()

# Obtain the component matching the embedding configuration

ingest = get_ingestion_component(
    storage_context=storage_context,
    embed_model=settings().embedding,
    transformations=IngestionHelper.default_transformations(),
    settings=settings(),
)

# Ingest every PDF in ./data/

folder = Path("data")
for file_path in folder.rglob("*.pdf"):
    ingest.ingest(file_path.name, file_path)

```

The `get_ingestion_component` factory (lines 90-115 of [`ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/ingest_component.py)) selects the appropriate implementation—`SimpleIngestComponent`, `BatchIngestComponent`, or others—based on `settings.embedding.ingest_mode`. Documents are parsed, chunked, embedded, and stored in the vector database, making them available for the `ChatService` retrieval pipeline.

## Summary

- **The high-level RAG chat API** centers on `ChatService` in [`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py), which orchestrates retrieval and generation.
- **Enable RAG context** by setting `use_context=True` in `chat()` or `stream_chat()` calls; this instantiates a `ContextChatEngine` that queries the vector store.
- **Configure retrieval behavior** via `RagSettings` in [`settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/settings.py), controlling `similarity_top_k`, similarity thresholds, and reranking options.
- **Scope searches** using `ContextFilter` to restrict retrieval to specific document IDs, supporting multi-tenant or filtered search scenarios.
- **Stream responses** using `stream_chat()` for real-time token generation while still receiving source attribution upon completion.

## Frequently Asked Questions

### How do I enable RAG context in PrivateGPT chat requests?

Set the `use_context` parameter to `True` when calling `ChatService.chat()` or `ChatService.stream_chat()`. This triggers the service to instantiate a `ContextChatEngine` that retrieves relevant document chunks from the vector store and injects them into the LLM prompt alongside your conversation history.

### What is the difference between chat() and stream_chat() in the high-level API?

The `chat()` method returns a `Completion` object synchronously containing the full response text and source chunks, making it ideal for batch processing. The `stream_chat()` method returns a `CompletionGen` generator that yields tokens in real-time as the LLM generates them, which is better suited for interactive user interfaces, while still providing source attribution after the stream completes.

### How can I restrict the chat to only search specific ingested documents?

Pass a `ContextFilter` object to the `context_filter` parameter of the chat methods. Create the filter with a list of specific document IDs: `ContextFilter(docs_ids=["your-doc-id"])`. This restricts the vector store retriever to only consider chunks from the specified documents, enabling multi-tenant scenarios or focused searches within a document subset.

### Where is the RAG retrieval behavior configured in PrivateGPT?

Retrieval parameters are controlled by the `RagSettings` model in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) (lines 99-108). Key settings include `similarity_top_k` (number of chunks to retrieve), `similarity_value` (minimum similarity threshold), and `rerank.enabled` (whether to apply a reranking model to improve result quality). These settings automatically configure the `ContextChatEngine` when `use_context=True` is invoked.