How to Implement Context Filtering in Private-GPT to Restrict RAG Responses to Specific Documents
Private-GPT enables document-scoped RAG responses through the ContextFilter mechanism, which limits vector-store retrieval to specific doc_id values passed via API or service-level calls.
The open-source zylon-ai/private-gpt repository provides a production-ready RAG pipeline that supports granular document access control through context filtering. By leveraging the ContextFilter Pydantic model and underlying metadata filtering capabilities, you can constrain retrieval-augmented generation to specific subsets of ingested documents without modifying the core embedding or LLM configuration.
Understanding the ContextFilter Architecture
Private-GPT implements context filtering through a coordinated three-layer architecture that translates high-level document IDs into vector-store specific metadata queries.
The flow originates in private_gpt/open_ai/extensions/context_filter.py, where the ContextFilter model defines a simple docs_ids list structure. This filter propagates through private_gpt/components/vector_store/vector_store_component.py, where the _doc_id_metadata_filter method converts the ID list into backend-specific metadata filters compatible with Qdrant, Chroma, or Postgres vector stores.
At the service layer, components such as private_gpt/server/chat/chat_service.py and private_gpt/server/recipes/summarize/summarize_service.py consume these filters to restrict retrieval engines and reference document sets, ensuring only authorized documents contribute to the generated response.
Step-by-Step Implementation Guide
Implementing context filtering requires collecting document IDs from the ingestion pipeline and passing them through the appropriate API or service interface.
Step 1: Identify Document IDs
When you ingest files via the ingestion endpoints defined in private_gpt/server/ingest/ingest_router.py (lines 51-57), the system returns unique doc_id values for each processed document. Retrieve these IDs through the list endpoint or capture them during the initial ingestion response.
Step 2: Construct the ContextFilter Object
Create a ContextFilter instance containing the target document IDs. The model accepts a list of UUID strings that identify specific documents in the vector store.
Step 3: Apply the Filter to RAG Endpoints
Pass the constructed filter to chat completion, completion, chunk retrieval, or summarization endpoints by including it in the request payload under the context_filter key.
Code Examples for Context Filtering
The following examples demonstrate context filtering across Python SDK usage, HTTP API calls, and internal service integration.
Python SDK Implementation
Instantiate the ContextFilter directly from the extensions module to restrict chat operations to specific documents:
from private_gpt.open_ai.extensions.context_filter import ContextFilter
# Target specific documents by their ingestion IDs
filter_obj = ContextFilter(docs_ids=[
"c202d5e6-7b69-4869-81cc-dd574ee8ee11",
"a7b3f2d9-3e1a-4d5c-9f8e-2b6a5e9c4d3f",
])
REST API Usage
Include the context_filter object in your JSON payload when calling the chat completions endpoint:
POST /v1/chat/completions HTTP/1.1
Content-Type: application/json
Authorization: Bearer <your-token>
{
"model": "gpt-4o-mini",
"messages": [
{"role": "user", "content": "Summarize the key points."}
],
"use_context": true,
"context_filter": {
"docs_ids": [
"c202d5e6-7b69-4869-81cc-dd574ee8ee11",
"a7b3f2d9-3e1a-4d5c-9f8e-2b6a5e9c4d3f"
]
}
}
Service-Level Integration
For custom service implementations, inject the ChatService or SummarizeService and pass the filter directly:
from private_gpt.server.chat.chat_service import ChatService
from private_gpt.open_ai.extensions.context_filter import ContextFilter
# Obtain service via dependency injection
chat_svc: ChatService = ...
messages = [...]
# Execute filtered chat
response = chat_svc.chat(
messages=messages,
use_context=True,
context_filter=ContextFilter(docs_ids=["doc-id-1", "doc-id-2"])
)
For summarization tasks specifically:
from private_gpt.server.recipes.summarize.summarize_service import SummarizeService
from private_gpt.open_ai.extensions.context_filter import ContextFilter
summarizer = SummarizeService(...)
summary = summarizer.summarize(
use_context=True,
context_filter=ContextFilter(docs_ids=["doc-id-3"])
)
How Context Filtering Works Under the Hood
The context filtering mechanism operates through two critical transformation stages that ensure efficient document scoping at the storage layer.
Vector Store Metadata Translation
In private_gpt/components/vector_store/vector_store_component.py, the _doc_id_metadata_filter method (lines 20-29) translates the ContextFilter.docs_ids list into MetadataFilter objects. These filters leverage the underlying vector store's native metadata querying capabilities—whether Qdrant's payload filters, Chroma's where clauses, or Postgres JSONB queries—to pre-filter chunks before similarity search execution.
Service Layer Enforcement
The ChatService class in private_gpt/server/chat/chat_service.py passes the context filter through its _chat_engine method to the VectorStoreComponent.get_retriever interface. Similarly, the SummarizeService in private_gpt/server/recipes/summarize/summarize_service.py utilizes _filter_ref_docs (lines 56-67) to apply document restrictions before generating summaries. This dual-layer approach ensures that both retrieval contexts and reference document sets respect the specified ID constraints.
Summary
- ContextFilter is defined in
private_gpt/open_ai/extensions/context_filter.pyas a Pydantic model containing adocs_idslist. - The filter converts to vector-store specific metadata filters via
_doc_id_metadata_filterin the vector store component. - Service implementations in
chat_service.pyandsummarize_service.pyenforce these constraints during RAG operations. - You can apply context filtering via REST API payloads, Python SDK calls, or direct service integration.
- This mechanism enables deterministic, document-scoped responses without modifying embedding models or re-indexing content.
Frequently Asked Questions
What format should document IDs use in the ContextFilter?
Document IDs must be valid UUID strings that match the doc_id values returned by the ingestion endpoints in private_gpt/server/ingest/ingest_router.py. The system expects a standard string representation of UUIDs within the docs_ids list.
Does context filtering work with all vector store backends?
Yes, the abstraction in vector_store_component.py handles backend-specific translations. Whether you configure Qdrant, Chroma, or Postgres as your vector store, the _doc_id_metadata_filter method generates the appropriate metadata query syntax for that backend.
Can I combine context filtering with other retrieval parameters?
Absolutely. The ContextFilter operates independently of other parameters like top_k or similarity thresholds. You can specify document IDs while adjusting retrieval counts and relevance cutoffs to fine-tune both scope and result quality.
Is there a performance impact when using context filtering?
Context filtering typically improves performance by reducing the search space. By pre-filtering the vector store metadata before similarity search, the system examines fewer chunks, resulting in faster retrieval times compared to scanning the entire document corpus.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →