How to Use the Low-Level Chunks API for Custom Retrieval Logic in RAG
The low-level Chunks API in PrivateGPT exposes the /v1/chunks endpoint to query indexed document fragments directly, returning similarity-scored chunks with optional surrounding context for building custom RAG pipelines without LLM inference overhead.
PrivateGPT stores every ingested document as a set of chunks (text fragments) indexed in a vector store. The low-level Chunks API lets you query that index directly, bypassing the LLM and returning only the most relevant fragments together with their surrounding context, making it ideal for custom retrieval-augmented generation workflows.
Architecture of the Low-Level Chunks API
The Chunks API follows a clean layered architecture that separates HTTP handling from vector-store operations.
HTTP Router and Request Schema
The entry point is defined in private_gpt/server/chunks/chunks_router.py, which registers the /v1/chunks POST endpoint. The request body uses the ChunksBody model containing:
text: The query string for similarity searchcontext_filter: An optionalContextFilterto limit searches to specific document IDslimit: Number of top results to returnprev_next_chunks: Number of surrounding fragments to fetch for context enrichment
The router resolves ChunksService from the dependency injection container and delegates to retrieve_relevant.
Service Layer and Vector Store Integration
In private_gpt/server/chunks/chunks_service.py, the ChunksService.retrieve_relevant method orchestrates the retrieval:
- Builds a
VectorStoreIndexbacked by the configured vector store (VectorStoreComponent) - Creates a retriever via
VectorStoreComponent.get_retriever, applying the optionalContextFilteras a document-ID filter - Executes a similarity-top-K search using
vector_index_retriever.retrieve(text)
The raw NodeWithScore objects from the vector store are sorted by similarity score and converted into public-facing Chunk models.
Chunk Conversion and Response Models
Each retrieved node is transformed into a Chunk object via Chunk.from_node (defined in chunks_service.py), which includes:
object: Always"context.chunk"score: The similarity score from the vector searchdocument: AnIngestedDocinstance (fromprivate_gpt/server/ingest/model.py) containingdoc_idand metadatatext: The fragment contentprevious_texts/next_texts: Optional sibling fragments obtained by traversing the node graph using_get_sibling_nodes_text
The router serializes a list of Chunk objects into ChunksResponse. Because no LLM call is involved, the endpoint delivers low-latency, cost-efficient retrieval ideal for custom preprocessing logic.
Querying the Chunks API
HTTP Request with curl
Send a POST request to retrieve the top 5 chunks related to sales figures, including one surrounding chunk on each side:
curl -X POST https://your-private-gpt-instance/v1/chunks \
-H "Authorization: Bearer <your-token>" \
-H "Content-Type: application/json" \
-d '{
"text": "Q3 2023 sales figures",
"limit": 5,
"prev_next_chunks": 1,
"context_filter": { "docs_ids": ["c202d5e6-7b69-4869-81cc-dd574ee8ee11"] }
}'
The response contains a data array of Chunk items, each with matched text, score, source document metadata, and surrounding context.
Python Client Implementation
For programmatic access, use the requests library to parse and filter results:
import requests
API_URL = "https://your-private-gpt-instance/v1/chunks"
TOKEN = "YOUR_API_KEY"
payload = {
"text": "How did our marketing spend affect sales?",
"limit": 3,
"prev_next_chunks": 2, # Two chunks before & after each hit
"context_filter": {"docs_ids": ["c202d5e6-7b69-4869-81cc-dd574ee8ee11"]},
}
headers = {
"Authorization": f"Bearer {TOKEN}",
"Content-Type": "application/json",
}
resp = requests.post(API_URL, json=payload, headers=headers)
chunks = resp.json()["data"]
for ch in chunks:
print(f"Score: {ch['score']:.4f}")
print(f"Doc ID: {ch['document']['doc_id']}")
print("Match:", ch["text"])
if ch["previous_texts"]:
print("Prev:", ch["previous_texts"])
if ch["next_texts"]:
print("Next:", ch["next_texts"])
print("-" * 40)
Direct Service Injection
When extending PrivateGPT with custom FastAPI routes or background workers, inject ChunksService directly to bypass HTTP overhead:
from private_gpt.server.chunks.chunks_service import ChunksService
def custom_retrieval(service: ChunksService, query: str):
# Retrieve top 10 chunks from all documents, with 2 surrounding pieces
chunks = service.retrieve_relevant(
text=query,
context_filter=None, # Search all documents
limit=10,
prev_next_chunks=2,
)
return chunks
Because ChunksService is a singleton managed by the DI container, you can reuse it across requests without recreating the vector index.
Extending Retrieval Logic
Filtering by Document ID
Restrict retrieval to a subset of documents using the ContextFilter (defined in private_gpt/open_ai/extensions/context_filter.py). Populate context_filter.docs_ids in the request body to limit the vector search scope to specific ingested documents, reducing noise and improving relevance for domain-specific queries.
Including Surrounding Context
Set prev_next_chunks to a positive integer to retrieve sibling fragments. The service calls _get_sibling_nodes_text to traverse the node graph and fetch preceding and following chunks, providing richer context for the LLM without requiring multiple API calls.
Custom Metadata Filters
To filter by metadata fields (e.g., only PDF files), extend the _doc_id_metadata_filter method in private_gpt/components/vector_store/vector_store_component.py. The default implementation only supports document ID filtering, but you can add MetadataFilter entries for custom keys like "file_type" or "author" to enable arbitrary metadata-based retrieval logic.
Summary
- The low-level Chunks API at
/v1/chunksprovides direct access to the vector store indexed chunks, bypassing LLM calls for fast, cheap retrieval. - Architecture:
chunks_router.pyhandles HTTP requests,ChunksService.retrieve_relevantinchunks_service.pyexecutes the search, andVectorStoreComponentmanages the retriever lifecycle. - Key models:
ChunksBodyfor requests,Chunkfor responses (containingscore,text,documentmetadata, andprevious_texts/next_texts), andContextFilterfor document scoping. - Extensibility: Configure
prev_next_chunksfor context windows, modifyvector_store_component.pyto add metadata filters, or injectChunksServicedirectly for server-side custom logic.
Frequently Asked Questions
How do I restrict the chunks API to search only specific documents?
Pass a context_filter object with a docs_ids array in your request body. The VectorStoreComponent applies this filter during get_retriever, limiting the similarity search to nodes belonging to those specific document IDs.
What is the performance difference between the chunks API and the standard chat completions API?
The chunks API executes only a vector similarity search against the VectorStoreIndex, avoiding LLM token generation entirely. This results in significantly lower latency and zero inference costs, making it suitable for preprocessing steps or high-volume retrieval tasks.
Can I retrieve chunks with their surrounding context in a single request?
Yes. Set the prev_next_chunks parameter to the number of sibling fragments you need. The service automatically traverses the node relationships via _get_sibling_nodes_text and populates previous_texts and next_texts in the response.
How do I implement custom metadata filtering beyond document IDs?
The current ContextFilter only supports docs_ids. To filter by other metadata (e.g., file type or creation date), modify the _doc_id_metadata_filter method in vector_store_component.py to construct additional MetadataFilter conditions for the vector store query.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →