How to Use the Low-Level Chunks API for Custom Retrieval Logic in RAG

The low-level Chunks API in PrivateGPT exposes the /v1/chunks endpoint to query indexed document fragments directly, returning similarity-scored chunks with optional surrounding context for building custom RAG pipelines without LLM inference overhead.

PrivateGPT stores every ingested document as a set of chunks (text fragments) indexed in a vector store. The low-level Chunks API lets you query that index directly, bypassing the LLM and returning only the most relevant fragments together with their surrounding context, making it ideal for custom retrieval-augmented generation workflows.

Architecture of the Low-Level Chunks API

The Chunks API follows a clean layered architecture that separates HTTP handling from vector-store operations.

HTTP Router and Request Schema

The entry point is defined in private_gpt/server/chunks/chunks_router.py, which registers the /v1/chunks POST endpoint. The request body uses the ChunksBody model containing:

  • text: The query string for similarity search
  • context_filter: An optional ContextFilter to limit searches to specific document IDs
  • limit: Number of top results to return
  • prev_next_chunks: Number of surrounding fragments to fetch for context enrichment

The router resolves ChunksService from the dependency injection container and delegates to retrieve_relevant.

Service Layer and Vector Store Integration

In private_gpt/server/chunks/chunks_service.py, the ChunksService.retrieve_relevant method orchestrates the retrieval:

  1. Builds a VectorStoreIndex backed by the configured vector store (VectorStoreComponent)
  2. Creates a retriever via VectorStoreComponent.get_retriever, applying the optional ContextFilter as a document-ID filter
  3. Executes a similarity-top-K search using vector_index_retriever.retrieve(text)

The raw NodeWithScore objects from the vector store are sorted by similarity score and converted into public-facing Chunk models.

Chunk Conversion and Response Models

Each retrieved node is transformed into a Chunk object via Chunk.from_node (defined in chunks_service.py), which includes:

  • object: Always "context.chunk"
  • score: The similarity score from the vector search
  • document: An IngestedDoc instance (from private_gpt/server/ingest/model.py) containing doc_id and metadata
  • text: The fragment content
  • previous_texts / next_texts: Optional sibling fragments obtained by traversing the node graph using _get_sibling_nodes_text

The router serializes a list of Chunk objects into ChunksResponse. Because no LLM call is involved, the endpoint delivers low-latency, cost-efficient retrieval ideal for custom preprocessing logic.

Querying the Chunks API

HTTP Request with curl

Send a POST request to retrieve the top 5 chunks related to sales figures, including one surrounding chunk on each side:

curl -X POST https://your-private-gpt-instance/v1/chunks \
  -H "Authorization: Bearer <your-token>" \
  -H "Content-Type: application/json" \
  -d '{
        "text": "Q3 2023 sales figures",
        "limit": 5,
        "prev_next_chunks": 1,
        "context_filter": { "docs_ids": ["c202d5e6-7b69-4869-81cc-dd574ee8ee11"] }
      }'

The response contains a data array of Chunk items, each with matched text, score, source document metadata, and surrounding context.

Python Client Implementation

For programmatic access, use the requests library to parse and filter results:

import requests

API_URL = "https://your-private-gpt-instance/v1/chunks"
TOKEN = "YOUR_API_KEY"

payload = {
    "text": "How did our marketing spend affect sales?",
    "limit": 3,
    "prev_next_chunks": 2,  # Two chunks before & after each hit

    "context_filter": {"docs_ids": ["c202d5e6-7b69-4869-81cc-dd574ee8ee11"]},
}

headers = {
    "Authorization": f"Bearer {TOKEN}",
    "Content-Type": "application/json",
}

resp = requests.post(API_URL, json=payload, headers=headers)
chunks = resp.json()["data"]

for ch in chunks:
    print(f"Score: {ch['score']:.4f}")
    print(f"Doc ID: {ch['document']['doc_id']}")
    print("Match:", ch["text"])
    if ch["previous_texts"]:
        print("Prev:", ch["previous_texts"])
    if ch["next_texts"]:
        print("Next:", ch["next_texts"])
    print("-" * 40)

Direct Service Injection

When extending PrivateGPT with custom FastAPI routes or background workers, inject ChunksService directly to bypass HTTP overhead:

from private_gpt.server.chunks.chunks_service import ChunksService

def custom_retrieval(service: ChunksService, query: str):
    # Retrieve top 10 chunks from all documents, with 2 surrounding pieces

    chunks = service.retrieve_relevant(
        text=query,
        context_filter=None,  # Search all documents

        limit=10,
        prev_next_chunks=2,
    )
    return chunks

Because ChunksService is a singleton managed by the DI container, you can reuse it across requests without recreating the vector index.

Extending Retrieval Logic

Filtering by Document ID

Restrict retrieval to a subset of documents using the ContextFilter (defined in private_gpt/open_ai/extensions/context_filter.py). Populate context_filter.docs_ids in the request body to limit the vector search scope to specific ingested documents, reducing noise and improving relevance for domain-specific queries.

Including Surrounding Context

Set prev_next_chunks to a positive integer to retrieve sibling fragments. The service calls _get_sibling_nodes_text to traverse the node graph and fetch preceding and following chunks, providing richer context for the LLM without requiring multiple API calls.

Custom Metadata Filters

To filter by metadata fields (e.g., only PDF files), extend the _doc_id_metadata_filter method in private_gpt/components/vector_store/vector_store_component.py. The default implementation only supports document ID filtering, but you can add MetadataFilter entries for custom keys like "file_type" or "author" to enable arbitrary metadata-based retrieval logic.

Summary

  • The low-level Chunks API at /v1/chunks provides direct access to the vector store indexed chunks, bypassing LLM calls for fast, cheap retrieval.
  • Architecture: chunks_router.py handles HTTP requests, ChunksService.retrieve_relevant in chunks_service.py executes the search, and VectorStoreComponent manages the retriever lifecycle.
  • Key models: ChunksBody for requests, Chunk for responses (containing score, text, document metadata, and previous_texts/next_texts), and ContextFilter for document scoping.
  • Extensibility: Configure prev_next_chunks for context windows, modify vector_store_component.py to add metadata filters, or inject ChunksService directly for server-side custom logic.

Frequently Asked Questions

How do I restrict the chunks API to search only specific documents?

Pass a context_filter object with a docs_ids array in your request body. The VectorStoreComponent applies this filter during get_retriever, limiting the similarity search to nodes belonging to those specific document IDs.

What is the performance difference between the chunks API and the standard chat completions API?

The chunks API executes only a vector similarity search against the VectorStoreIndex, avoiding LLM token generation entirely. This results in significantly lower latency and zero inference costs, making it suitable for preprocessing steps or high-volume retrieval tasks.

Can I retrieve chunks with their surrounding context in a single request?

Yes. Set the prev_next_chunks parameter to the number of sibling fragments you need. The service automatically traverses the node relationships via _get_sibling_nodes_text and populates previous_texts and next_texts in the response.

How do I implement custom metadata filtering beyond document IDs?

The current ContextFilter only supports docs_ids. To filter by other metadata (e.g., file type or creation date), modify the _doc_id_metadata_filter method in vector_store_component.py to construct additional MetadataFilter conditions for the vector store query.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →