# How to Use the Low-Level Chunks API for Custom Retrieval Logic in RAG

> Master the PrivateGPT Chunks API to build custom RAG pipelines. Query document fragments directly for efficient retrieval logic without LLM inference.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: how-to-guide
- Published: 2026-03-06

---

**The low-level Chunks API in PrivateGPT exposes the `/v1/chunks` endpoint to query indexed document fragments directly, returning similarity-scored chunks with optional surrounding context for building custom RAG pipelines without LLM inference overhead.**

PrivateGPT stores every ingested document as a set of **chunks** (text fragments) indexed in a vector store. The low-level Chunks API lets you query that index directly, bypassing the LLM and returning only the most relevant fragments together with their surrounding context, making it ideal for custom retrieval-augmented generation workflows.

## Architecture of the Low-Level Chunks API

The Chunks API follows a clean layered architecture that separates HTTP handling from vector-store operations.

### HTTP Router and Request Schema

The entry point is defined in [`private_gpt/server/chunks/chunks_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chunks/chunks_router.py), which registers the `/v1/chunks` POST endpoint. The request body uses the `ChunksBody` model containing:
- `text`: The query string for similarity search
- `context_filter`: An optional `ContextFilter` to limit searches to specific document IDs
- `limit`: Number of top results to return
- `prev_next_chunks`: Number of surrounding fragments to fetch for context enrichment

The router resolves `ChunksService` from the dependency injection container and delegates to `retrieve_relevant`.

### Service Layer and Vector Store Integration

In [`private_gpt/server/chunks/chunks_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chunks/chunks_service.py), the `ChunksService.retrieve_relevant` method orchestrates the retrieval:

1. Builds a **`VectorStoreIndex`** backed by the configured vector store (`VectorStoreComponent`)
2. Creates a retriever via `VectorStoreComponent.get_retriever`, applying the optional `ContextFilter` as a document-ID filter
3. Executes a similarity-top-K search using `vector_index_retriever.retrieve(text)`

The raw `NodeWithScore` objects from the vector store are sorted by similarity score and converted into public-facing `Chunk` models.

### Chunk Conversion and Response Models

Each retrieved node is transformed into a `Chunk` object via `Chunk.from_node` (defined in [`chunks_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/chunks_service.py)), which includes:
- `object`: Always `"context.chunk"`
- `score`: The similarity score from the vector search
- `document`: An `IngestedDoc` instance (from [`private_gpt/server/ingest/model.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/ingest/model.py)) containing `doc_id` and metadata
- `text`: The fragment content
- `previous_texts` / `next_texts`: Optional sibling fragments obtained by traversing the node graph using `_get_sibling_nodes_text`

The router serializes a list of `Chunk` objects into `ChunksResponse`. Because no LLM call is involved, the endpoint delivers low-latency, cost-efficient retrieval ideal for custom preprocessing logic.

## Querying the Chunks API

### HTTP Request with curl

Send a POST request to retrieve the top 5 chunks related to sales figures, including one surrounding chunk on each side:

```bash
curl -X POST https://your-private-gpt-instance/v1/chunks \
  -H "Authorization: Bearer <your-token>" \
  -H "Content-Type: application/json" \
  -d '{
        "text": "Q3 2023 sales figures",
        "limit": 5,
        "prev_next_chunks": 1,
        "context_filter": { "docs_ids": ["c202d5e6-7b69-4869-81cc-dd574ee8ee11"] }
      }'

```

The response contains a `data` array of `Chunk` items, each with matched text, score, source document metadata, and surrounding context.

### Python Client Implementation

For programmatic access, use the `requests` library to parse and filter results:

```python
import requests

API_URL = "https://your-private-gpt-instance/v1/chunks"
TOKEN = "YOUR_API_KEY"

payload = {
    "text": "How did our marketing spend affect sales?",
    "limit": 3,
    "prev_next_chunks": 2,  # Two chunks before & after each hit

    "context_filter": {"docs_ids": ["c202d5e6-7b69-4869-81cc-dd574ee8ee11"]},
}

headers = {
    "Authorization": f"Bearer {TOKEN}",
    "Content-Type": "application/json",
}

resp = requests.post(API_URL, json=payload, headers=headers)
chunks = resp.json()["data"]

for ch in chunks:
    print(f"Score: {ch['score']:.4f}")
    print(f"Doc ID: {ch['document']['doc_id']}")
    print("Match:", ch["text"])
    if ch["previous_texts"]:
        print("Prev:", ch["previous_texts"])
    if ch["next_texts"]:
        print("Next:", ch["next_texts"])
    print("-" * 40)

```

### Direct Service Injection

When extending PrivateGPT with custom FastAPI routes or background workers, inject `ChunksService` directly to bypass HTTP overhead:

```python
from private_gpt.server.chunks.chunks_service import ChunksService

def custom_retrieval(service: ChunksService, query: str):
    # Retrieve top 10 chunks from all documents, with 2 surrounding pieces

    chunks = service.retrieve_relevant(
        text=query,
        context_filter=None,  # Search all documents

        limit=10,
        prev_next_chunks=2,
    )
    return chunks

```

Because `ChunksService` is a singleton managed by the DI container, you can reuse it across requests without recreating the vector index.

## Extending Retrieval Logic

### Filtering by Document ID

Restrict retrieval to a subset of documents using the `ContextFilter` (defined in [`private_gpt/open_ai/extensions/context_filter.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/open_ai/extensions/context_filter.py)). Populate `context_filter.docs_ids` in the request body to limit the vector search scope to specific ingested documents, reducing noise and improving relevance for domain-specific queries.

### Including Surrounding Context

Set `prev_next_chunks` to a positive integer to retrieve sibling fragments. The service calls `_get_sibling_nodes_text` to traverse the node graph and fetch preceding and following chunks, providing richer context for the LLM without requiring multiple API calls.

### Custom Metadata Filters

To filter by metadata fields (e.g., only PDF files), extend the `_doc_id_metadata_filter` method in [`private_gpt/components/vector_store/vector_store_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/vector_store/vector_store_component.py). The default implementation only supports document ID filtering, but you can add `MetadataFilter` entries for custom keys like `"file_type"` or `"author"` to enable arbitrary metadata-based retrieval logic.

## Summary

- The **low-level Chunks API** at `/v1/chunks` provides direct access to the vector store indexed chunks, bypassing LLM calls for fast, cheap retrieval.
- **Architecture**: [`chunks_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/chunks_router.py) handles HTTP requests, `ChunksService.retrieve_relevant` in [`chunks_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/chunks_service.py) executes the search, and `VectorStoreComponent` manages the retriever lifecycle.
- **Key models**: `ChunksBody` for requests, `Chunk` for responses (containing `score`, `text`, `document` metadata, and `previous_texts`/`next_texts`), and `ContextFilter` for document scoping.
- **Extensibility**: Configure `prev_next_chunks` for context windows, modify [`vector_store_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/vector_store_component.py) to add metadata filters, or inject `ChunksService` directly for server-side custom logic.

## Frequently Asked Questions

### How do I restrict the chunks API to search only specific documents?

Pass a `context_filter` object with a `docs_ids` array in your request body. The `VectorStoreComponent` applies this filter during `get_retriever`, limiting the similarity search to nodes belonging to those specific document IDs.

### What is the performance difference between the chunks API and the standard chat completions API?

The chunks API executes only a vector similarity search against the `VectorStoreIndex`, avoiding LLM token generation entirely. This results in significantly lower latency and zero inference costs, making it suitable for preprocessing steps or high-volume retrieval tasks.

### Can I retrieve chunks with their surrounding context in a single request?

Yes. Set the `prev_next_chunks` parameter to the number of sibling fragments you need. The service automatically traverses the node relationships via `_get_sibling_nodes_text` and populates `previous_texts` and `next_texts` in the response.

### How do I implement custom metadata filtering beyond document IDs?

The current `ContextFilter` only supports `docs_ids`. To filter by other metadata (e.g., file type or creation date), modify the `_doc_id_metadata_filter` method in [`vector_store_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/vector_store_component.py) to construct additional `MetadataFilter` conditions for the vector store query.