# How Reranking Is Implemented and Configured in the PrivateGPT RAG Pipeline with Sentence Transformers

> Learn how PrivateGPT implements reranking in its RAG pipeline with Sentence Transformers. Discover how to configure this feature to reorder documents by relevance before LLM context assembly.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: how-to-guide
- Published: 2026-03-06

---

**PrivateGPT implements optional reranking in its RAG pipeline using a configurable Sentence Transformer cross-encoder that re-orders retrieved documents by relevance after vector search but before LLM context assembly.**

The PrivateGPT repository enhances Retrieval-Augmented Generation (RAG) accuracy by integrating a Sentence Transformers-based reranking layer. This component refines the candidate document set using a cross-encoder model, ensuring only the most contextually relevant passages reach the language model for generation.

## Configuration Schema for Reranking

All reranking behavior is governed by the **`RerankSettings`** Pydantic model located in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) (lines 83-96). This configuration namespace supports three primary parameters:

- **`enabled`** (bool) – Toggles the reranking step on or off.
- **`model`** (str) – The Hugging Face identifier for the Sentence Transformer cross-encoder. The default value is `cross-encoder/ms-marco-MiniLM-L-2-v2`. Only cross-encoder architectures are supported because they output a single scalar relevance score per query-document pair.
- **`top_n`** (int) – Specifies how many top-ranked documents to retain after reranking and forward to the LLM context window.

These settings are nested under the `rag` configuration key, allowing environment-specific overrides via YAML or environment variables.

## Pipeline Integration in the Chat Service

The dynamic assembly of the RAG pipeline occurs in [`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py) (lines 130-136). During query processing, the chat service conditionally instantiates the reranker based on the `settings.rag.rerank.enabled` flag:

```python
if settings.rag.rerank.enabled:
    rerank_postprocessor = SentenceTransformerRerank(
        model=settings.rag.rerank.model,
        top_n=settings.rag.rerank.top_n,
    )
    node_postprocessors.append(rerank_postprocessor)

```

When enabled, the `SentenceTransformerRerank` instance is appended to the `node_postprocessors` list. This list executes sequentially after the initial vector-store retrieval, allowing the reranker to receive the full candidate set and filter it before context construction.

## Execution Flow and Technical Implementation

The reranking step operates as a **node post-processor** within the LlamaIndex-based query engine. Understanding its internal mechanics clarifies its performance characteristics and configuration requirements.

### Cross-Encoder Scoring Mechanics

The `SentenceTransformerRerank` class wraps the **sentence-transformers** library to load the specified cross-encoder model. During execution:

1. The model receives pairs of `(query, document_text)` for every candidate retrieved from the vector store.
2. The cross-encoder outputs a single scalar relevance score for each pair, representing the semantic similarity between the question and the passage.
3. Scores are sorted in descending order, and only the highest-scoring `top_n` documents are retained.

Unlike bi-encoders used in the initial retrieval phase, cross-encoders attend to both inputs simultaneously, yielding more accurate relevance rankings at the cost of increased computational overhead.

### Performance and Resource Considerations

Because cross-encoders evaluate query-document pairs individually rather than through pre-computed embeddings, latency scales linearly with the number of candidates. The `top_n` parameter serves as a critical performance guardrail:

- **Default trade-off:** The repository defaults to `top_n=2` to balance relevance gains against inference latency.
- **Memory:** The model is loaded into memory when the chat service initializes the post-processor, consuming GPU or CPU resources proportional to the model size (e.g., MiniLM variants use approximately 50-100 MB).
- **Caching:** The sentence-transformers library automatically downloads and caches the specified model on first use to `~/.cache/torch/sentence_transformers/`.

## Enabling and Configuring Reranking

To activate reranking with a custom model and return the top three most relevant documents, modify the settings programmatically or via the [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) file:

```python
from private_gpt.settings import Settings

# Load configuration

settings = Settings()

# Enable Sentence Transformer reranking

settings.rag.rerank.enabled = True
settings.rag.rerank.model = "cross-encoder/ms-marco-MiniLM-L-12-v2"
settings.rag.rerank.top_n = 3

# The chat service will now rerank retrieved nodes before LLM generation

```

For YAML-based configuration, set the corresponding keys under the `rag` section:

```yaml
rag:
  rerank:
    enabled: true
    model: "cross-encoder/ms-marco-MiniLM-L-12-v2"
    top_n: 3

```

## Summary

- **Configuration location:** Reranking settings are defined in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) (lines 83-96) via the `RerankSettings` Pydantic model.
- **Pipeline injection:** The `SentenceTransformerRerank` post-processor is instantiated and appended to `node_postprocessors` in [`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py) (lines 130-136) when `settings.rag.rerank.enabled` is true.
- **Model requirements:** Only cross-encoder Sentence Transformer models are supported, defaulting to `cross-encoder/ms-marco-MiniLM-L-2-v2`.
- **Filtering mechanism:** The `top_n` parameter controls how many reranked documents are passed to the LLM, defaulting to 2 for latency optimization.
- **Scoring method:** The cross-encoder generates scalar relevance scores for each query-document pair, enabling finer-grained ranking than vector similarity alone.

## Frequently Asked Questions

### What is the default reranking model used by PrivateGPT?

The default model is `cross-encoder/ms-marco-MiniLM-L-2-v2`, a lightweight cross-encoder optimized for MS MARCO passage ranking tasks. This 6-layer MiniLM model provides a balance between ranking accuracy and inference speed suitable for local deployments.

### How does reranking differ from the initial vector search in PrivateGPT?

The initial vector search uses a bi-encoder embedding model to retrieve the top `k` documents via approximate nearest neighbor (ANN) search in the vector store. Reranking occurs subsequently, using a cross-encoder to rescore and reorder only those `k` candidates with higher precision, typically retaining a smaller subset (`top_n`) for the LLM context.

### Can I use a custom Hugging Face cross-encoder model for reranking?

Yes, any Sentence Transformer cross-encoder compatible with the `sentence-transformers` library can be specified via the `model` field in `RerankSettings`. Simply provide the Hugging Face repository identifier (e.g., `cross-encoder/stsb-roberta-large`) and ensure the model follows the cross-encoder architecture that outputs single-value relevance scores.

### Where is the reranking logic instantiated in the PrivateGPT codebase?

The reranking post-processor is instantiated conditionally in [`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py) at lines 130-136. This code checks the `settings.rag.rerank.enabled` boolean and, if true, creates a `SentenceTransformerRerank` instance with the configured model and `top_n` parameters before appending it to the query engine's post-processor chain.