How Reranking Is Implemented and Configured in the PrivateGPT RAG Pipeline with Sentence Transformers

PrivateGPT implements optional reranking in its RAG pipeline using a configurable Sentence Transformer cross-encoder that re-orders retrieved documents by relevance after vector search but before LLM context assembly.

The PrivateGPT repository enhances Retrieval-Augmented Generation (RAG) accuracy by integrating a Sentence Transformers-based reranking layer. This component refines the candidate document set using a cross-encoder model, ensuring only the most contextually relevant passages reach the language model for generation.

Configuration Schema for Reranking

All reranking behavior is governed by the RerankSettings Pydantic model located in private_gpt/settings/settings.py (lines 83-96). This configuration namespace supports three primary parameters:

  • enabled (bool) – Toggles the reranking step on or off.
  • model (str) – The Hugging Face identifier for the Sentence Transformer cross-encoder. The default value is cross-encoder/ms-marco-MiniLM-L-2-v2. Only cross-encoder architectures are supported because they output a single scalar relevance score per query-document pair.
  • top_n (int) – Specifies how many top-ranked documents to retain after reranking and forward to the LLM context window.

These settings are nested under the rag configuration key, allowing environment-specific overrides via YAML or environment variables.

Pipeline Integration in the Chat Service

The dynamic assembly of the RAG pipeline occurs in private_gpt/server/chat/chat_service.py (lines 130-136). During query processing, the chat service conditionally instantiates the reranker based on the settings.rag.rerank.enabled flag:

if settings.rag.rerank.enabled:
    rerank_postprocessor = SentenceTransformerRerank(
        model=settings.rag.rerank.model,
        top_n=settings.rag.rerank.top_n,
    )
    node_postprocessors.append(rerank_postprocessor)

When enabled, the SentenceTransformerRerank instance is appended to the node_postprocessors list. This list executes sequentially after the initial vector-store retrieval, allowing the reranker to receive the full candidate set and filter it before context construction.

Execution Flow and Technical Implementation

The reranking step operates as a node post-processor within the LlamaIndex-based query engine. Understanding its internal mechanics clarifies its performance characteristics and configuration requirements.

Cross-Encoder Scoring Mechanics

The SentenceTransformerRerank class wraps the sentence-transformers library to load the specified cross-encoder model. During execution:

  1. The model receives pairs of (query, document_text) for every candidate retrieved from the vector store.
  2. The cross-encoder outputs a single scalar relevance score for each pair, representing the semantic similarity between the question and the passage.
  3. Scores are sorted in descending order, and only the highest-scoring top_n documents are retained.

Unlike bi-encoders used in the initial retrieval phase, cross-encoders attend to both inputs simultaneously, yielding more accurate relevance rankings at the cost of increased computational overhead.

Performance and Resource Considerations

Because cross-encoders evaluate query-document pairs individually rather than through pre-computed embeddings, latency scales linearly with the number of candidates. The top_n parameter serves as a critical performance guardrail:

  • Default trade-off: The repository defaults to top_n=2 to balance relevance gains against inference latency.
  • Memory: The model is loaded into memory when the chat service initializes the post-processor, consuming GPU or CPU resources proportional to the model size (e.g., MiniLM variants use approximately 50-100 MB).
  • Caching: The sentence-transformers library automatically downloads and caches the specified model on first use to ~/.cache/torch/sentence_transformers/.

Enabling and Configuring Reranking

To activate reranking with a custom model and return the top three most relevant documents, modify the settings programmatically or via the settings.yaml file:

from private_gpt.settings import Settings

# Load configuration

settings = Settings()

# Enable Sentence Transformer reranking

settings.rag.rerank.enabled = True
settings.rag.rerank.model = "cross-encoder/ms-marco-MiniLM-L-12-v2"
settings.rag.rerank.top_n = 3

# The chat service will now rerank retrieved nodes before LLM generation

For YAML-based configuration, set the corresponding keys under the rag section:

rag:
  rerank:
    enabled: true
    model: "cross-encoder/ms-marco-MiniLM-L-12-v2"
    top_n: 3

Summary

  • Configuration location: Reranking settings are defined in private_gpt/settings/settings.py (lines 83-96) via the RerankSettings Pydantic model.
  • Pipeline injection: The SentenceTransformerRerank post-processor is instantiated and appended to node_postprocessors in private_gpt/server/chat/chat_service.py (lines 130-136) when settings.rag.rerank.enabled is true.
  • Model requirements: Only cross-encoder Sentence Transformer models are supported, defaulting to cross-encoder/ms-marco-MiniLM-L-2-v2.
  • Filtering mechanism: The top_n parameter controls how many reranked documents are passed to the LLM, defaulting to 2 for latency optimization.
  • Scoring method: The cross-encoder generates scalar relevance scores for each query-document pair, enabling finer-grained ranking than vector similarity alone.

Frequently Asked Questions

What is the default reranking model used by PrivateGPT?

The default model is cross-encoder/ms-marco-MiniLM-L-2-v2, a lightweight cross-encoder optimized for MS MARCO passage ranking tasks. This 6-layer MiniLM model provides a balance between ranking accuracy and inference speed suitable for local deployments.

How does reranking differ from the initial vector search in PrivateGPT?

The initial vector search uses a bi-encoder embedding model to retrieve the top k documents via approximate nearest neighbor (ANN) search in the vector store. Reranking occurs subsequently, using a cross-encoder to rescore and reorder only those k candidates with higher precision, typically retaining a smaller subset (top_n) for the LLM context.

Can I use a custom Hugging Face cross-encoder model for reranking?

Yes, any Sentence Transformer cross-encoder compatible with the sentence-transformers library can be specified via the model field in RerankSettings. Simply provide the Hugging Face repository identifier (e.g., cross-encoder/stsb-roberta-large) and ensure the model follows the cross-encoder architecture that outputs single-value relevance scores.

Where is the reranking logic instantiated in the PrivateGPT codebase?

The reranking post-processor is instantiated conditionally in private_gpt/server/chat/chat_service.py at lines 130-136. This code checks the settings.rag.rerank.enabled boolean and, if true, creates a SentenceTransformerRerank instance with the configured model and top_n parameters before appending it to the query engine's post-processor chain.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →