How to Configure Similarity Thresholds for Filtering RAG Retrieval Results in PrivateGPT

PrivateGPT controls RAG retrieval relevance through two tunable settings: similarity_top_k sets the maximum candidate documents fetched from the vector store, while similarity_value establishes a minimum cosine similarity cutoff that automatically triggers a SimilarityPostprocessor in the chat pipeline.

Configuring similarity thresholds for filtering RAG retrieval results ensures that only contextually relevant chunks reach the LLM, reducing hallucinations and improving response accuracy. PrivateGPT implements this through the RagSettings model, which exposes precision controls for both the breadth of initial retrieval and the strictness of post-processing filtration according to the source code in zylon-ai/private-gpt.

Understanding the Two Similarity Parameters

PrivateGPT separates retrieval control into two distinct parameters that work sequentially in the pipeline.

similarity_top_k: Limiting Candidate Volume

The similarity_top_k parameter defines the maximum number of candidate documents the vector store returns before any additional filtering occurs. Defined in the RagSettings model in private_gpt/settings/settings.py (lines 399-405), this value defaults to 2 and is passed directly to the vector index retriever.

In private_gpt/components/vector_store/vector_store_component.py (lines 199-207), the VectorStoreComponent.get_retriever method receives this parameter and instantiates a VectorIndexRetriever configured to fetch up to similarity_top_k vectors from the underlying store.

similarity_value: Enforcing Minimum Cosine Similarity

The similarity_value parameter specifies the minimum cosine similarity score (ranging from 0 to 1) that a document must achieve to be considered a valid match. When this value is set, PrivateGPT automatically injects a SimilarityPostprocessor into the chat engine pipeline. This post-processor discards any retrieved documents falling below the threshold before they are passed to the LLM.

According to the implementation in private_gpt/server/chat/chat_service.py (lines 24-29), the ChatService._chat_engine method checks if settings.rag.similarity_value is truthy, and if so, appends a SimilarityPostprocessor initialized with similarity_cutoff=settings.rag.similarity_value to the list of node post-processors.

Configuration Methods

You can define these thresholds through YAML configuration files, environment variables, or programmatically at runtime.

YAML Configuration

Add the parameters to your settings.yaml file under the rag key:

rag:
  # Retrieve the top 5 most similar vectors from the store

  similarity_top_k: 5
  # Discard any vectors with cosine similarity below 0.78

  similarity_value: 0.78
  rerank:
    enabled: false

Environment Variable Overrides

PrivateGPT supports overriding YAML values using environment variables prefixed with PGPT_RAG_:

export PGPT_RAG_SIMILARITY_TOP_K=5
export PGPT_RAG_SIMILARITY_VALUE=0.78

These variables are merged into the Settings object when the application initializes.

Runtime Programmatic Configuration

For testing or custom scripts, modify the settings object directly after initialization:

from private_gpt.settings.settings import Settings

# Load current configuration (merges YAML and environment)

settings = Settings()

# Adjust similarity thresholds

settings.rag.similarity_top_k = 8
settings.rag.similarity_value = 0.85

# ChatService automatically uses these updated values

Verifying the Post-Processor Pipeline

To confirm that the similarity filtering is active, inspect the chat engine's post-processors:

from private_gpt.server.chat.chat_service import ChatService

service = ChatService(...)
engine = service._chat_engine(use_context=True, context_filter=None)

# Inspect attached post-processors

print([type(p).__name__ for p in engine.node_postprocessors])

# Output includes 'SimilarityPostprocessor' when similarity_value is configured

If SimilarityPostprocessor appears in the list, the cutoff filter is actively screening retrieval results.

Summary

  • similarity_top_k controls the initial retrieval volume from the vector store (default: 2), configured in VectorStoreComponent.get_retriever.
  • similarity_value enables post-retrieval filtering by cosine similarity (default: None/disabled), implemented via SimilarityPostprocessor in ChatService._chat_engine.
  • Both parameters are defined in the RagSettings Pydantic model in private_gpt/settings/settings.py.
  • Configuration persists through settings.yaml, environment variables (PGPT_RAG_*), or direct modification of the Settings object.

Frequently Asked Questions

What is the default similarity filtering behavior in PrivateGPT?

By default, similarity_top_k is set to 2 and similarity_value is None. This means the system retrieves only the two most similar vectors from the store and applies no minimum similarity cutoff, passing all retrieved candidates to the LLM regardless of their relevance scores.

How does similarity_value interact with reranking modules?

The SimilarityPostprocessor executes after initial retrieval but before reranking (if enabled). Documents below the similarity_value threshold are discarded immediately, reducing the candidate pool before any reranking model processes the results. This ensures rerankers only evaluate sufficiently relevant chunks.

Can similarity filtering be disabled after being enabled?

Yes. Set similarity_value to null in your YAML configuration or omit the environment variable PGPT_RAG_SIMILARITY_VALUE. When the value is falsy (None, 0, or empty), ChatService._chat_engine skips adding the SimilarityPostprocessor, effectively disabling the cutoff filter while retaining the similarity_top_k limit.

Which source file controls the minimum similarity cutoff logic?

The enforcement logic resides in private_gpt/server/chat/chat_service.py. Specifically, lines 24-29 check for the presence of settings.rag.similarity_value and conditionally instantiate the SimilarityPostprocessor with the configured similarity_cutoff, appending it to the chat engine's node post-processors list.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →