How to Configure Similarity Thresholds for Filtering RAG Retrieval Results in PrivateGPT
PrivateGPT controls RAG retrieval relevance through two tunable settings: similarity_top_k sets the maximum candidate documents fetched from the vector store, while similarity_value establishes a minimum cosine similarity cutoff that automatically triggers a SimilarityPostprocessor in the chat pipeline.
Configuring similarity thresholds for filtering RAG retrieval results ensures that only contextually relevant chunks reach the LLM, reducing hallucinations and improving response accuracy. PrivateGPT implements this through the RagSettings model, which exposes precision controls for both the breadth of initial retrieval and the strictness of post-processing filtration according to the source code in zylon-ai/private-gpt.
Understanding the Two Similarity Parameters
PrivateGPT separates retrieval control into two distinct parameters that work sequentially in the pipeline.
similarity_top_k: Limiting Candidate Volume
The similarity_top_k parameter defines the maximum number of candidate documents the vector store returns before any additional filtering occurs. Defined in the RagSettings model in private_gpt/settings/settings.py (lines 399-405), this value defaults to 2 and is passed directly to the vector index retriever.
In private_gpt/components/vector_store/vector_store_component.py (lines 199-207), the VectorStoreComponent.get_retriever method receives this parameter and instantiates a VectorIndexRetriever configured to fetch up to similarity_top_k vectors from the underlying store.
similarity_value: Enforcing Minimum Cosine Similarity
The similarity_value parameter specifies the minimum cosine similarity score (ranging from 0 to 1) that a document must achieve to be considered a valid match. When this value is set, PrivateGPT automatically injects a SimilarityPostprocessor into the chat engine pipeline. This post-processor discards any retrieved documents falling below the threshold before they are passed to the LLM.
According to the implementation in private_gpt/server/chat/chat_service.py (lines 24-29), the ChatService._chat_engine method checks if settings.rag.similarity_value is truthy, and if so, appends a SimilarityPostprocessor initialized with similarity_cutoff=settings.rag.similarity_value to the list of node post-processors.
Configuration Methods
You can define these thresholds through YAML configuration files, environment variables, or programmatically at runtime.
YAML Configuration
Add the parameters to your settings.yaml file under the rag key:
rag:
# Retrieve the top 5 most similar vectors from the store
similarity_top_k: 5
# Discard any vectors with cosine similarity below 0.78
similarity_value: 0.78
rerank:
enabled: false
Environment Variable Overrides
PrivateGPT supports overriding YAML values using environment variables prefixed with PGPT_RAG_:
export PGPT_RAG_SIMILARITY_TOP_K=5
export PGPT_RAG_SIMILARITY_VALUE=0.78
These variables are merged into the Settings object when the application initializes.
Runtime Programmatic Configuration
For testing or custom scripts, modify the settings object directly after initialization:
from private_gpt.settings.settings import Settings
# Load current configuration (merges YAML and environment)
settings = Settings()
# Adjust similarity thresholds
settings.rag.similarity_top_k = 8
settings.rag.similarity_value = 0.85
# ChatService automatically uses these updated values
Verifying the Post-Processor Pipeline
To confirm that the similarity filtering is active, inspect the chat engine's post-processors:
from private_gpt.server.chat.chat_service import ChatService
service = ChatService(...)
engine = service._chat_engine(use_context=True, context_filter=None)
# Inspect attached post-processors
print([type(p).__name__ for p in engine.node_postprocessors])
# Output includes 'SimilarityPostprocessor' when similarity_value is configured
If SimilarityPostprocessor appears in the list, the cutoff filter is actively screening retrieval results.
Summary
similarity_top_kcontrols the initial retrieval volume from the vector store (default:2), configured inVectorStoreComponent.get_retriever.similarity_valueenables post-retrieval filtering by cosine similarity (default:None/disabled), implemented viaSimilarityPostprocessorinChatService._chat_engine.- Both parameters are defined in the
RagSettingsPydantic model inprivate_gpt/settings/settings.py. - Configuration persists through
settings.yaml, environment variables (PGPT_RAG_*), or direct modification of theSettingsobject.
Frequently Asked Questions
What is the default similarity filtering behavior in PrivateGPT?
By default, similarity_top_k is set to 2 and similarity_value is None. This means the system retrieves only the two most similar vectors from the store and applies no minimum similarity cutoff, passing all retrieved candidates to the LLM regardless of their relevance scores.
How does similarity_value interact with reranking modules?
The SimilarityPostprocessor executes after initial retrieval but before reranking (if enabled). Documents below the similarity_value threshold are discarded immediately, reducing the candidate pool before any reranking model processes the results. This ensures rerankers only evaluate sufficiently relevant chunks.
Can similarity filtering be disabled after being enabled?
Yes. Set similarity_value to null in your YAML configuration or omit the environment variable PGPT_RAG_SIMILARITY_VALUE. When the value is falsy (None, 0, or empty), ChatService._chat_engine skips adding the SimilarityPostprocessor, effectively disabling the cutoff filter while retaining the similarity_top_k limit.
Which source file controls the minimum similarity cutoff logic?
The enforcement logic resides in private_gpt/server/chat/chat_service.py. Specifically, lines 24-29 check for the presence of settings.rag.similarity_value and conditionally instantiate the SimilarityPostprocessor with the configured similarity_cutoff, appending it to the chat engine's node post-processors list.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →