# How to Configure Similarity Thresholds for Filtering RAG Retrieval Results in PrivateGPT

> Configure PrivateGPT similarity thresholds to filter RAG retrieval results. Learn how similarity_top_k and similarity_value enhance relevancy and optimize your chatbot's performance.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: how-to-guide
- Published: 2026-03-06

---

**PrivateGPT controls RAG retrieval relevance through two tunable settings: `similarity_top_k` sets the maximum candidate documents fetched from the vector store, while `similarity_value` establishes a minimum cosine similarity cutoff that automatically triggers a `SimilarityPostprocessor` in the chat pipeline.**

Configuring similarity thresholds for filtering RAG retrieval results ensures that only contextually relevant chunks reach the LLM, reducing hallucinations and improving response accuracy. PrivateGPT implements this through the `RagSettings` model, which exposes precision controls for both the breadth of initial retrieval and the strictness of post-processing filtration according to the source code in `zylon-ai/private-gpt`.

## Understanding the Two Similarity Parameters

PrivateGPT separates retrieval control into two distinct parameters that work sequentially in the pipeline.

### similarity_top_k: Limiting Candidate Volume

The **`similarity_top_k`** parameter defines the maximum number of candidate documents the vector store returns before any additional filtering occurs. Defined in the `RagSettings` model in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) (lines 399-405), this value defaults to `2` and is passed directly to the vector index retriever.

In [`private_gpt/components/vector_store/vector_store_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/vector_store/vector_store_component.py) (lines 199-207), the `VectorStoreComponent.get_retriever` method receives this parameter and instantiates a `VectorIndexRetriever` configured to fetch up to `similarity_top_k` vectors from the underlying store.

### similarity_value: Enforcing Minimum Cosine Similarity

The **`similarity_value`** parameter specifies the minimum cosine similarity score (ranging from 0 to 1) that a document must achieve to be considered a valid match. When this value is set, PrivateGPT automatically injects a **`SimilarityPostprocessor`** into the chat engine pipeline. This post-processor discards any retrieved documents falling below the threshold before they are passed to the LLM.

According to the implementation in [`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py) (lines 24-29), the `ChatService._chat_engine` method checks if `settings.rag.similarity_value` is truthy, and if so, appends a `SimilarityPostprocessor` initialized with `similarity_cutoff=settings.rag.similarity_value` to the list of node post-processors.

## Configuration Methods

You can define these thresholds through YAML configuration files, environment variables, or programmatically at runtime.

### YAML Configuration

Add the parameters to your [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) file under the `rag` key:

```yaml
rag:
  # Retrieve the top 5 most similar vectors from the store

  similarity_top_k: 5
  # Discard any vectors with cosine similarity below 0.78

  similarity_value: 0.78
  rerank:
    enabled: false

```

### Environment Variable Overrides

PrivateGPT supports overriding YAML values using environment variables prefixed with `PGPT_RAG_`:

```bash
export PGPT_RAG_SIMILARITY_TOP_K=5
export PGPT_RAG_SIMILARITY_VALUE=0.78

```

These variables are merged into the `Settings` object when the application initializes.

### Runtime Programmatic Configuration

For testing or custom scripts, modify the settings object directly after initialization:

```python
from private_gpt.settings.settings import Settings

# Load current configuration (merges YAML and environment)

settings = Settings()

# Adjust similarity thresholds

settings.rag.similarity_top_k = 8
settings.rag.similarity_value = 0.85

# ChatService automatically uses these updated values

```

## Verifying the Post-Processor Pipeline

To confirm that the similarity filtering is active, inspect the chat engine's post-processors:

```python
from private_gpt.server.chat.chat_service import ChatService

service = ChatService(...)
engine = service._chat_engine(use_context=True, context_filter=None)

# Inspect attached post-processors

print([type(p).__name__ for p in engine.node_postprocessors])

# Output includes 'SimilarityPostprocessor' when similarity_value is configured

```

If `SimilarityPostprocessor` appears in the list, the cutoff filter is actively screening retrieval results.

## Summary

- **`similarity_top_k`** controls the initial retrieval volume from the vector store (default: `2`), configured in `VectorStoreComponent.get_retriever`.
- **`similarity_value`** enables post-retrieval filtering by cosine similarity (default: `None`/`disabled`), implemented via `SimilarityPostprocessor` in `ChatService._chat_engine`.
- Both parameters are defined in the `RagSettings` Pydantic model in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py).
- Configuration persists through [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml), environment variables (`PGPT_RAG_*`), or direct modification of the `Settings` object.

## Frequently Asked Questions

### What is the default similarity filtering behavior in PrivateGPT?

By default, `similarity_top_k` is set to `2` and `similarity_value` is `None`. This means the system retrieves only the two most similar vectors from the store and applies no minimum similarity cutoff, passing all retrieved candidates to the LLM regardless of their relevance scores.

### How does `similarity_value` interact with reranking modules?

The `SimilarityPostprocessor` executes after initial retrieval but before reranking (if enabled). Documents below the `similarity_value` threshold are discarded immediately, reducing the candidate pool before any reranking model processes the results. This ensures rerankers only evaluate sufficiently relevant chunks.

### Can similarity filtering be disabled after being enabled?

Yes. Set `similarity_value` to `null` in your YAML configuration or omit the environment variable `PGPT_RAG_SIMILARITY_VALUE`. When the value is falsy (None, 0, or empty), `ChatService._chat_engine` skips adding the `SimilarityPostprocessor`, effectively disabling the cutoff filter while retaining the `similarity_top_k` limit.

### Which source file controls the minimum similarity cutoff logic?

The enforcement logic resides in [`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py). Specifically, lines 24-29 check for the presence of `settings.rag.similarity_value` and conditionally instantiate the `SimilarityPostprocessor` with the configured `similarity_cutoff`, appending it to the chat engine's node post-processors list.