# How to Use Contextual Compression in RAG: A Complete Implementation Guide

> Master Contextual Compression in RAG with our implementation guide. Filter irrelevant text using an LLM compressor to enhance RAG performance and reduce token usage for better results.

- Repository: [NirDiamant/RAG_Techniques](https://github.com/nirdiamant/rag_techniques)
- Tags: how-to-guide
- Published: 2026-02-19

---

**Contextual Compression in RAG** filters irrelevant text from retrieved documents using an LLM compressor before generation, reducing token usage while preserving query-relevant content.

Contextual Compression is an advanced retrieval technique that shrinks the context window by removing redundant passages while keeping semantically relevant segments. In the **NirDiamant/RAG_Techniques** repository, this pattern is implemented by wrapping a base vector retriever with LangChain's `ContextualCompressionRetriever` and an `LLMChainExtractor`. This guide walks through the exact implementation found in the source code, including the pipeline architecture and runnable examples.

## How Contextual Compression Works in the Pipeline

The implementation in [`all_rag_techniques_runnable_scripts/contextual_compression.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/contextual_compression.py) chains four distinct stages to compress context before feeding it to the LLM.

### Document Ingestion and Embedding

First, the PDF is processed using the `encode_pdf` helper from [`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py). This splits the document into chunks, generates embeddings, and stores them in a vector database. In lines 42–48 of the main script, the pipeline initializes the vector store with these embeddings, preparing the foundation for similarity search.

### Base Retrieval with Vector Store

The vector store’s `as_retriever()` method returns the top-k most similar chunks for any given query. According to lines 45–47 in [`contextual_compression.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/contextual_compression.py), this base retriever fetches an initial broad set of candidate documents that likely contain relevant information, even if they include extraneous text.

### LLM-Based Compression

Here is where the compression occurs. Lines 49–51 instantiate `LLMChainExtractor.from_llm(llm)`, which builds a lightweight LLM chain tasked with extracting only the query-relevant content from each retrieved chunk. This component analyzes each document against the original question, discarding paragraphs that do not contribute to the answer.

### Contextual Compression Retriever and QA Chain

Lines 52–57 wrap the base retriever and compressor inside `ContextualCompressionRetriever`. This combined retriever automatically summarizes or filters chunks before they reach the final QA chain. Finally, lines 58–63 configure `RetrievalQA.from_chain_type`, which runs the LLM on the compressed context and returns the answer along with the trimmed source documents.

## Implementing Contextual Compression in RAG

Below is a minimal, end-to-end example you can run from the repository’s root. It assumes your OpenAI key is stored in a `.env` file.

```python

# example_contextual_compression.py

import os
from dotenv import load_dotenv
from all_rag_techniques_runnable_scripts.contextual_compression import ContextualCompressionRAG

# Load .env (contains OPENAI_API_KEY)

load_dotenv()
os.environ["OPENAI_API_KEY"] = os.getenv("OPENAI_API_KEY")

# 1️⃣ Initialise the RAG pipeline

rag = ContextualCompressionRAG(
    path="data/Understanding_Climate_Change.pdf",   # any PDF you want to query

    model_name="gpt-4o-mini",                      # cheap, fast model

    temperature=0.0,
    max_tokens=4000,
)

# 2️⃣ Run a query – the compressor will trim the retrieved chunks

query = "What are the main causes of climate change according to the report?"
answer, latency = rag.run(query)

print("\n=== ANSWER ===")
print(answer["result"])
print("\n=== SOURCE DOCUMENTS (excerpt) ===")
for doc in answer["source_documents"]:
    print("-" * 40)
    print(doc.page_content[:300] + "...")
print(f"\nExecution time: {latency:.2f}s")

```

Run the script with:

```bash
python example_contextual_compression.py

```

For an interactive walkthrough, use the notebook at `all_rag_techniques/contextual_compression.ipynb`, which visualizes retrieved versus compressed chunks with step-by-step commentary.

## Key Source Files and Architecture

Understanding the repository structure helps customize the compression logic for your use case.

- **[`all_rag_techniques_runnable_scripts/contextual_compression.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/contextual_compression.py)** – Contains the `ContextualCompressionRAG` class that wires together the vector store, base retriever, `LLMChainExtractor`, and QA chain.
- **[`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py)** – Provides `encode_pdf` and other utilities for document chunking and embedding.
- **`all_rag_techniques/contextual_compression.ipynb`** – Interactive notebook demonstrating the technique with visualizations.
- **[`evaluation/evalute_rag.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/evaluation/evalute_rag.py)** – Evaluation utilities to benchmark compression against baseline RAG approaches.

The architectural flow follows this pattern: PDF ingestion → chunking → embedding storage → base retrieval → LLM compression → final QA generation. This ensures only the most salient text reaches the expensive LLM call.

## Summary

- **Contextual Compression in RAG** reduces token costs by filtering retrieved documents through an LLM compressor before generation.
- The pipeline uses `LLMChainExtractor` to extract query-relevant snippets and `ContextualCompressionRetriever` to orchestrate the compression.
- Implementation resides in [`all_rag_techniques_runnable_scripts/contextual_compression.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/contextual_compression.py), utilizing `encode_pdf` for ingestion and `RetrievalQA` for final answer synthesis.
- Adjust `retriever.search_kwargs["k"]` to balance between retrieval recall and compression speed.

## Frequently Asked Questions

### What is the difference between standard RAG and Contextual Compression RAG?

Standard RAG feeds entire retrieved chunks to the LLM, often including irrelevant text that consumes tokens and dilutes focus. Contextual Compression RAG adds an intermediate LLM step that extracts only the sentences relevant to the query, resulting in more focused answers and lower API costs.

### Which LangChain classes handle the compression logic?

The implementation relies on two primary LangChain components: `LLMChainExtractor.from_llm()` builds the compressor that trims individual documents, and `ContextualCompressionRetriever` wraps the base retriever to apply this compression automatically across all retrieved chunks.

### How does the `encode_pdf` helper function support the pipeline?

Located in [`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py), `encode_pdf` handles the initial document ingestion phase by splitting PDFs into chunks, generating embeddings, and storing them in the vector database. This prepares the data layer that the `ContextualCompressionRetriever` queries during the retrieval phase.

### Can I adjust the number of chunks retrieved before compression?

Yes. Modify `retriever.search_kwargs["k"]` in the `ContextualCompressionRAG.__init__` method to control how many initial chunks the base retriever fetches. Fetching fewer chunks speeds up the compression step but may miss useful context, while more chunks increase recall at the cost of higher compression latency.