How to Use Contextual Compression in RAG: A Complete Implementation Guide

Contextual Compression in RAG filters irrelevant text from retrieved documents using an LLM compressor before generation, reducing token usage while preserving query-relevant content.

Contextual Compression is an advanced retrieval technique that shrinks the context window by removing redundant passages while keeping semantically relevant segments. In the NirDiamant/RAG_Techniques repository, this pattern is implemented by wrapping a base vector retriever with LangChain's ContextualCompressionRetriever and an LLMChainExtractor. This guide walks through the exact implementation found in the source code, including the pipeline architecture and runnable examples.

How Contextual Compression Works in the Pipeline

The implementation in all_rag_techniques_runnable_scripts/contextual_compression.py chains four distinct stages to compress context before feeding it to the LLM.

Document Ingestion and Embedding

First, the PDF is processed using the encode_pdf helper from helper_functions.py. This splits the document into chunks, generates embeddings, and stores them in a vector database. In lines 42–48 of the main script, the pipeline initializes the vector store with these embeddings, preparing the foundation for similarity search.

Base Retrieval with Vector Store

The vector store’s as_retriever() method returns the top-k most similar chunks for any given query. According to lines 45–47 in contextual_compression.py, this base retriever fetches an initial broad set of candidate documents that likely contain relevant information, even if they include extraneous text.

LLM-Based Compression

Here is where the compression occurs. Lines 49–51 instantiate LLMChainExtractor.from_llm(llm), which builds a lightweight LLM chain tasked with extracting only the query-relevant content from each retrieved chunk. This component analyzes each document against the original question, discarding paragraphs that do not contribute to the answer.

Contextual Compression Retriever and QA Chain

Lines 52–57 wrap the base retriever and compressor inside ContextualCompressionRetriever. This combined retriever automatically summarizes or filters chunks before they reach the final QA chain. Finally, lines 58–63 configure RetrievalQA.from_chain_type, which runs the LLM on the compressed context and returns the answer along with the trimmed source documents.

Implementing Contextual Compression in RAG

Below is a minimal, end-to-end example you can run from the repository’s root. It assumes your OpenAI key is stored in a .env file.


# example_contextual_compression.py

import os
from dotenv import load_dotenv
from all_rag_techniques_runnable_scripts.contextual_compression import ContextualCompressionRAG

# Load .env (contains OPENAI_API_KEY)

load_dotenv()
os.environ["OPENAI_API_KEY"] = os.getenv("OPENAI_API_KEY")

# 1️⃣ Initialise the RAG pipeline

rag = ContextualCompressionRAG(
    path="data/Understanding_Climate_Change.pdf",   # any PDF you want to query

    model_name="gpt-4o-mini",                      # cheap, fast model

    temperature=0.0,
    max_tokens=4000,
)

# 2️⃣ Run a query – the compressor will trim the retrieved chunks

query = "What are the main causes of climate change according to the report?"
answer, latency = rag.run(query)

print("\n=== ANSWER ===")
print(answer["result"])
print("\n=== SOURCE DOCUMENTS (excerpt) ===")
for doc in answer["source_documents"]:
    print("-" * 40)
    print(doc.page_content[:300] + "...")
print(f"\nExecution time: {latency:.2f}s")

Run the script with:

python example_contextual_compression.py

For an interactive walkthrough, use the notebook at all_rag_techniques/contextual_compression.ipynb, which visualizes retrieved versus compressed chunks with step-by-step commentary.

Key Source Files and Architecture

Understanding the repository structure helps customize the compression logic for your use case.

  • all_rag_techniques_runnable_scripts/contextual_compression.py – Contains the ContextualCompressionRAG class that wires together the vector store, base retriever, LLMChainExtractor, and QA chain.
  • helper_functions.py – Provides encode_pdf and other utilities for document chunking and embedding.
  • all_rag_techniques/contextual_compression.ipynb – Interactive notebook demonstrating the technique with visualizations.
  • evaluation/evalute_rag.py – Evaluation utilities to benchmark compression against baseline RAG approaches.

The architectural flow follows this pattern: PDF ingestion → chunking → embedding storage → base retrieval → LLM compression → final QA generation. This ensures only the most salient text reaches the expensive LLM call.

Summary

  • Contextual Compression in RAG reduces token costs by filtering retrieved documents through an LLM compressor before generation.
  • The pipeline uses LLMChainExtractor to extract query-relevant snippets and ContextualCompressionRetriever to orchestrate the compression.
  • Implementation resides in all_rag_techniques_runnable_scripts/contextual_compression.py, utilizing encode_pdf for ingestion and RetrievalQA for final answer synthesis.
  • Adjust retriever.search_kwargs["k"] to balance between retrieval recall and compression speed.

Frequently Asked Questions

What is the difference between standard RAG and Contextual Compression RAG?

Standard RAG feeds entire retrieved chunks to the LLM, often including irrelevant text that consumes tokens and dilutes focus. Contextual Compression RAG adds an intermediate LLM step that extracts only the sentences relevant to the query, resulting in more focused answers and lower API costs.

Which LangChain classes handle the compression logic?

The implementation relies on two primary LangChain components: LLMChainExtractor.from_llm() builds the compressor that trims individual documents, and ContextualCompressionRetriever wraps the base retriever to apply this compression automatically across all retrieved chunks.

How does the encode_pdf helper function support the pipeline?

Located in helper_functions.py, encode_pdf handles the initial document ingestion phase by splitting PDFs into chunks, generating embeddings, and storing them in the vector database. This prepares the data layer that the ContextualCompressionRetriever queries during the retrieval phase.

Can I adjust the number of chunks retrieved before compression?

Yes. Modify retriever.search_kwargs["k"] in the ContextualCompressionRAG.__init__ method to control how many initial chunks the base retriever fetches. Fetching fewer chunks speeds up the compression step but may miss useful context, while more chunks increase recall at the cost of higher compression latency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →