# How to Use RAPTOR for RAG: Implementing Hierarchical Tree-Based Retrieval

> Learn how to use RAPTOR for RAG. Implement hierarchical tree-based retrieval for precise, context-aware answers. Explore this advanced RAG technique.

- Repository: [NirDiamant/RAG_Techniques](https://github.com/nirdiamant/rag_techniques)
- Tags: tutorial
- Published: 2026-02-19

---

**RAPTOR (Recursive Abstractive Processing and Thematic Organization for Retrieval) builds a hierarchical tree of document summaries using Gaussian Mixture Model clustering and performs contextual-compression retrieval to answer queries with structured precision.**

RAPTOR is an end-to-end retrieval-augmented generation architecture available in the [NirDiamant/RAG_Techniques](https://github.com/NirDiamant/RAG_Techniques) repository. This technique recursively abstracts document chunks into a multi-level tree structure, storing every level in a vectorstore to enable both high-level thematic retrieval and granular detail extraction. Understanding how to use RAPTOR for RAG is essential when working with large corpora where relationships between concepts span multiple documents and require hierarchical navigation.

## What is RAPTOR and Why Use It?

RAPTOR addresses the limitations of flat document chunking by creating a **recursive summarization hierarchy**. The system embeds raw text, clusters embeddings using a Gaussian Mixture Model (GMM) for soft assignments, and summarizes each cluster into parent nodes. This process repeats for `max_levels` iterations, building a tree where each node contains metadata linking it to parents and children.

Unlike standard RAG that retrieves isolated chunks, RAPTOR's **hierarchical retrieval** starts at the highest abstraction level and descends through the tree only when relevant clusters are found. This approach, combined with **contextual compression** via an `LLMChainExtractor`, ensures only relevant snippets reach the final language model, improving answer accuracy while reducing token consumption.

## Step-by-Step RAPTOR Implementation

### 1. Load and Chunk Documents

The pipeline begins by ingesting source documents. In `all_rag_techniques/raptor.ipynb`, the implementation uses `PyPDFLoader` to extract text from PDFs and splits content into raw page-level strings suitable for tree construction.

```python
from langchain_community.document_loaders import PyPDFLoader

loader = PyPDFLoader("data/Understanding_Climate_Change.pdf")
documents = loader.load()
texts = [doc.page_content for doc in documents]

```

### 2. Build the Hierarchical Summary Tree

The core logic resides in `build_raptor_tree` within [`all_rag_techniques_runnable_scripts/raptor.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/raptor.py). This function iterates through `max_levels` (default 3), performing three operations at each level:

- **Embedding**: `embed_texts` generates vector representations of current nodes
- **Clustering**: `perform_clustering` applies Gaussian Mixture Model clustering to group related embeddings with soft probabilistic assignments
- **Summarization**: `summarize_texts` invokes an LLM to generate abstractions for each cluster

```python
from all_rag_techniques_runnable_scripts.raptor import build_raptor_tree

tree_df = build_raptor_tree(
    texts=texts,
    max_levels=3,
    embedding_model="text-embedding-3-small"
)

```

The resulting `pandas.DataFrame` stores original texts, embeddings, cluster IDs, and critical **metadata** including `level`, `origin`, `parent_id`, and `child_ids`, enabling tree traversal without reprocessing embeddings.

### 3. Assemble the Global Vectorstore

Once the tree is complete, `build_vectorstore` (lines 81-97 in [`raptor.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/raptor.py)) flattens all hierarchical levels into a unified FAISS index. Every node becomes a LangChain `Document` object containing both the text content and its hierarchical metadata.

```python
from all_rag_techniques_runnable_scripts.raptor import build_vectorstore

vectorstore = build_vectorstore(tree_df)
base_retriever = vectorstore.as_retriever(search_kwargs={"k": 5})

```

### 4. Configure Contextual-Compression Retrieval

To filter irrelevant content before final generation, the system wraps the base FAISS retriever with an `LLMChainExtractor`. The `create_retriever` function (lines 69-89) implements this compression layer.

```python
from all_rag_techniques_runnable_scripts.raptor import create_retriever
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4", temperature=0)
compression_retriever = create_retriever(base_retriever, llm)

```

You can customize the extraction prompt by modifying the template in `create_retriever` to specify exactly which information to preserve.

### 5. Execute Hierarchical Retrieval

The `hierarchical_retrieval` function (lines 100-128) implements the tree-traversal logic. Starting from the highest level, it queries the vectorstore with the user question **and** level-specific metadata filters. When relevant documents are found, the function appends child IDs to the query and descends to the next level, effectively drilling down from abstract summaries to concrete details.

```python
from all_rag_techniques_runnable_scripts.raptor import hierarchical_retrieval

retrieved_docs = hierarchical_retrieval(
    query="What is the greenhouse effect?",
    retriever=compression_retriever,
    tree_metadata=tree_df
)

```

### 6. Generate the Final Answer

Finally, `raptor_query` (lines 140-170) concatenates the retrieved hierarchical snippets into a context string and prompts a `ChatOpenAI` instance to synthesize the answer. The `RAPTORMethod` class encapsulates this entire workflow.

```python
from all_rag_techniques_runnable_scripts.raptor import RAPTORMethod

raptor = RAPTORMethod(texts, max_levels=3)
result = raptor.run(query="Explain the main causes of climate change.", k=5)

print(result["answer"])
print(f"Documents retrieved: {len(result['retrieved_documents'])}")

```

For debugging, `print_query_details` (lines 184-206) outputs the tree levels of retrieved documents and their parent-child relationships.

## Complete Code Examples

### CLI Quick Start

Execute the full pipeline from the command line using the runnable script:

```bash
pip install faiss-cpu langchain langchain-openai scikit-learn python-dotenv

python all_rag_techniques_runnable_scripts/raptor.py \
    --path data/Understanding_Climate_Change.pdf \
    --query "What is the greenhouse effect?" \
    --max_levels 3

```

The script loads environment variables from `.env`, builds the tree, initializes the vectorstore, and prints the contextualized answer.

### Python API Usage

For programmatic control, instantiate the `RAPTORMethod` class:

```python
from all_rag_techniques_runnable_scripts.raptor import RAPTORMethod

# Initialize with raw texts and build tree automatically

raptor = RAPTORMethod(
    texts=["Document paragraph 1...", "Document paragraph 2..."],
    max_levels=3
)

# Run complete pipeline

result = raptor.run(
    query="Analyze the economic impacts of climate policy",
    k=5  # chunks per level

)

print("Answer:", result["answer"])
print("Sources used:", len(result["retrieved_documents"]))

```

The `__init__` method automatically calls `build_raptor_tree()` (lines 18-26), while `run()` orchestrates vectorstore creation, retrieval, and answer generation (lines 79-105).

### Customizing the Retriever Prompt

Modify the contextual-compression behavior by editing the prompt template in `create_retriever`:

```python
from langchain.prompts import ChatPromptTemplate

custom_prompt = ChatPromptTemplate.from_template(
    "Extract ONLY sentences containing quantitative data.\n"
    "Context: {context}\n"
    "Question: {question}\n"
    "Extracted data:"
)

```

Pass this prompt to your retriever configuration to enforce strict content filtering before answer generation.

## Summary

- **RAPTOR creates hierarchical trees** by recursively clustering and summarizing documents using Gaussian Mixture Models, storing parent-child relationships in node metadata.
- **Implementation requires six steps**: document loading, tree building via `build_raptor_tree`, vectorstore assembly, contextual-compression retriever setup, hierarchical traversal, and answer generation.
- **Key code locations**: Tree logic resides in [`all_rag_techniques_runnable_scripts/raptor.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/raptor.py), while the interactive demonstration is available in `all_rag_techniques/raptor.ipynb`.
- **The `RAPTORMethod` class** provides a high-level API that encapsulates the entire pipeline, automatically handling embeddings, clustering, and retrieval.
- **Contextual compression** via `LLMChainExtractor` filters irrelevant content before final answer generation, optimizing token usage and accuracy.

## Frequently Asked Questions

### What makes RAPTOR different from standard RAG?

Standard RAG retrieves flat chunks based on vector similarity, potentially missing broader thematic connections. RAPTOR organizes information into a **multi-level summary tree**, enabling retrieval to navigate from high-level concepts down to specific details. According to the `NirDiamant/RAG_Techniques` source code, this hierarchical approach uses Gaussian Mixture Model clustering to create soft thematic groupings that pure similarity search cannot capture.

### How does RAPTOR handle large document collections?

RAPTOR manages scale through **recursive abstraction**. Each tree level reduces the document set to a smaller set of summaries, allowing the system to index massive corpora without linearly increasing retrieval complexity. The `max_levels` parameter controls tree depth, and the `build_raptor_tree` function in [`raptor.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/raptor.py) processes only the current level's embeddings at each iteration, keeping memory usage bounded.

### Can I use different embedding models or LLMs with RAPTOR?

Yes. The implementation uses LangChain abstractions for all model interactions. You can swap the embedding model by modifying the `embed_texts` function calls, and change the LLM for summarization or answer generation by passing different model instances to `RAPTORMethod` or the individual functions. The code supports any LangChain-compatible embeddings or chat models.

### What are the computational costs of building the RAPTOR tree?

Tree construction involves **O(n)** embedding calls for the initial level, followed by clustering and summarization for each subsequent level. Gaussian Mixture Model clustering in `perform_clustering` adds computational overhead compared to k-means, but provides superior soft clustering. The trade-off is front-loaded: once built, retrieval is fast due to the FAISS index, though storage costs increase because the vectorstore maintains embeddings for every tree level simultaneously.