How to Use RAPTOR for RAG: Implementing Hierarchical Tree-Based Retrieval

RAPTOR (Recursive Abstractive Processing and Thematic Organization for Retrieval) builds a hierarchical tree of document summaries using Gaussian Mixture Model clustering and performs contextual-compression retrieval to answer queries with structured precision.

RAPTOR is an end-to-end retrieval-augmented generation architecture available in the NirDiamant/RAG_Techniques repository. This technique recursively abstracts document chunks into a multi-level tree structure, storing every level in a vectorstore to enable both high-level thematic retrieval and granular detail extraction. Understanding how to use RAPTOR for RAG is essential when working with large corpora where relationships between concepts span multiple documents and require hierarchical navigation.

What is RAPTOR and Why Use It?

RAPTOR addresses the limitations of flat document chunking by creating a recursive summarization hierarchy. The system embeds raw text, clusters embeddings using a Gaussian Mixture Model (GMM) for soft assignments, and summarizes each cluster into parent nodes. This process repeats for max_levels iterations, building a tree where each node contains metadata linking it to parents and children.

Unlike standard RAG that retrieves isolated chunks, RAPTOR's hierarchical retrieval starts at the highest abstraction level and descends through the tree only when relevant clusters are found. This approach, combined with contextual compression via an LLMChainExtractor, ensures only relevant snippets reach the final language model, improving answer accuracy while reducing token consumption.

Step-by-Step RAPTOR Implementation

1. Load and Chunk Documents

The pipeline begins by ingesting source documents. In all_rag_techniques/raptor.ipynb, the implementation uses PyPDFLoader to extract text from PDFs and splits content into raw page-level strings suitable for tree construction.

from langchain_community.document_loaders import PyPDFLoader

loader = PyPDFLoader("data/Understanding_Climate_Change.pdf")
documents = loader.load()
texts = [doc.page_content for doc in documents]

2. Build the Hierarchical Summary Tree

The core logic resides in build_raptor_tree within all_rag_techniques_runnable_scripts/raptor.py. This function iterates through max_levels (default 3), performing three operations at each level:

  • Embedding: embed_texts generates vector representations of current nodes
  • Clustering: perform_clustering applies Gaussian Mixture Model clustering to group related embeddings with soft probabilistic assignments
  • Summarization: summarize_texts invokes an LLM to generate abstractions for each cluster
from all_rag_techniques_runnable_scripts.raptor import build_raptor_tree

tree_df = build_raptor_tree(
    texts=texts,
    max_levels=3,
    embedding_model="text-embedding-3-small"
)

The resulting pandas.DataFrame stores original texts, embeddings, cluster IDs, and critical metadata including level, origin, parent_id, and child_ids, enabling tree traversal without reprocessing embeddings.

3. Assemble the Global Vectorstore

Once the tree is complete, build_vectorstore (lines 81-97 in raptor.py) flattens all hierarchical levels into a unified FAISS index. Every node becomes a LangChain Document object containing both the text content and its hierarchical metadata.

from all_rag_techniques_runnable_scripts.raptor import build_vectorstore

vectorstore = build_vectorstore(tree_df)
base_retriever = vectorstore.as_retriever(search_kwargs={"k": 5})

4. Configure Contextual-Compression Retrieval

To filter irrelevant content before final generation, the system wraps the base FAISS retriever with an LLMChainExtractor. The create_retriever function (lines 69-89) implements this compression layer.

from all_rag_techniques_runnable_scripts.raptor import create_retriever
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4", temperature=0)
compression_retriever = create_retriever(base_retriever, llm)

You can customize the extraction prompt by modifying the template in create_retriever to specify exactly which information to preserve.

5. Execute Hierarchical Retrieval

The hierarchical_retrieval function (lines 100-128) implements the tree-traversal logic. Starting from the highest level, it queries the vectorstore with the user question and level-specific metadata filters. When relevant documents are found, the function appends child IDs to the query and descends to the next level, effectively drilling down from abstract summaries to concrete details.

from all_rag_techniques_runnable_scripts.raptor import hierarchical_retrieval

retrieved_docs = hierarchical_retrieval(
    query="What is the greenhouse effect?",
    retriever=compression_retriever,
    tree_metadata=tree_df
)

6. Generate the Final Answer

Finally, raptor_query (lines 140-170) concatenates the retrieved hierarchical snippets into a context string and prompts a ChatOpenAI instance to synthesize the answer. The RAPTORMethod class encapsulates this entire workflow.

from all_rag_techniques_runnable_scripts.raptor import RAPTORMethod

raptor = RAPTORMethod(texts, max_levels=3)
result = raptor.run(query="Explain the main causes of climate change.", k=5)

print(result["answer"])
print(f"Documents retrieved: {len(result['retrieved_documents'])}")

For debugging, print_query_details (lines 184-206) outputs the tree levels of retrieved documents and their parent-child relationships.

Complete Code Examples

CLI Quick Start

Execute the full pipeline from the command line using the runnable script:

pip install faiss-cpu langchain langchain-openai scikit-learn python-dotenv

python all_rag_techniques_runnable_scripts/raptor.py \
    --path data/Understanding_Climate_Change.pdf \
    --query "What is the greenhouse effect?" \
    --max_levels 3

The script loads environment variables from .env, builds the tree, initializes the vectorstore, and prints the contextualized answer.

Python API Usage

For programmatic control, instantiate the RAPTORMethod class:

from all_rag_techniques_runnable_scripts.raptor import RAPTORMethod

# Initialize with raw texts and build tree automatically

raptor = RAPTORMethod(
    texts=["Document paragraph 1...", "Document paragraph 2..."],
    max_levels=3
)

# Run complete pipeline

result = raptor.run(
    query="Analyze the economic impacts of climate policy",
    k=5  # chunks per level

)

print("Answer:", result["answer"])
print("Sources used:", len(result["retrieved_documents"]))

The __init__ method automatically calls build_raptor_tree() (lines 18-26), while run() orchestrates vectorstore creation, retrieval, and answer generation (lines 79-105).

Customizing the Retriever Prompt

Modify the contextual-compression behavior by editing the prompt template in create_retriever:

from langchain.prompts import ChatPromptTemplate

custom_prompt = ChatPromptTemplate.from_template(
    "Extract ONLY sentences containing quantitative data.\n"
    "Context: {context}\n"
    "Question: {question}\n"
    "Extracted data:"
)

Pass this prompt to your retriever configuration to enforce strict content filtering before answer generation.

Summary

  • RAPTOR creates hierarchical trees by recursively clustering and summarizing documents using Gaussian Mixture Models, storing parent-child relationships in node metadata.
  • Implementation requires six steps: document loading, tree building via build_raptor_tree, vectorstore assembly, contextual-compression retriever setup, hierarchical traversal, and answer generation.
  • Key code locations: Tree logic resides in all_rag_techniques_runnable_scripts/raptor.py, while the interactive demonstration is available in all_rag_techniques/raptor.ipynb.
  • The RAPTORMethod class provides a high-level API that encapsulates the entire pipeline, automatically handling embeddings, clustering, and retrieval.
  • Contextual compression via LLMChainExtractor filters irrelevant content before final answer generation, optimizing token usage and accuracy.

Frequently Asked Questions

What makes RAPTOR different from standard RAG?

Standard RAG retrieves flat chunks based on vector similarity, potentially missing broader thematic connections. RAPTOR organizes information into a multi-level summary tree, enabling retrieval to navigate from high-level concepts down to specific details. According to the NirDiamant/RAG_Techniques source code, this hierarchical approach uses Gaussian Mixture Model clustering to create soft thematic groupings that pure similarity search cannot capture.

How does RAPTOR handle large document collections?

RAPTOR manages scale through recursive abstraction. Each tree level reduces the document set to a smaller set of summaries, allowing the system to index massive corpora without linearly increasing retrieval complexity. The max_levels parameter controls tree depth, and the build_raptor_tree function in raptor.py processes only the current level's embeddings at each iteration, keeping memory usage bounded.

Can I use different embedding models or LLMs with RAPTOR?

Yes. The implementation uses LangChain abstractions for all model interactions. You can swap the embedding model by modifying the embed_texts function calls, and change the LLM for summarization or answer generation by passing different model instances to RAPTORMethod or the individual functions. The code supports any LangChain-compatible embeddings or chat models.

What are the computational costs of building the RAPTOR tree?

Tree construction involves O(n) embedding calls for the initial level, followed by clustering and summarization for each subsequent level. Gaussian Mixture Model clustering in perform_clustering adds computational overhead compared to k-means, but provides superior soft clustering. The trade-off is front-loaded: once built, retrieval is fast due to the FAISS index, though storage costs increase because the vectorstore maintains embeddings for every tree level simultaneously.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →