RAG Dependencies in AI Agent Book: Complete Package Guide for RAPTOR and GraphRAG

The AI Agent Book repository requires distinct package sets for RAPTOR (tree-structured retrieval) and GraphRAG (knowledge-graph retrieval), with both sharing common vector storage, LLM, and document processing libraries defined in chapter3/structured-index/requirements.txt.

The bojieli/ai-agent-book repository implements production-ready Retrieval-Augmented Generation (RAG) pipelines using two advanced indexing strategies. Understanding the dependencies for RAG is essential for replicating the book's examples and deploying the provided FastAPI service. All required packages are centralized in chapter3/structured-index/requirements.txt, which separates concerns between the two indexing approaches and their shared infrastructure.

RAPTOR-Specific Dependencies

According to the source code, lines 2-8 of chapter3/structured-index/requirements.txt declare the core packages for RAPTOR (Recursive Abstraction Processing for Tree-Organized Retrieval). This approach builds a hierarchical tree index that recursively summarizes documents for fast retrieval.

Install these packages to enable RAPTOR functionality:

pip install raptor-rag==0.3.0 openai>=1.12.0 scikit-learn>=1.3.0 numpy>=1.24.0 tiktoken>=0.5.0 umap-learn>=0.5.4

The raptor-rag package provides the tree construction algorithms, while umap-learn and scikit-learn handle dimensionality reduction and clustering operations during index creation.

GraphRAG-Specific Dependencies

Lines 10-16 of the requirements file define the stack for GraphRAG, a knowledge-graph style index that extracts entities and relations to support multi-hop traversals across document collections.

Install the GraphRAG dependency set:

pip install graphrag>=0.3.0 azure-search-documents>=11.4.0 azure-storage-blob>=12.19.0 networkx>=3.0 pyarrow>=15.0.0 pyyaml>=6.0 rich>=13.0.0

Here, graphrag provides the core knowledge-graph indexing engine, networkx handles graph traversals, and azure-search-documents enables optional Azure Cognitive Search integration for hybrid retrieval scenarios.

Shared Infrastructure Dependencies

Both RAG variants rely on a common stack for vector operations, LLM access, and utility functions. Lines 19-50 of chapter3/structured-index/requirements.txt contain these shared dependencies organized by functional category.

Vector Storage and Embeddings

Lines 19-21 specify the embedding and similarity search stack:

  • faiss-cpu – Facebook AI Similarity Search for efficient vector indexing
  • sentence-transformers – Local embedding model inference
  • transformers – Hugging Face model pipeline support

HTTP API and Environment

Lines 27-30 declare the serving stack:

  • fastapi and uvicorn – ASGI server for the HTTP API
  • pydantic – Data validation for API request/response models
  • httpx and requests – HTTP client libraries for external API calls
  • python-dotenv – Environment variable management for API keys

Document Processing

Lines 33-38 and 40-42 list the document ingestion utilities used by the DocumentProcessor class:

  • pypdf and pdfplumber – PDF text extraction
  • beautifulsoup4 and lxml – HTML parsing
  • markdown – Markdown document handling
  • aiofiles – Asynchronous file operations
  • tabulate – Table formatting for Intel-manual parsing helpers

Logging and Monitoring

Lines 44-45 provide observability tools:

  • loguru – Structured logging throughout the pipeline
  • tqdm – Progress bars for long-running indexing operations

Testing Utilities

Lines 48-50 enable the test suite:

  • pytest and pytest-asyncio – Test framework for async RAG operations

Key Implementation Files

The dependencies support specific modules within chapter3/structured-index/:

  • raptor_indexer.py – Implements RaptorIndexer using the RAPTOR package stack
  • graphrag_indexer.py – Implements GraphRAGIndexer with NetworkX and Azure integrations
  • document_processor.py – Coordinates PDF/HTML parsing using BeautifulSoup and PyPDF
  • api_service.py – FastAPI wrapper relying on Pydantic and Uvicorn
  • main.py – CLI entry point importing build_indexes and query_indexes functions

Usage Examples

Install all dependencies at once from the repository root:

pip install -r chapter3/structured-index/requirements.txt

Build both indexes for a document:


# demo.py (run from the repository root)

import asyncio
from pathlib import Path
from chapter3.structured_index.main import build_indexes

async def run_demo():
    await build_indexes(
        file_path=Path("sample_document.pdf"),
        index_type="both",          # build RAPTOR and GraphRAG

        output="index_stats.json"  # optional JSON with statistics

    )

if __name__ == "__main__":
    asyncio.run(run_demo())

Query with multi-hop GraphRAG traversal:

import asyncio
from chapter3.structured_index.main import query_indexes

async def query_demo():
    results = await query_indexes(
        query="What are the safety requirements for a nuclear reactor?",
        index_type="both",
        top_k=5,
        multi_hop=2               # follow up to 2 relation hops

    )
    print(results)

if __name__ == "__main__":
    asyncio.run(query_demo())

Start the HTTP API service:

python -m chapter3.structured_index.main serve

The service exposes /query and /build endpoints on http://127.0.0.1:8000 and requires an OPENAI_API_KEY environment variable for LLM access.

Summary

  • RAPTOR dependencies (lines 2-8): raptor-rag==0.3.0, openai>=1.12.0, scikit-learn>=1.3.0, numpy>=1.24.0, tiktoken>=0.5.0, umap-learn>=0.5.4
  • GraphRAG dependencies (lines 10-16): graphrag>=0.3.0, azure-search-documents>=11.4.0, azure-storage-blob>=12.19.0, networkx>=3.0, pyarrow>=15.0.0, pyyaml>=6.0, rich>=13.0.0
  • Common stack (lines 19-50): faiss-cpu, sentence-transformers, transformers, fastapi, uvicorn, pypdf, beautifulsoup4, loguru, pytest
  • All packages are pinned or constrained in chapter3/structured-index/requirements.txt to ensure compatibility between the tree-based and graph-based retrieval systems.

Frequently Asked Questions

What is the minimum set of dependencies to run only the RAPTOR indexer?

Install lines 2-8 and 19-25 from chapter3/structured-index/requirements.txt. You need raptor-rag==0.3.0, openai>=1.12.0, scikit-learn>=1.3.0, numpy>=1.24.0, tiktoken>=0.5.0, umap-learn>=0.5.4, plus the shared vector libraries faiss-cpu, sentence-transformers, and transformers. This excludes GraphRAG-specific packages like azure-search-documents and networkx.

Does the GraphRAG implementation require Microsoft Azure services?

No. While azure-search-documents>=11.4.0 and azure-storage-blob>=12.19.0 are listed as dependencies for hybrid search capabilities, the core GraphRAGIndexer in chapter3/structured-index/graphrag_indexer.py operates locally using networkx>=3.0 for graph storage and traversal. Azure integration is optional for cloud deployment scenarios.

Which dependencies handle PDF and HTML document ingestion?

The DocumentProcessor class relies on pypdf and pdfplumber for PDF extraction, beautifulsoup4 with lxml for HTML parsing, and markdown for MD files. These are declared in lines 33-38 of the requirements file. aiofiles enables non-blocking file I/O during batch processing, while tabulate (lines 40-42) assists with Intel-manual parsing helpers.

How do I start the FastAPI service for the RAG pipeline?

Install the API-specific dependencies from lines 27-30: fastapi, uvicorn, pydantic, and httpx. Then execute python -m chapter3.structured_index.main serve. The server requires an OPENAI_API_KEY environment variable for LLM access during retrieval, though local embedding models via sentence-transformers can operate without API calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →