RAG Dependencies in AI Agent Book: Complete Package Guide for RAPTOR and GraphRAG
The AI Agent Book repository requires distinct package sets for RAPTOR (tree-structured retrieval) and GraphRAG (knowledge-graph retrieval), with both sharing common vector storage, LLM, and document processing libraries defined in chapter3/structured-index/requirements.txt.
The bojieli/ai-agent-book repository implements production-ready Retrieval-Augmented Generation (RAG) pipelines using two advanced indexing strategies. Understanding the dependencies for RAG is essential for replicating the book's examples and deploying the provided FastAPI service. All required packages are centralized in chapter3/structured-index/requirements.txt, which separates concerns between the two indexing approaches and their shared infrastructure.
RAPTOR-Specific Dependencies
According to the source code, lines 2-8 of chapter3/structured-index/requirements.txt declare the core packages for RAPTOR (Recursive Abstraction Processing for Tree-Organized Retrieval). This approach builds a hierarchical tree index that recursively summarizes documents for fast retrieval.
Install these packages to enable RAPTOR functionality:
pip install raptor-rag==0.3.0 openai>=1.12.0 scikit-learn>=1.3.0 numpy>=1.24.0 tiktoken>=0.5.0 umap-learn>=0.5.4
The raptor-rag package provides the tree construction algorithms, while umap-learn and scikit-learn handle dimensionality reduction and clustering operations during index creation.
GraphRAG-Specific Dependencies
Lines 10-16 of the requirements file define the stack for GraphRAG, a knowledge-graph style index that extracts entities and relations to support multi-hop traversals across document collections.
Install the GraphRAG dependency set:
pip install graphrag>=0.3.0 azure-search-documents>=11.4.0 azure-storage-blob>=12.19.0 networkx>=3.0 pyarrow>=15.0.0 pyyaml>=6.0 rich>=13.0.0
Here, graphrag provides the core knowledge-graph indexing engine, networkx handles graph traversals, and azure-search-documents enables optional Azure Cognitive Search integration for hybrid retrieval scenarios.
Shared Infrastructure Dependencies
Both RAG variants rely on a common stack for vector operations, LLM access, and utility functions. Lines 19-50 of chapter3/structured-index/requirements.txt contain these shared dependencies organized by functional category.
Vector Storage and Embeddings
Lines 19-21 specify the embedding and similarity search stack:
faiss-cpu– Facebook AI Similarity Search for efficient vector indexingsentence-transformers– Local embedding model inferencetransformers– Hugging Face model pipeline support
HTTP API and Environment
Lines 27-30 declare the serving stack:
fastapianduvicorn– ASGI server for the HTTP APIpydantic– Data validation for API request/response modelshttpxandrequests– HTTP client libraries for external API callspython-dotenv– Environment variable management for API keys
Document Processing
Lines 33-38 and 40-42 list the document ingestion utilities used by the DocumentProcessor class:
pypdfandpdfplumber– PDF text extractionbeautifulsoup4andlxml– HTML parsingmarkdown– Markdown document handlingaiofiles– Asynchronous file operationstabulate– Table formatting for Intel-manual parsing helpers
Logging and Monitoring
Lines 44-45 provide observability tools:
loguru– Structured logging throughout the pipelinetqdm– Progress bars for long-running indexing operations
Testing Utilities
Lines 48-50 enable the test suite:
pytestandpytest-asyncio– Test framework for async RAG operations
Key Implementation Files
The dependencies support specific modules within chapter3/structured-index/:
raptor_indexer.py– ImplementsRaptorIndexerusing the RAPTOR package stackgraphrag_indexer.py– ImplementsGraphRAGIndexerwith NetworkX and Azure integrationsdocument_processor.py– Coordinates PDF/HTML parsing using BeautifulSoup and PyPDFapi_service.py– FastAPI wrapper relying on Pydantic and Uvicornmain.py– CLI entry point importingbuild_indexesandquery_indexesfunctions
Usage Examples
Install all dependencies at once from the repository root:
pip install -r chapter3/structured-index/requirements.txt
Build both indexes for a document:
# demo.py (run from the repository root)
import asyncio
from pathlib import Path
from chapter3.structured_index.main import build_indexes
async def run_demo():
await build_indexes(
file_path=Path("sample_document.pdf"),
index_type="both", # build RAPTOR and GraphRAG
output="index_stats.json" # optional JSON with statistics
)
if __name__ == "__main__":
asyncio.run(run_demo())
Query with multi-hop GraphRAG traversal:
import asyncio
from chapter3.structured_index.main import query_indexes
async def query_demo():
results = await query_indexes(
query="What are the safety requirements for a nuclear reactor?",
index_type="both",
top_k=5,
multi_hop=2 # follow up to 2 relation hops
)
print(results)
if __name__ == "__main__":
asyncio.run(query_demo())
Start the HTTP API service:
python -m chapter3.structured_index.main serve
The service exposes /query and /build endpoints on http://127.0.0.1:8000 and requires an OPENAI_API_KEY environment variable for LLM access.
Summary
- RAPTOR dependencies (lines 2-8):
raptor-rag==0.3.0,openai>=1.12.0,scikit-learn>=1.3.0,numpy>=1.24.0,tiktoken>=0.5.0,umap-learn>=0.5.4 - GraphRAG dependencies (lines 10-16):
graphrag>=0.3.0,azure-search-documents>=11.4.0,azure-storage-blob>=12.19.0,networkx>=3.0,pyarrow>=15.0.0,pyyaml>=6.0,rich>=13.0.0 - Common stack (lines 19-50):
faiss-cpu,sentence-transformers,transformers,fastapi,uvicorn,pypdf,beautifulsoup4,loguru,pytest - All packages are pinned or constrained in
chapter3/structured-index/requirements.txtto ensure compatibility between the tree-based and graph-based retrieval systems.
Frequently Asked Questions
What is the minimum set of dependencies to run only the RAPTOR indexer?
Install lines 2-8 and 19-25 from chapter3/structured-index/requirements.txt. You need raptor-rag==0.3.0, openai>=1.12.0, scikit-learn>=1.3.0, numpy>=1.24.0, tiktoken>=0.5.0, umap-learn>=0.5.4, plus the shared vector libraries faiss-cpu, sentence-transformers, and transformers. This excludes GraphRAG-specific packages like azure-search-documents and networkx.
Does the GraphRAG implementation require Microsoft Azure services?
No. While azure-search-documents>=11.4.0 and azure-storage-blob>=12.19.0 are listed as dependencies for hybrid search capabilities, the core GraphRAGIndexer in chapter3/structured-index/graphrag_indexer.py operates locally using networkx>=3.0 for graph storage and traversal. Azure integration is optional for cloud deployment scenarios.
Which dependencies handle PDF and HTML document ingestion?
The DocumentProcessor class relies on pypdf and pdfplumber for PDF extraction, beautifulsoup4 with lxml for HTML parsing, and markdown for MD files. These are declared in lines 33-38 of the requirements file. aiofiles enables non-blocking file I/O during batch processing, while tabulate (lines 40-42) assists with Intel-manual parsing helpers.
How do I start the FastAPI service for the RAG pipeline?
Install the API-specific dependencies from lines 27-30: fastapi, uvicorn, pydantic, and httpx. Then execute python -m chapter3.structured_index.main serve. The server requires an OPENAI_API_KEY environment variable for LLM access during retrieval, though local embedding models via sentence-transformers can operate without API calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →