# RAG Dependencies in AI Agent Book: Complete Package Guide for RAPTOR and GraphRAG

> Explore RAG dependencies in the AI Agent Book. Understand RAPTOR and GraphRAG package sets, common vector storage, LLM, and document processing libraries.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: complete-package-guide
- Published: 2026-08-22

---

**The AI Agent Book repository requires distinct package sets for RAPTOR (tree-structured retrieval) and GraphRAG (knowledge-graph retrieval), with both sharing common vector storage, LLM, and document processing libraries defined in [`chapter3/structured-index/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/requirements.txt).**

The `bojieli/ai-agent-book` repository implements production-ready Retrieval-Augmented Generation (RAG) pipelines using two advanced indexing strategies. Understanding the dependencies for RAG is essential for replicating the book's examples and deploying the provided FastAPI service. All required packages are centralized in [`chapter3/structured-index/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/requirements.txt), which separates concerns between the two indexing approaches and their shared infrastructure.

## RAPTOR-Specific Dependencies

According to the source code, lines 2-8 of [`chapter3/structured-index/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/requirements.txt) declare the core packages for **RAPTOR** (Recursive Abstraction Processing for Tree-Organized Retrieval). This approach builds a hierarchical tree index that recursively summarizes documents for fast retrieval.

Install these packages to enable RAPTOR functionality:

```bash
pip install raptor-rag==0.3.0 openai>=1.12.0 scikit-learn>=1.3.0 numpy>=1.24.0 tiktoken>=0.5.0 umap-learn>=0.5.4

```

The `raptor-rag` package provides the tree construction algorithms, while `umap-learn` and `scikit-learn` handle dimensionality reduction and clustering operations during index creation.

## GraphRAG-Specific Dependencies

Lines 10-16 of the requirements file define the stack for **GraphRAG**, a knowledge-graph style index that extracts entities and relations to support multi-hop traversals across document collections.

Install the GraphRAG dependency set:

```bash
pip install graphrag>=0.3.0 azure-search-documents>=11.4.0 azure-storage-blob>=12.19.0 networkx>=3.0 pyarrow>=15.0.0 pyyaml>=6.0 rich>=13.0.0

```

Here, `graphrag` provides the core knowledge-graph indexing engine, `networkx` handles graph traversals, and `azure-search-documents` enables optional Azure Cognitive Search integration for hybrid retrieval scenarios.

## Shared Infrastructure Dependencies

Both RAG variants rely on a common stack for vector operations, LLM access, and utility functions. Lines 19-50 of [`chapter3/structured-index/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/requirements.txt) contain these shared dependencies organized by functional category.

### Vector Storage and Embeddings

Lines 19-21 specify the embedding and similarity search stack:

- **`faiss-cpu`** – Facebook AI Similarity Search for efficient vector indexing
- **`sentence-transformers`** – Local embedding model inference
- **`transformers`** – Hugging Face model pipeline support

### HTTP API and Environment

Lines 27-30 declare the serving stack:

- **`fastapi`** and **`uvicorn`** – ASGI server for the HTTP API
- **`pydantic`** – Data validation for API request/response models
- **`httpx`** and **`requests`** – HTTP client libraries for external API calls
- **`python-dotenv`** – Environment variable management for API keys

### Document Processing

Lines 33-38 and 40-42 list the document ingestion utilities used by the `DocumentProcessor` class:

- **`pypdf`** and **`pdfplumber`** – PDF text extraction
- **`beautifulsoup4`** and **`lxml`** – HTML parsing
- **`markdown`** – Markdown document handling
- **`aiofiles`** – Asynchronous file operations
- **`tabulate`** – Table formatting for Intel-manual parsing helpers

### Logging and Monitoring

Lines 44-45 provide observability tools:

- **`loguru`** – Structured logging throughout the pipeline
- **`tqdm`** – Progress bars for long-running indexing operations

### Testing Utilities

Lines 48-50 enable the test suite:

- **`pytest`** and **`pytest-asyncio`** – Test framework for async RAG operations

## Key Implementation Files

The dependencies support specific modules within `chapter3/structured-index/`:

- **[`raptor_indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/raptor_indexer.py)** – Implements `RaptorIndexer` using the RAPTOR package stack
- **[`graphrag_indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/graphrag_indexer.py)** – Implements `GraphRAGIndexer` with NetworkX and Azure integrations
- **[`document_processor.py`](https://github.com/bojieli/ai-agent-book/blob/main/document_processor.py)** – Coordinates PDF/HTML parsing using BeautifulSoup and PyPDF
- **[`api_service.py`](https://github.com/bojieli/ai-agent-book/blob/main/api_service.py)** – FastAPI wrapper relying on Pydantic and Uvicorn
- **[`main.py`](https://github.com/bojieli/ai-agent-book/blob/main/main.py)** – CLI entry point importing `build_indexes` and `query_indexes` functions

## Usage Examples

Install all dependencies at once from the repository root:

```bash
pip install -r chapter3/structured-index/requirements.txt

```

Build both indexes for a document:

```python

# demo.py (run from the repository root)

import asyncio
from pathlib import Path
from chapter3.structured_index.main import build_indexes

async def run_demo():
    await build_indexes(
        file_path=Path("sample_document.pdf"),
        index_type="both",          # build RAPTOR and GraphRAG

        output="index_stats.json"  # optional JSON with statistics

    )

if __name__ == "__main__":
    asyncio.run(run_demo())

```

Query with multi-hop GraphRAG traversal:

```python
import asyncio
from chapter3.structured_index.main import query_indexes

async def query_demo():
    results = await query_indexes(
        query="What are the safety requirements for a nuclear reactor?",
        index_type="both",
        top_k=5,
        multi_hop=2               # follow up to 2 relation hops

    )
    print(results)

if __name__ == "__main__":
    asyncio.run(query_demo())

```

Start the HTTP API service:

```bash
python -m chapter3.structured_index.main serve

```

The service exposes `/query` and `/build` endpoints on `http://127.0.0.1:8000` and requires an `OPENAI_API_KEY` environment variable for LLM access.

## Summary

- **RAPTOR dependencies** (lines 2-8): `raptor-rag==0.3.0`, `openai>=1.12.0`, `scikit-learn>=1.3.0`, `numpy>=1.24.0`, `tiktoken>=0.5.0`, `umap-learn>=0.5.4`
- **GraphRAG dependencies** (lines 10-16): `graphrag>=0.3.0`, `azure-search-documents>=11.4.0`, `azure-storage-blob>=12.19.0`, `networkx>=3.0`, `pyarrow>=15.0.0`, `pyyaml>=6.0`, `rich>=13.0.0`
- **Common stack** (lines 19-50): `faiss-cpu`, `sentence-transformers`, `transformers`, `fastapi`, `uvicorn`, `pypdf`, `beautifulsoup4`, `loguru`, `pytest`
- All packages are pinned or constrained in [`chapter3/structured-index/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/requirements.txt) to ensure compatibility between the tree-based and graph-based retrieval systems.

## Frequently Asked Questions

### What is the minimum set of dependencies to run only the RAPTOR indexer?

Install lines 2-8 and 19-25 from [`chapter3/structured-index/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/requirements.txt). You need `raptor-rag==0.3.0`, `openai>=1.12.0`, `scikit-learn>=1.3.0`, `numpy>=1.24.0`, `tiktoken>=0.5.0`, `umap-learn>=0.5.4`, plus the shared vector libraries `faiss-cpu`, `sentence-transformers`, and `transformers`. This excludes GraphRAG-specific packages like `azure-search-documents` and `networkx`.

### Does the GraphRAG implementation require Microsoft Azure services?

No. While `azure-search-documents>=11.4.0` and `azure-storage-blob>=12.19.0` are listed as dependencies for hybrid search capabilities, the core `GraphRAGIndexer` in [`chapter3/structured-index/graphrag_indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py) operates locally using `networkx>=3.0` for graph storage and traversal. Azure integration is optional for cloud deployment scenarios.

### Which dependencies handle PDF and HTML document ingestion?

The `DocumentProcessor` class relies on `pypdf` and `pdfplumber` for PDF extraction, `beautifulsoup4` with `lxml` for HTML parsing, and `markdown` for MD files. These are declared in lines 33-38 of the requirements file. `aiofiles` enables non-blocking file I/O during batch processing, while `tabulate` (lines 40-42) assists with Intel-manual parsing helpers.

### How do I start the FastAPI service for the RAG pipeline?

Install the API-specific dependencies from lines 27-30: `fastapi`, `uvicorn`, `pydantic`, and `httpx`. Then execute `python -m chapter3.structured_index.main serve`. The server requires an `OPENAI_API_KEY` environment variable for LLM access during retrieval, though local embedding models via `sentence-transformers` can operate without API calls.