# How to Integrate PageIndex into Existing RAG Pipelines: A Complete Technical Guide

> Integrate PageIndex into RAG pipelines by replacing chunking with hierarchical retrieval. Use page_index to build document trees and retrieve precise contextual sections for enhanced RAG performance.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: how-to-guide
- Published: 2026-02-16

---

**You can integrate PageIndex into existing RAG pipelines by replacing traditional vector-store chunking with hierarchical tree-based retrieval, using the `page_index` function to generate structured document trees and implementing reasoning-based node selection to retrieve precise contextual sections.**

PageIndex is an open-source library from VectifyAI that transforms long PDFs and Markdown files into hierarchical tree indices, serving as a structural retrieval backbone for Retrieval-Augmented Generation (RAG) systems. When you integrate PageIndex into existing RAG pipelines, you replace fixed-size semantic chunking with document-native sections enriched with page ranges, enabling LLM-driven reasoning over chapter hierarchies before text retrieval.

## Understanding the PageIndex Architecture

PageIndex processes documents through a seven-stage pipeline defined in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) and [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py). Understanding this flow helps you optimize integration points.

**Stage 1: PDF Tokenization**

The system reads documents page-by-page using `get_page_tokens` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 13-24) to calculate token counts and establish physical page boundaries.

**Stage 2: TOC Detection**

The `toc_detector_single_page` and `find_toc_pages` functions (lines 4-63 in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)) analyze the first `toc_check_page_num` pages to identify existing table-of-contents structures using LLM prompts.

**Stage 3: TOC Extraction and Transformation**

When a TOC exists, `toc_extractor` retrieves raw text and `toc_transformer` (lines 19-36) converts it into structured JSON containing `structure`, `title`, and `physical_index` fields.

**Stage 4: Physical Page Mapping**

The `toc_index_extractor` function (lines 40-66) verifies existing page numbers or infers section start pages through LLM reasoning when no TOC is present via `process_no_toc`.

**Stage 5: Consistency Verification**

`verify_toc` and `fix_incorrect_toc` (lines 90-140) re-check random sections and automatically correct inaccurate indices to ensure retrieval precision.

**Stage 6: Post-Processing**

The `post_processing` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 60-80) enriches nodes with `start_index` and `end_index` boundaries, generates optional summaries, and assigns unique `node_id` values.

**Stage 7: Final Output**

`page_index_main` in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) (lines 55-68) returns a JSON object containing `doc_name`, `doc_description`, and the hierarchical `structure` ready for RAG integration.

## Integration Patterns for RAG Pipelines

Traditional RAG architectures follow this pattern:

```

PDF → Chunk → Embedding → Vector DB → Retriever → LLM

```

When you integrate PageIndex into existing RAG pipelines, you replace the chunking and vector storage layers with structural reasoning:

```

PDF → PageIndex (tree) → Node selector (LLM reasoning) → Node text → LLM

```

**Structural Retrieval Benefits**

- **Semantic Integrity**: Unlike fixed-size chunks that split logical sections, PageIndex preserves document hierarchy (chapters, sections, subsections) with exact page ranges.
- **Explainable Paths**: Node titles and optional summaries create human-readable retrieval traces, showing exactly which document sections informed the generation.
- **Hierarchical Reasoning**: The LLM can navigate parent-child relationships ("Look in Chapter 3, specifically Section 2.1") before consuming text, reducing context window usage.

**Hybrid Architectures**

You can combine PageIndex with vector stores by using the tree for coarse retrieval (selecting relevant chapters) and embeddings for fine-grained search within specific nodes, optimizing both precision and recall.

## Step-by-Step Implementation

### Installing and Configuring PageIndex

Clone the repository and install dependencies:

```bash
git clone https://github.com/VectifyAI/PageIndex.git
cd PageIndex
pip install -r requirements.txt

```

Configure your LLM provider by creating a `.env` file:

```bash
echo "CHATGPT_API_KEY=sk-your-api-key" > .env

```

PageIndex supports model-agnostic configuration through the `ConfigLoader` class in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py), defaulting to `gpt-4o-2024-11-20`.

### Generating the Hierarchical Tree

Use the `page_index` function from [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) to process PDFs:

```python
from pageindex import page_index

tree = page_index(
    doc="reports/2023-annual-report.pdf",
    model="gpt-4o-2024-11-20",
    toc_check_page_num=30,
    max_page_num_each_node=12,
    max_token_num_each_node=25000,
    if_add_node_id="yes",
    if_add_node_summary="yes",
    if_add_doc_description="yes",
    if_add_node_text="yes"
)

```

The returned dictionary contains `doc_name`, `doc_description`, and `structure` (the hierarchical tree). Each node includes `title`, `node_id`, `start_index`, `end_index`, and optional `summary` and `text` fields.

Flatten the tree for easier processing using `structure_to_list` from [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py):

```python
from pageindex.utils import structure_to_list

nodes = structure_to_list(tree["structure"])

```

### Implementing Reasoning-Based Retrieval

Replace vector similarity search with LLM-driven node selection:

```python
import openai
import os

openai.api_key = os.getenv("CHATGPT_API_KEY")

def select_nodes(query, nodes):
    # Build prompt with titles and summaries

    titles = "\n".join(
        f"{i}. {n['title']} – {n.get('summary', '')}" 
        for i, n in enumerate(nodes)
    )
    
    prompt = f"""You are given a list of document sections with short summaries:
{titles}

Which section(s) best answer the user question?
Question: {query}
Respond with the index numbers (comma-separated) only."""

    response = openai.ChatCompletion.create(
        model="gpt-4o-2024-11-20",
        messages=[{"role": "user", "content": prompt}]
    )
    
    answer = response.choices[0].message.content.strip()
    return [int(i) for i in answer.split(",") if i.strip().isdigit()]

# Usage

chosen_indices = select_nodes("How did the Fed's policy change in 2023?", nodes)
context = "\n".join(nodes[i]["text"] for i in chosen_indices)

```

### Hybrid Retrieval with Vector Embeddings

Combine PageIndex with embeddings for fine-grained search within nodes:

```python
from sentence_transformers import SentenceTransformer
import numpy as np

embedder = SentenceTransformer("all-MiniLM-L6-v2")

def hybrid_retrieve(question, nodes):
    # Step 1: Select candidate nodes using PageIndex structure

    candidate_ids = select_nodes(question, nodes)
    
    # Step 2: Embed question

    q_vec = embedder.encode(question, normalize_embeddings=True)
    
    # Step 3: Find best match within candidates using embeddings

    best_node, best_score = None, -np.inf
    for nid in candidate_ids:
        node_text = nodes[nid]["text"]
        doc_vec = embedder.encode(node_text, normalize_embeddings=True)
        score = np.dot(q_vec, doc_vec)
        if score > best_score:
            best_score, best_node = score, nid
            
    return nodes[best_node]["text"]

# Usage

context = hybrid_retrieve("What were the Q3 revenue figures?", nodes)

```

### Indexing Markdown Documentation

For Markdown files, use the `md_to_tree` function from [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py):

```python
import asyncio
from pageindex.page_index_md import md_to_tree

async def index_markdown(md_path):
    result = await md_to_tree(
        md_path=md_path,
        if_thinning=True,
        min_token_threshold=5000,
        if_add_node_summary="yes",
        summary_token_threshold=200,
        model="gpt-4o-2024-11-20",
        if_add_doc_description="yes",
        if_add_node_text="yes",
        if_add_node_id="yes"
    )
    return result

# Run

tree_md = asyncio.run(index_markdown("docs/technical-spec.md"))

```

## Summary

- **PageIndex generates hierarchical tree structures** from PDFs and Markdown using `page_index` and `md_to_tree`, preserving document sections with exact page ranges instead of arbitrary chunks.

- **Integration replaces vector-store chunking** with reasoning-based retrieval, where LLMs navigate node titles and summaries to select relevant sections before consuming text.

- **Core functions** include `page_index_main` in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) for orchestration, `structure_to_list` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) for flattening trees, and `select_nodes` patterns for LLM-driven retrieval.

- **Hybrid approaches** combine PageIndex for coarse retrieval with embeddings for fine-grained search within selected nodes, optimizing both precision and computational efficiency.

- **Configuration** supports model-agnostic LLM selection through `ConfigLoader`, with customizable prompts for TOC detection and transformation in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py).

## Frequently Asked Questions

### How does PageIndex handle PDFs without a table of contents?

When no TOC is detected, PageIndex uses the `process_no_toc` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) to infer section boundaries by prompting the LLM to analyze document structure and identify logical chapter breaks based on content flow and formatting cues.

### Can I use PageIndex with local LLMs instead of OpenAI?

Yes, PageIndex is model-agnostic. The `ConfigLoader` class in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) manages LLM configuration, and you can modify the `ChatGPT_API` wrapper functions to route requests to local endpoints by changing the base URL and model parameters while maintaining the same JSON response format.

### What is the difference between `page_index` and `page_index_main`?

`page_index` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) is the core processing function that accepts configuration parameters directly, while `page_index_main` in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) serves as the CLI entry point that builds a configuration object from command-line arguments before invoking the core function.

### How do I merge PageIndex trees from multiple documents?

Since PageIndex outputs standard JSON objects containing `doc_name`, `doc_description`, and `structure` arrays, you can merge trees by combining the `structure` lists under a new parent node or maintaining separate document roots in a dictionary keyed by `doc_name`, then using `structure_to_list` from [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) to flatten specific branches for retrieval.