How to Integrate PageIndex into Existing RAG Pipelines: A Complete Technical Guide

You can integrate PageIndex into existing RAG pipelines by replacing traditional vector-store chunking with hierarchical tree-based retrieval, using the page_index function to generate structured document trees and implementing reasoning-based node selection to retrieve precise contextual sections.

PageIndex is an open-source library from VectifyAI that transforms long PDFs and Markdown files into hierarchical tree indices, serving as a structural retrieval backbone for Retrieval-Augmented Generation (RAG) systems. When you integrate PageIndex into existing RAG pipelines, you replace fixed-size semantic chunking with document-native sections enriched with page ranges, enabling LLM-driven reasoning over chapter hierarchies before text retrieval.

Understanding the PageIndex Architecture

PageIndex processes documents through a seven-stage pipeline defined in pageindex/page_index.py and pageindex/utils.py. Understanding this flow helps you optimize integration points.

Stage 1: PDF Tokenization

The system reads documents page-by-page using get_page_tokens in pageindex/utils.py (lines 13-24) to calculate token counts and establish physical page boundaries.

Stage 2: TOC Detection

The toc_detector_single_page and find_toc_pages functions (lines 4-63 in pageindex/page_index.py) analyze the first toc_check_page_num pages to identify existing table-of-contents structures using LLM prompts.

Stage 3: TOC Extraction and Transformation

When a TOC exists, toc_extractor retrieves raw text and toc_transformer (lines 19-36) converts it into structured JSON containing structure, title, and physical_index fields.

Stage 4: Physical Page Mapping

The toc_index_extractor function (lines 40-66) verifies existing page numbers or infers section start pages through LLM reasoning when no TOC is present via process_no_toc.

Stage 5: Consistency Verification

verify_toc and fix_incorrect_toc (lines 90-140) re-check random sections and automatically correct inaccurate indices to ensure retrieval precision.

Stage 6: Post-Processing

The post_processing function in pageindex/utils.py (lines 60-80) enriches nodes with start_index and end_index boundaries, generates optional summaries, and assigns unique node_id values.

Stage 7: Final Output

page_index_main in run_pageindex.py (lines 55-68) returns a JSON object containing doc_name, doc_description, and the hierarchical structure ready for RAG integration.

Integration Patterns for RAG Pipelines

Traditional RAG architectures follow this pattern:


PDF → Chunk → Embedding → Vector DB → Retriever → LLM

When you integrate PageIndex into existing RAG pipelines, you replace the chunking and vector storage layers with structural reasoning:


PDF → PageIndex (tree) → Node selector (LLM reasoning) → Node text → LLM

Structural Retrieval Benefits

  • Semantic Integrity: Unlike fixed-size chunks that split logical sections, PageIndex preserves document hierarchy (chapters, sections, subsections) with exact page ranges.
  • Explainable Paths: Node titles and optional summaries create human-readable retrieval traces, showing exactly which document sections informed the generation.
  • Hierarchical Reasoning: The LLM can navigate parent-child relationships ("Look in Chapter 3, specifically Section 2.1") before consuming text, reducing context window usage.

Hybrid Architectures

You can combine PageIndex with vector stores by using the tree for coarse retrieval (selecting relevant chapters) and embeddings for fine-grained search within specific nodes, optimizing both precision and recall.

Step-by-Step Implementation

Installing and Configuring PageIndex

Clone the repository and install dependencies:

git clone https://github.com/VectifyAI/PageIndex.git
cd PageIndex
pip install -r requirements.txt

Configure your LLM provider by creating a .env file:

echo "CHATGPT_API_KEY=sk-your-api-key" > .env

PageIndex supports model-agnostic configuration through the ConfigLoader class in pageindex/utils.py, defaulting to gpt-4o-2024-11-20.

Generating the Hierarchical Tree

Use the page_index function from pageindex/page_index.py to process PDFs:

from pageindex import page_index

tree = page_index(
    doc="reports/2023-annual-report.pdf",
    model="gpt-4o-2024-11-20",
    toc_check_page_num=30,
    max_page_num_each_node=12,
    max_token_num_each_node=25000,
    if_add_node_id="yes",
    if_add_node_summary="yes",
    if_add_doc_description="yes",
    if_add_node_text="yes"
)

The returned dictionary contains doc_name, doc_description, and structure (the hierarchical tree). Each node includes title, node_id, start_index, end_index, and optional summary and text fields.

Flatten the tree for easier processing using structure_to_list from pageindex/utils.py:

from pageindex.utils import structure_to_list

nodes = structure_to_list(tree["structure"])

Implementing Reasoning-Based Retrieval

Replace vector similarity search with LLM-driven node selection:

import openai
import os

openai.api_key = os.getenv("CHATGPT_API_KEY")

def select_nodes(query, nodes):
    # Build prompt with titles and summaries

    titles = "\n".join(
        f"{i}. {n['title']} – {n.get('summary', '')}" 
        for i, n in enumerate(nodes)
    )
    
    prompt = f"""You are given a list of document sections with short summaries:
{titles}

Which section(s) best answer the user question?
Question: {query}
Respond with the index numbers (comma-separated) only."""

    response = openai.ChatCompletion.create(
        model="gpt-4o-2024-11-20",
        messages=[{"role": "user", "content": prompt}]
    )
    
    answer = response.choices[0].message.content.strip()
    return [int(i) for i in answer.split(",") if i.strip().isdigit()]

# Usage

chosen_indices = select_nodes("How did the Fed's policy change in 2023?", nodes)
context = "\n".join(nodes[i]["text"] for i in chosen_indices)

Hybrid Retrieval with Vector Embeddings

Combine PageIndex with embeddings for fine-grained search within nodes:

from sentence_transformers import SentenceTransformer
import numpy as np

embedder = SentenceTransformer("all-MiniLM-L6-v2")

def hybrid_retrieve(question, nodes):
    # Step 1: Select candidate nodes using PageIndex structure

    candidate_ids = select_nodes(question, nodes)
    
    # Step 2: Embed question

    q_vec = embedder.encode(question, normalize_embeddings=True)
    
    # Step 3: Find best match within candidates using embeddings

    best_node, best_score = None, -np.inf
    for nid in candidate_ids:
        node_text = nodes[nid]["text"]
        doc_vec = embedder.encode(node_text, normalize_embeddings=True)
        score = np.dot(q_vec, doc_vec)
        if score > best_score:
            best_score, best_node = score, nid
            
    return nodes[best_node]["text"]

# Usage

context = hybrid_retrieve("What were the Q3 revenue figures?", nodes)

Indexing Markdown Documentation

For Markdown files, use the md_to_tree function from pageindex/page_index_md.py:

import asyncio
from pageindex.page_index_md import md_to_tree

async def index_markdown(md_path):
    result = await md_to_tree(
        md_path=md_path,
        if_thinning=True,
        min_token_threshold=5000,
        if_add_node_summary="yes",
        summary_token_threshold=200,
        model="gpt-4o-2024-11-20",
        if_add_doc_description="yes",
        if_add_node_text="yes",
        if_add_node_id="yes"
    )
    return result

# Run

tree_md = asyncio.run(index_markdown("docs/technical-spec.md"))

Summary

  • PageIndex generates hierarchical tree structures from PDFs and Markdown using page_index and md_to_tree, preserving document sections with exact page ranges instead of arbitrary chunks.

  • Integration replaces vector-store chunking with reasoning-based retrieval, where LLMs navigate node titles and summaries to select relevant sections before consuming text.

  • Core functions include page_index_main in run_pageindex.py for orchestration, structure_to_list in pageindex/utils.py for flattening trees, and select_nodes patterns for LLM-driven retrieval.

  • Hybrid approaches combine PageIndex for coarse retrieval with embeddings for fine-grained search within selected nodes, optimizing both precision and computational efficiency.

  • Configuration supports model-agnostic LLM selection through ConfigLoader, with customizable prompts for TOC detection and transformation in pageindex/page_index.py.

Frequently Asked Questions

How does PageIndex handle PDFs without a table of contents?

When no TOC is detected, PageIndex uses the process_no_toc function in pageindex/page_index.py to infer section boundaries by prompting the LLM to analyze document structure and identify logical chapter breaks based on content flow and formatting cues.

Can I use PageIndex with local LLMs instead of OpenAI?

Yes, PageIndex is model-agnostic. The ConfigLoader class in pageindex/utils.py manages LLM configuration, and you can modify the ChatGPT_API wrapper functions to route requests to local endpoints by changing the base URL and model parameters while maintaining the same JSON response format.

What is the difference between page_index and page_index_main?

page_index in pageindex/page_index.py is the core processing function that accepts configuration parameters directly, while page_index_main in run_pageindex.py serves as the CLI entry point that builds a configuration object from command-line arguments before invoking the core function.

How do I merge PageIndex trees from multiple documents?

Since PageIndex outputs standard JSON objects containing doc_name, doc_description, and structure arrays, you can merge trees by combining the structure lists under a new parent node or maintaining separate document roots in a dictionary keyed by doc_name, then using structure_to_list from pageindex/utils.py to flatten specific branches for retrieval.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →