How to Integrate PageIndex into Existing RAG Pipelines: A Complete Technical Guide
You can integrate PageIndex into existing RAG pipelines by replacing traditional vector-store chunking with hierarchical tree-based retrieval, using the page_index function to generate structured document trees and implementing reasoning-based node selection to retrieve precise contextual sections.
PageIndex is an open-source library from VectifyAI that transforms long PDFs and Markdown files into hierarchical tree indices, serving as a structural retrieval backbone for Retrieval-Augmented Generation (RAG) systems. When you integrate PageIndex into existing RAG pipelines, you replace fixed-size semantic chunking with document-native sections enriched with page ranges, enabling LLM-driven reasoning over chapter hierarchies before text retrieval.
Understanding the PageIndex Architecture
PageIndex processes documents through a seven-stage pipeline defined in pageindex/page_index.py and pageindex/utils.py. Understanding this flow helps you optimize integration points.
Stage 1: PDF Tokenization
The system reads documents page-by-page using get_page_tokens in pageindex/utils.py (lines 13-24) to calculate token counts and establish physical page boundaries.
Stage 2: TOC Detection
The toc_detector_single_page and find_toc_pages functions (lines 4-63 in pageindex/page_index.py) analyze the first toc_check_page_num pages to identify existing table-of-contents structures using LLM prompts.
Stage 3: TOC Extraction and Transformation
When a TOC exists, toc_extractor retrieves raw text and toc_transformer (lines 19-36) converts it into structured JSON containing structure, title, and physical_index fields.
Stage 4: Physical Page Mapping
The toc_index_extractor function (lines 40-66) verifies existing page numbers or infers section start pages through LLM reasoning when no TOC is present via process_no_toc.
Stage 5: Consistency Verification
verify_toc and fix_incorrect_toc (lines 90-140) re-check random sections and automatically correct inaccurate indices to ensure retrieval precision.
Stage 6: Post-Processing
The post_processing function in pageindex/utils.py (lines 60-80) enriches nodes with start_index and end_index boundaries, generates optional summaries, and assigns unique node_id values.
Stage 7: Final Output
page_index_main in run_pageindex.py (lines 55-68) returns a JSON object containing doc_name, doc_description, and the hierarchical structure ready for RAG integration.
Integration Patterns for RAG Pipelines
Traditional RAG architectures follow this pattern:
PDF → Chunk → Embedding → Vector DB → Retriever → LLM
When you integrate PageIndex into existing RAG pipelines, you replace the chunking and vector storage layers with structural reasoning:
PDF → PageIndex (tree) → Node selector (LLM reasoning) → Node text → LLM
Structural Retrieval Benefits
- Semantic Integrity: Unlike fixed-size chunks that split logical sections, PageIndex preserves document hierarchy (chapters, sections, subsections) with exact page ranges.
- Explainable Paths: Node titles and optional summaries create human-readable retrieval traces, showing exactly which document sections informed the generation.
- Hierarchical Reasoning: The LLM can navigate parent-child relationships ("Look in Chapter 3, specifically Section 2.1") before consuming text, reducing context window usage.
Hybrid Architectures
You can combine PageIndex with vector stores by using the tree for coarse retrieval (selecting relevant chapters) and embeddings for fine-grained search within specific nodes, optimizing both precision and recall.
Step-by-Step Implementation
Installing and Configuring PageIndex
Clone the repository and install dependencies:
git clone https://github.com/VectifyAI/PageIndex.git
cd PageIndex
pip install -r requirements.txt
Configure your LLM provider by creating a .env file:
echo "CHATGPT_API_KEY=sk-your-api-key" > .env
PageIndex supports model-agnostic configuration through the ConfigLoader class in pageindex/utils.py, defaulting to gpt-4o-2024-11-20.
Generating the Hierarchical Tree
Use the page_index function from pageindex/page_index.py to process PDFs:
from pageindex import page_index
tree = page_index(
doc="reports/2023-annual-report.pdf",
model="gpt-4o-2024-11-20",
toc_check_page_num=30,
max_page_num_each_node=12,
max_token_num_each_node=25000,
if_add_node_id="yes",
if_add_node_summary="yes",
if_add_doc_description="yes",
if_add_node_text="yes"
)
The returned dictionary contains doc_name, doc_description, and structure (the hierarchical tree). Each node includes title, node_id, start_index, end_index, and optional summary and text fields.
Flatten the tree for easier processing using structure_to_list from pageindex/utils.py:
from pageindex.utils import structure_to_list
nodes = structure_to_list(tree["structure"])
Implementing Reasoning-Based Retrieval
Replace vector similarity search with LLM-driven node selection:
import openai
import os
openai.api_key = os.getenv("CHATGPT_API_KEY")
def select_nodes(query, nodes):
# Build prompt with titles and summaries
titles = "\n".join(
f"{i}. {n['title']} – {n.get('summary', '')}"
for i, n in enumerate(nodes)
)
prompt = f"""You are given a list of document sections with short summaries:
{titles}
Which section(s) best answer the user question?
Question: {query}
Respond with the index numbers (comma-separated) only."""
response = openai.ChatCompletion.create(
model="gpt-4o-2024-11-20",
messages=[{"role": "user", "content": prompt}]
)
answer = response.choices[0].message.content.strip()
return [int(i) for i in answer.split(",") if i.strip().isdigit()]
# Usage
chosen_indices = select_nodes("How did the Fed's policy change in 2023?", nodes)
context = "\n".join(nodes[i]["text"] for i in chosen_indices)
Hybrid Retrieval with Vector Embeddings
Combine PageIndex with embeddings for fine-grained search within nodes:
from sentence_transformers import SentenceTransformer
import numpy as np
embedder = SentenceTransformer("all-MiniLM-L6-v2")
def hybrid_retrieve(question, nodes):
# Step 1: Select candidate nodes using PageIndex structure
candidate_ids = select_nodes(question, nodes)
# Step 2: Embed question
q_vec = embedder.encode(question, normalize_embeddings=True)
# Step 3: Find best match within candidates using embeddings
best_node, best_score = None, -np.inf
for nid in candidate_ids:
node_text = nodes[nid]["text"]
doc_vec = embedder.encode(node_text, normalize_embeddings=True)
score = np.dot(q_vec, doc_vec)
if score > best_score:
best_score, best_node = score, nid
return nodes[best_node]["text"]
# Usage
context = hybrid_retrieve("What were the Q3 revenue figures?", nodes)
Indexing Markdown Documentation
For Markdown files, use the md_to_tree function from pageindex/page_index_md.py:
import asyncio
from pageindex.page_index_md import md_to_tree
async def index_markdown(md_path):
result = await md_to_tree(
md_path=md_path,
if_thinning=True,
min_token_threshold=5000,
if_add_node_summary="yes",
summary_token_threshold=200,
model="gpt-4o-2024-11-20",
if_add_doc_description="yes",
if_add_node_text="yes",
if_add_node_id="yes"
)
return result
# Run
tree_md = asyncio.run(index_markdown("docs/technical-spec.md"))
Summary
-
PageIndex generates hierarchical tree structures from PDFs and Markdown using
page_indexandmd_to_tree, preserving document sections with exact page ranges instead of arbitrary chunks. -
Integration replaces vector-store chunking with reasoning-based retrieval, where LLMs navigate node titles and summaries to select relevant sections before consuming text.
-
Core functions include
page_index_maininrun_pageindex.pyfor orchestration,structure_to_listinpageindex/utils.pyfor flattening trees, andselect_nodespatterns for LLM-driven retrieval. -
Hybrid approaches combine PageIndex for coarse retrieval with embeddings for fine-grained search within selected nodes, optimizing both precision and computational efficiency.
-
Configuration supports model-agnostic LLM selection through
ConfigLoader, with customizable prompts for TOC detection and transformation inpageindex/page_index.py.
Frequently Asked Questions
How does PageIndex handle PDFs without a table of contents?
When no TOC is detected, PageIndex uses the process_no_toc function in pageindex/page_index.py to infer section boundaries by prompting the LLM to analyze document structure and identify logical chapter breaks based on content flow and formatting cues.
Can I use PageIndex with local LLMs instead of OpenAI?
Yes, PageIndex is model-agnostic. The ConfigLoader class in pageindex/utils.py manages LLM configuration, and you can modify the ChatGPT_API wrapper functions to route requests to local endpoints by changing the base URL and model parameters while maintaining the same JSON response format.
What is the difference between page_index and page_index_main?
page_index in pageindex/page_index.py is the core processing function that accepts configuration parameters directly, while page_index_main in run_pageindex.py serves as the CLI entry point that builds a configuration object from command-line arguments before invoking the core function.
How do I merge PageIndex trees from multiple documents?
Since PageIndex outputs standard JSON objects containing doc_name, doc_description, and structure arrays, you can merge trees by combining the structure lists under a new parent node or maintaining separate document roots in a dictionary keyed by doc_name, then using structure_to_list from pageindex/utils.py to flatten specific branches for retrieval.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →