How PageIndex's Reasoning-Based Retrieval Compares to Vector Similarity Search
PageIndex replaces traditional vector similarity search with a reasoning-based retrieval system that uses an LLM to navigate a hierarchical table-of-contents tree, achieving 98.7% accuracy on FinanceBench without requiring embeddings or vector databases.
PageIndex, an open-source project by VectifyAI, introduces a novel approach to document retrieval that eliminates the need for vector embeddings entirely. Unlike conventional retrieval-augmented generation (RAG) systems that rely on semantic similarity between query and text chunks, PageIndex implements reasoning-based retrieval that mimics how humans navigate physical documents—by reading the table of contents first. This architectural shift moves from opaque vector arithmetic to explicit logical reasoning over document structure.
What Is Reasoning-Based Retrieval in PageIndex?
Reasoning-based retrieval in PageIndex is a multi-stage pipeline where an LLM reasons over a document's hierarchical structure rather than comparing embedding vectors. The system constructs a tree index that mirrors the document's logical organization—chapters, sections, and subsections—complete with explicit page numbers and LLM-generated summaries.
In pageindex/page_index.py, the core orchestration happens through functions like page_index_main and tree_parser. The retrieval process treats the query as a reasoning task: the LLM receives the question plus the current subtree, then decides whether to descend into child nodes or return the current node as relevant. This decision-making is logged in thinking fields within prompts like check_title_appearance and check_title_appearance_in_start_concurrent.
Reasoning-Based Retrieval vs. Vector Similarity Search: Key Differences
The architectural divergence between these approaches impacts every layer of the retrieval stack:
| Aspect | Vector Similarity Search | PageIndex Reasoning-Based Retrieval |
|---|---|---|
| Index Structure | Flat high-dimensional vectors stored in a vector DB (e.g., FAISS, Pinecone). | A tree index mirroring document hierarchy (sections, subsections, pages). The TOC is generated automatically by the LLM and stored as JSON (see pageindex/page_index_md.py and pageindex/page_index.py). |
| Retrieval Query | Embedding of the query → nearest-neighbor lookup in vector space. | LLM receives the question + TOC tree and reasons over the tree to decide which nodes to explore, then reads actual pages only for selected nodes. |
| Data Stored | Dense embeddings (often opaque). | Explicit page numbers (physical_index) and section titles, plus optional LLM-generated summaries (summary). No vectors persisted. |
| Explainability | Retrieval scores are numeric; hard to justify why a chunk matched. | LLM reasoning steps returned as JSON (thinking fields). Final TOC contains human-readable titles and page references, making the path traceable. |
| Dependencies | Requires separate vector DB service, embedding model, and chunking logic. | Zero external DB – only OpenAI LLM API. No chunking; sections delimited by headings and page tags (<physical_index_X>). |
| Accuracy | Limited by embedding quality; "semantic similarity ≠ relevance". | ≈98.7% accuracy on FinanceBench by forcing LLM to reason over document structure. |
| Scalability | Scales with vector DB indexing; large corpora need more storage. | Works on single PDF without extra storage; tree size grows linearly with sections, not token count. |
How Reasoning-Based Retrieval Works Under the Hood
The PageIndex pipeline in pageindex/page_index.py implements a six-stage reasoning architecture:
Step 1: TOC Detection and Extraction
The system first identifies whether the PDF contains a table of contents using toc_detector_single_page and find_toc_pages. If a TOC exists, toc_extractor parses the raw text into structured data. If not, the system falls back to process_no_toc to generate a hierarchical structure purely from the document's heading patterns.
Step 2: Tree Construction and Physical Index Assignment
The toc_transformer converts raw TOC entries into a JSON tree structure with nested children arrays. Then toc_index_extractor (via LLM prompts) assigns physical_index values to each node, tagging content with markers like <physical_index_5> to map logical sections to actual PDF pages. The validate_and_truncate_physical_indices function ensures page numbers remain within document bounds.
Step 3: LLM-Driven Tree Search
The core reasoning happens in tree_parser. Instead of vector lookup, this function performs a tree search where the LLM evaluates whether a question matches the current node's title or requires descending into children. The prompt check_title_appearance_in_start_concurrent asks the LLM to reason about semantic relevance to section titles, returning JSON with thinking fields that explain the decision. Only when the LLM identifies a relevant node does the system retrieve the actual page content via get_page_tokens from pageindex/utils.py.
Implementation: Building and Querying a PageIndex Tree
Generate a PageIndex Tree from a PDF
The page_index function in pageindex/page_index.py serves as the main entry point for constructing the reasoning index:
from pageindex.page_index import page_index
# Minimal call – the function reads the .env for the OpenAI key
tree = page_index(
doc="tests/pdfs/2023-annual-report.pdf", # any PDF path
model="gpt-4o-2024-11-20",
toc_check_page_num=20,
max_page_num_each_node=10,
max_token_num_each_node=20000,
if_add_node_id="yes",
if_add_node_summary="yes",
if_add_doc_description="yes",
if_add_node_text="no"
)
print(tree["structure"]) # hierarchical JSON with titles, page numbers, summaries
This generates a JSON tree where each node contains physical_index mappings and optional LLM-generated summaries via generate_summaries_for_structure.
Perform a Reasoning-Based Query
To query the index, use tree_parser which implements the reasoning-based retrieval logic:
import asyncio
from pageindex.page_index import tree_parser, get_page_tokens
# Load the PDF pages (token lengths are needed for internal chunking)
pages = get_page_tokens("tests/pdfs/2023-annual-report.pdf")
# Assume `toc_tree` is the output of the previous step
async def ask_question(question: str):
# The tree_parser internally does a tree-search using the LLM
result = await tree_parser(pages, opt=None) # opt contains model config
# The library wraps the question into LLM prompts that reason over the TOC
return result
answer = asyncio.run(ask_question("What are the main risks highlighted in the 2023 annual report?"))
print(answer) # Returns the most relevant node(s) with start/end indices
The tree_parser function coordinates with meta_processor to traverse the tree, using prompts like check_title_appearance to validate relevance at each level.
Handle Documents Without a TOC
For PDFs lacking a table of contents, PageIndex automatically falls back to pure generation mode:
from pageindex.page_index import page_index
tree = page_index(
doc="tests/pdfs/four-lectures.pdf",
model="gpt-4o-2024-11-20",
# No TOC in this PDF; the library will auto-generate a hierarchical index.
)
This fallback is implemented in process_no_toc (lines 590-620 in pageindex/page_index.py), which analyzes heading patterns to construct the tree structure without explicit TOC pages.
Summary
- PageIndex implements reasoning-based retrieval that replaces vector embeddings with LLM-driven navigation over a hierarchical document tree.
- The system constructs a TOC tree with explicit
physical_indexmappings via functions liketoc_extractorandtoc_index_extractorinpageindex/page_index.py. - Retrieval occurs through
tree_parser, which prompts the LLM to reason about which tree nodes contain the answer, rather than performing nearest-neighbor vector lookup. - The approach achieves 98.7% accuracy on FinanceBench without requiring vector databases, embedding models, or chunking logic.
- Explainability is built-in through JSON
thinkingfields that trace the LLM's reasoning path, while fallback mechanisms likeprocess_no_tochandle documents without explicit TOCs.
Frequently Asked Questions
What makes reasoning-based retrieval more accurate than vector search?
Traditional vector similarity search relies on embedding models that approximate semantic similarity, which often conflates related concepts with actually relevant content. PageIndex's reasoning-based retrieval forces the LLM to explicitly reason about document structure and section relevance through prompts like check_title_appearance_in_start_concurrent, yielding 98.7% accuracy on FinanceBench compared to typical vector RAG baselines that struggle with precise page-level retrieval.
Does PageIndex require a vector database?
No. PageIndex operates with zero external database dependencies. Unlike conventional RAG systems that require FAISS, Pinecone, or similar vector stores, PageIndex stores only a JSON tree structure with physical_index references and optional summaries. All retrieval logic happens via the OpenAI API through functions like tree_parser and meta_processor in pageindex/page_index.py.
How does PageIndex handle documents without a table of contents?
When no TOC is detected via find_toc_pages and toc_detector_single_page, PageIndex automatically invokes process_no_toc (implemented around lines 590-620 in pageindex/page_index.py). This fallback mode analyzes heading hierarchies and font patterns to generate a synthetic TOC tree, ensuring reasoning-based retrieval works even on unstructured PDFs that lack explicit navigation metadata.
Is reasoning-based retrieval slower than vector similarity search?
While vector search relies on fast nearest-neighbor algorithms (often sub-millisecond), PageIndex trades raw speed for precision through LLM inference. The tree_parser function makes sequential API calls to reason about tree nodes, which introduces latency proportional to tree depth. However, this cost buys explainability (via thinking JSON fields) and higher accuracy (98.7% on FinanceBench), making it suitable for high-stakes domain-specific retrieval where precision outweighs microsecond latency requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →