How PageIndex Achieves State-of-the-Art Accuracy on FinanceBench

PageIndex reaches 98.7% accuracy on FinanceBench by replacing vector-based retrieval with a hierarchical tree index and reasoning-based navigation, eliminating arbitrary chunking and semantic similarity assumptions that degrade performance on complex financial reasoning tasks.

PageIndex is an open-source document indexing system developed by VectifyAI that reimagines how large language models interact with long professional documents. Unlike traditional retrieval-augmented generation (RAG) systems that rely on vector similarity, PageIndex constructs a semantic table-of-contents tree that mirrors how human experts navigate financial reports. This architecture enables the state-of-the-art FinanceBench accuracy by allowing multi-step logical deduction rather than surface-level text matching.

The 98.7% FinanceBench Accuracy Breakdown

FinanceBench is a rigorous benchmark testing financial question-answering across complex, multi-step reasoning scenarios involving SEC filings and earnings reports. PageIndex achieves 98.7% accuracy according to the project README【^73†L73-L75】 by abandoning the two core assumptions that limit traditional RAG systems: that semantic similarity equals relevance, and that documents should be split into arbitrary token-length chunks.

Instead, PageIndex treats the document as a navigable hierarchy. The system parses PDFs into a tree-structured index where each node represents a natural document section, preserving the original author-intended organization. This allows the LLM to perform reasoning-based retrieval—traversing the tree to locate specific fiscal-year-end balances or corresponding footnotes through logical deduction rather than vector matching.

Core Architecture Components

Tree-Structured Index Generation

The foundation of PageIndex FinanceBench accuracy lies in its semantic parsing pipeline. The system first detects whether a table of contents exists using check_toc and toc_extractor in pageindex/page_index.py【^19†L19-L36】. If no TOC is present, the process_no_toc function【^68†L68-L84】 initiates hierarchical generation via generate_toc_init【^334†L334-L359】.

This process creates a coarse-to-fine outline where parent nodes represent major sections (e.g., "Financial Statements") and leaf nodes contain specific disclosures. By preserving the document's natural hierarchy, the system eliminates the context fragmentation that occurs when arbitrary chunking splits related financial data across unrelated blocks.

Physical Page Tagging

Every segment of raw PDF text is wrapped with <physical_index_X> tags that mark exact page numbers. This implementation in add_page_number_to_toc【^53†L53-L71】 and the toc_extractor【^19†L19-L36】 enables precise mapping from tree nodes to original document locations.

Physical page tagging eliminates "off-by-one" errors that plague vector-based similarity search, where retrieved chunks might come from adjacent but irrelevant pages. For FinanceBench questions requiring exact figure verification (e.g., "What was the depreciation expense in Q3 2023?"), this precision ensures the LLM retrieves the correct table cell rather than a similar-sounding paragraph from a different quarter.

LLM-Driven Title Verification

PageIndex guarantees structural accuracy through check_title_appearance【^13†L13-L46】, an asynchronous function that verifies whether a node's title actually appears at the start of its candidate page. This validation step prevents phantom sections—structural nodes that don't correspond to actual document content—from contaminating the index.

By ensuring inferred structure matches source text, this component maintains the integrity required for complex financial reasoning. When navigating to "Note 5: Income Taxes," the system confirms this header exists before treating the subsequent content as tax-related disclosures.

Iterative Validation and Repair

The system implements a robust feedback loop through verify_toc【^91†L91-L124】 and fix_incorrect_toc_with_retries【^70†L70-L87】. After initial tree construction, PageIndex samples nodes and re-checks them with the LLM, automatically repairing mismatches through iterative refinement.

This validation layer reduces error propagation from early parsing mistakes. For lengthy SEC 10-K filings spanning hundreds of pages, this ensures the final index accurately represents the document's true structure, enabling reliable retrieval of specific financial metrics required by FinanceBench questions.

Why PageIndex Abandons Vector Databases

Traditional RAG systems assume that semantic similarity equals relevance—that text chunks with similar embeddings to the query must contain the answer. PageIndex explicitly rejects this assumption, as documented in the repository's design overview【^73†L73-L74】.

The system never splits documents into arbitrary token-length chunks nor stores embeddings in a vector store. This elimination of chunking prevents the fragmentation of tabular financial data and multi-page disclosures that require holistic context. For FinanceBench tasks involving cross-references between financial statements and footnotes, maintaining document-level coherence is essential for accurate reasoning.

Reasoning-Based Retrieval Pipeline

During query execution, PageIndex performs a tree search over the hierarchical index, allowing the LLM to reason about which branch best answers the question. This navigation occurs through the tree_parser and meta_processor components【^21†L21-L33】【^55†L55-L66】.

Unlike vector retrieval, which returns isolated chunks, this approach enables multi-step logical deduction. For example, answering "What was the depreciation expense in fiscal year 2023?" requires locating the fiscal year definition, finding the balance sheet, identifying the property and equipment line item, then cross-referencing the corresponding footnote. The tree structure allows the LLM to traverse this logical path precisely, mimicking how financial analysts navigate reports.

Implementation Example

The following examples demonstrate how to invoke PageIndex on financial PDFs and retrieve the hierarchical tree structure with optional node summaries.


# Basic indexing without summaries

from pageindex.page_index import page_index

# Path to a financial report PDF (e.g., an SEC 10-K)

pdf_path = "tests/pdfs/q1-fy25-earnings.pdf"

# Run PageIndex with default settings (gpt-4o, no node text, no summaries)

result = page_index(pdf_path)

print("Document name:", result["doc_name"])
print("Top-level sections:")
for node in result["structure"]:
    print(f"- {node['title']} (pages {node['start_index']}-{node['end_index']})")

# Full indexing with node text and summaries for downstream QA

from pageindex.page_index import page_index

pdf_path = "tests/pdfs/2023-annual-report.pdf"

# Enable node text and LLM-generated summaries

result = page_index(
    pdf_path,
    if_add_node_text="yes",
    if_add_node_summary="yes",
    model="gpt-4o-2024-11-20",          # same model used in the benchmark

    max_page_num_each_node=15,
    max_token_num_each_node=20000,
)

# The returned structure now contains text and summary fields for each node

for node in result["structure"]:
    print(f"=== {node['title']} ===")
    print("Summary:", node.get("summary", "(no summary)"))
    # Access the full extracted text if needed

    # print(node["text"])

Both examples automatically execute the complete pipeline: detecting the table of contents via check_toc and toc_extractor, building the hierarchical index through generate_toc_init, validating structure with verify_toc and fix_incorrect_toc_with_retries, and optionally generating summaries via generate_node_summary in pageindex/utils.py.

Summary

PageIndex achieves 98.7% accuracy on FinanceBench through a fundamentally different approach to document retrieval:

  • Hierarchical tree indexing replaces arbitrary chunking, preserving document structure via generate_toc_init and process_no_toc in pageindex/page_index.py.
  • Physical page tagging ensures precise location mapping through <physical_index_X> tags, eliminating off-by-one errors common in vector search.
  • Iterative LLM validation guarantees structural integrity via check_title_appearance, verify_toc, and fix_incorrect_toc_with_retries.
  • Reasoning-based retrieval enables multi-step logical deduction through tree navigation, mimicking expert financial analyst behavior rather than relying on semantic similarity.
  • Zero vector dependencies removes the "similarity equals relevance" assumption that degrades performance on domain-specific financial reasoning tasks.

Frequently Asked Questions

How does PageIndex differ from traditional RAG systems?

Traditional RAG systems split documents into arbitrary token-length chunks and retrieve via vector similarity, assuming that semantically similar text contains the answer. PageIndex eliminates both practices, instead building a hierarchical tree index that preserves the original document structure and using reasoning-based navigation to locate precise answers. This architectural difference explains the 98.7% FinanceBench accuracy versus typical RAG performance on multi-step financial reasoning.

What specific validation steps ensure the tree structure matches the actual PDF content?

PageIndex implements a three-layer validation system. First, check_title_appearance verifies that inferred section titles actually appear at the start of candidate pages. Second, verify_toc samples nodes from the initial tree and re-checks them with the LLM. Third, fix_incorrect_toc_with_retries automatically repairs any mismatches through iterative refinement. This pipeline prevents phantom sections and ensures the final index accurately represents the source document.

Can PageIndex handle PDFs that already contain a table of contents?

Yes. The system first attempts to extract existing TOCs via check_toc and toc_extractor in pageindex/page_index.py. If a valid TOC is detected, PageIndex uses it as the foundation for the hierarchical index. If no TOC exists or the existing one is incomplete, the system falls back to process_no_toc and generate_toc_init to automatically construct the semantic hierarchy from scratch.

Why does eliminating vector databases improve financial document QA accuracy?

Vector databases rely on the assumption that semantic similarity equates to relevance, which fails on domain-specific financial reasoning tasks. Financial questions often require locating specific figures that share no semantic similarity with the query terms (e.g., finding a depreciation expense value). By removing vector stores and arbitrary chunking, PageIndex preserves the full context of tables, footnotes, and cross-references, enabling precise retrieval through structural navigation rather than approximate similarity matching.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →