# How PageIndex's Reasoning-Based Retrieval Compares to Vector Similarity Search

> Discover how PageIndex reasoning-based retrieval outperforms vector similarity search. Achieve 98.7% accuracy on FinanceBench without embeddings or vector databases. Learn more about this LLM-powered innovation.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: comparison
- Published: 2026-02-16

---

**PageIndex replaces traditional vector similarity search with a reasoning-based retrieval system that uses an LLM to navigate a hierarchical table-of-contents tree, achieving 98.7% accuracy on FinanceBench without requiring embeddings or vector databases.**

PageIndex, an open-source project by VectifyAI, introduces a novel approach to document retrieval that eliminates the need for vector embeddings entirely. Unlike conventional retrieval-augmented generation (RAG) systems that rely on semantic similarity between query and text chunks, PageIndex implements **reasoning-based retrieval** that mimics how humans navigate physical documents—by reading the table of contents first. This architectural shift moves from opaque vector arithmetic to explicit logical reasoning over document structure.

## What Is Reasoning-Based Retrieval in PageIndex?

Reasoning-based retrieval in PageIndex is a multi-stage pipeline where an LLM reasons over a document's hierarchical structure rather than comparing embedding vectors. The system constructs a **tree index** that mirrors the document's logical organization—chapters, sections, and subsections—complete with explicit page numbers and LLM-generated summaries.

In [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), the core orchestration happens through functions like `page_index_main` and `tree_parser`. The retrieval process treats the query as a reasoning task: the LLM receives the question plus the current subtree, then decides whether to descend into child nodes or return the current node as relevant. This decision-making is logged in `thinking` fields within prompts like `check_title_appearance` and `check_title_appearance_in_start_concurrent`.

## Reasoning-Based Retrieval vs. Vector Similarity Search: Key Differences

The architectural divergence between these approaches impacts every layer of the retrieval stack:

| Aspect | Vector Similarity Search | PageIndex Reasoning-Based Retrieval |
|--------|--------------------------|-------------------------------------|
| **Index Structure** | Flat high-dimensional vectors stored in a vector DB (e.g., FAISS, Pinecone). | A **tree index** mirroring document hierarchy (sections, subsections, pages). The TOC is generated automatically by the LLM and stored as JSON (see [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) and [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)). |
| **Retrieval Query** | Embedding of the query → nearest-neighbor lookup in vector space. | LLM receives the **question + TOC tree** and **reasons** over the tree to decide which nodes to explore, then reads actual pages only for selected nodes. |
| **Data Stored** | Dense embeddings (often opaque). | **Explicit page numbers** (`physical_index`) and section titles, plus optional LLM-generated summaries (`summary`). No vectors persisted. |
| **Explainability** | Retrieval scores are numeric; hard to justify *why* a chunk matched. | LLM reasoning steps returned as JSON (`thinking` fields). Final TOC contains human-readable titles and page references, making the path traceable. |
| **Dependencies** | Requires separate vector DB service, embedding model, and chunking logic. | **Zero external DB** – only OpenAI LLM API. No chunking; sections delimited by headings and page tags (`<physical_index_X>`). |
| **Accuracy** | Limited by embedding quality; "semantic similarity ≠ relevance". | **≈98.7% accuracy** on FinanceBench by forcing LLM to reason over document structure. |
| **Scalability** | Scales with vector DB indexing; large corpora need more storage. | Works on single PDF without extra storage; tree size grows linearly with sections, not token count. |

## How Reasoning-Based Retrieval Works Under the Hood

The PageIndex pipeline in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) implements a six-stage reasoning architecture:

### Step 1: TOC Detection and Extraction

The system first identifies whether the PDF contains a table of contents using `toc_detector_single_page` and `find_toc_pages`. If a TOC exists, `toc_extractor` parses the raw text into structured data. If not, the system falls back to `process_no_toc` to generate a hierarchical structure purely from the document's heading patterns.

### Step 2: Tree Construction and Physical Index Assignment

The `toc_transformer` converts raw TOC entries into a JSON tree structure with nested `children` arrays. Then `toc_index_extractor` (via LLM prompts) assigns `physical_index` values to each node, tagging content with markers like `<physical_index_5>` to map logical sections to actual PDF pages. The `validate_and_truncate_physical_indices` function ensures page numbers remain within document bounds.

### Step 3: LLM-Driven Tree Search

The core reasoning happens in `tree_parser`. Instead of vector lookup, this function performs a **tree search** where the LLM evaluates whether a question matches the current node's title or requires descending into children. The prompt `check_title_appearance_in_start_concurrent` asks the LLM to reason about semantic relevance to section titles, returning JSON with `thinking` fields that explain the decision. Only when the LLM identifies a relevant node does the system retrieve the actual page content via `get_page_tokens` from [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py).

## Implementation: Building and Querying a PageIndex Tree

### Generate a PageIndex Tree from a PDF

The `page_index` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) serves as the main entry point for constructing the reasoning index:

```python
from pageindex.page_index import page_index

# Minimal call – the function reads the .env for the OpenAI key

tree = page_index(
    doc="tests/pdfs/2023-annual-report.pdf",   # any PDF path

    model="gpt-4o-2024-11-20",
    toc_check_page_num=20,
    max_page_num_each_node=10,
    max_token_num_each_node=20000,
    if_add_node_id="yes",
    if_add_node_summary="yes",
    if_add_doc_description="yes",
    if_add_node_text="no"
)

print(tree["structure"])   # hierarchical JSON with titles, page numbers, summaries

```

This generates a JSON tree where each node contains `physical_index` mappings and optional LLM-generated summaries via `generate_summaries_for_structure`.

### Perform a Reasoning-Based Query

To query the index, use `tree_parser` which implements the reasoning-based retrieval logic:

```python
import asyncio
from pageindex.page_index import tree_parser, get_page_tokens

# Load the PDF pages (token lengths are needed for internal chunking)

pages = get_page_tokens("tests/pdfs/2023-annual-report.pdf")

# Assume `toc_tree` is the output of the previous step

async def ask_question(question: str):
    # The tree_parser internally does a tree-search using the LLM

    result = await tree_parser(pages, opt=None)   # opt contains model config

    # The library wraps the question into LLM prompts that reason over the TOC

    return result

answer = asyncio.run(ask_question("What are the main risks highlighted in the 2023 annual report?"))
print(answer)   # Returns the most relevant node(s) with start/end indices

```

The `tree_parser` function coordinates with `meta_processor` to traverse the tree, using prompts like `check_title_appearance` to validate relevance at each level.

### Handle Documents Without a TOC

For PDFs lacking a table of contents, PageIndex automatically falls back to pure generation mode:

```python
from pageindex.page_index import page_index

tree = page_index(
    doc="tests/pdfs/four-lectures.pdf",
    model="gpt-4o-2024-11-20",
    # No TOC in this PDF; the library will auto-generate a hierarchical index.

)

```

This fallback is implemented in `process_no_toc` (lines 590-620 in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)), which analyzes heading patterns to construct the tree structure without explicit TOC pages.

## Summary

- **PageIndex** implements **reasoning-based retrieval** that replaces vector embeddings with LLM-driven navigation over a hierarchical document tree.
- The system constructs a **TOC tree** with explicit `physical_index` mappings via functions like `toc_extractor` and `toc_index_extractor` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py).
- **Retrieval** occurs through `tree_parser`, which prompts the LLM to reason about which tree nodes contain the answer, rather than performing nearest-neighbor vector lookup.
- The approach achieves **98.7% accuracy** on FinanceBench without requiring vector databases, embedding models, or chunking logic.
- **Explainability** is built-in through JSON `thinking` fields that trace the LLM's reasoning path, while **fallback mechanisms** like `process_no_toc` handle documents without explicit TOCs.

## Frequently Asked Questions

### What makes reasoning-based retrieval more accurate than vector search?

Traditional vector similarity search relies on embedding models that approximate semantic similarity, which often conflates related concepts with actually relevant content. PageIndex's reasoning-based retrieval forces the LLM to explicitly reason about document structure and section relevance through prompts like `check_title_appearance_in_start_concurrent`, yielding **98.7% accuracy** on FinanceBench compared to typical vector RAG baselines that struggle with precise page-level retrieval.

### Does PageIndex require a vector database?

No. PageIndex operates with **zero external database dependencies**. Unlike conventional RAG systems that require FAISS, Pinecone, or similar vector stores, PageIndex stores only a JSON tree structure with `physical_index` references and optional summaries. All retrieval logic happens via the OpenAI API through functions like `tree_parser` and `meta_processor` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py).

### How does PageIndex handle documents without a table of contents?

When no TOC is detected via `find_toc_pages` and `toc_detector_single_page`, PageIndex automatically invokes `process_no_toc` (implemented around lines 590-620 in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)). This fallback mode analyzes heading hierarchies and font patterns to generate a synthetic TOC tree, ensuring reasoning-based retrieval works even on unstructured PDFs that lack explicit navigation metadata.

### Is reasoning-based retrieval slower than vector similarity search?

While vector search relies on fast nearest-neighbor algorithms (often sub-millisecond), PageIndex trades raw speed for precision through LLM inference. The `tree_parser` function makes sequential API calls to reason about tree nodes, which introduces latency proportional to tree depth. However, this cost buys **explainability** (via `thinking` JSON fields) and **higher accuracy** (98.7% on FinanceBench), making it suitable for high-stakes domain-specific retrieval where precision outweighs microsecond latency requirements.