# How the Verification Process in PageIndex Ensures Accurate Page Indices

> Discover how the PageIndex verification process uses LLM fuzzy matching to ensure accurate page indices. Learn about its iterative correction for high-confidence results.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: architecture
- Published: 2026-02-16

---

**The verification process in PageIndex validates every generated page number by cross-referencing candidate table-of-contents entries against actual document content using LLM-based fuzzy matching, achieving high-confidence indices through iterative correction and accuracy thresholds.**

The VectifyAI/PageIndex repository implements a robust verification pipeline that guarantees precise page-number assignments in automatically generated tables of contents. By combining candidate generation with rigorous content validation, the verification process in PageIndex eliminates hallucinated page references and ensures that each section title actually appears on its assigned page.

## Overview of the PageIndex Verification Pipeline

The verification workflow operates as a four-stage pipeline implemented primarily in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py). This architecture separates candidate generation from validation, allowing the system to detect and correct errors before finalizing the table of contents.

The pipeline handles edge cases such as out-of-range page numbers, sections that span multiple pages, and documents where the initial TOC generation produces low-confidence results. By employing concurrent LLM checks and statistical sampling, the system balances thoroughness with computational efficiency.

## Step-by-Step Verification Workflow

### 1. Generating the Candidate TOC

The process begins in `meta_processor`, which creates an initial table of contents with potential `physical_index` values (page numbers) using one of three strategies: `process_toc_with_page_numbers`, `process_toc_no_page_numbers`, or `process_no_toc`. These functions populate the candidate list that subsequent stages will validate.

### 2. Validating Page-Index Bounds

Before content verification begins, `validate_and_truncate_physical_indices` filters the candidate list to remove any `physical_index` entries that exceed the document length. This prevents out-of-range errors during the content-checking phase and ensures that verification only proceeds against valid page references.

### 3. Executing the Verification Check

The core verification logic resides in `verify_toc`, which implements a multi-step validation process:

**Early Abort Conditions:** The function first identifies the highest non-`None` page number in the candidate list. If no valid indices exist, or if the highest index is suspiciously small (less than half the document length), verification aborts early and returns zero accuracy.

**Sampling Strategy:** Depending on configuration, `verify_toc` either checks all TOC items or selects a random subset of size `N`. Each sampled item is annotated with its original list position (`list_index`) to maintain traceability.

**Concurrent Content Checking:** For each sampled item, `check_title_appearance` executes concurrently. This helper constructs a prompt asking the LLM to determine whether the section title **appears or starts** on the specified page text, using fuzzy matching that ignores whitespace variations. The LLM returns a JSON response with `answer: "yes"` or `"no"`.

**Accuracy Calculation:** Results are aggregated to compute an accuracy score equal to `correct_count / checked_count`. Items receiving `"no"` responses are collected as `incorrect_results` for potential correction.

### 4. Handling Post-Verification Results

Back in `meta_processor`, the returned `accuracy` and `incorrect_results` drive subsequent actions:

- **Perfect Accuracy (1.0):** The TOC is accepted immediately without modification.
- **High Accuracy (> 0.6):** When accuracy exceeds 60% but incorrect items exist, `fix_incorrect_toc_with_retries` re-runs the LLM on failing entries, attempting up to three retries to resolve mismatches.
- **Low Accuracy:** If accuracy falls below the threshold, the processor falls back to an alternative mode (such as generating a TOC without page numbers) and repeats the entire pipeline.

## Code Implementation Details

The verification system relies on three critical functions defined in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py):

- **`verify_toc(pages, toc, start_index, model, toc_check_page_num, ...)`**: Orchestrates the entire verification workflow, implementing sampling logic and concurrency management.
- **`check_title_appearance(page_text, title, model)`**: Constructs the LLM prompt for fuzzy title matching and parses the JSON response to determine page presence.
- **`validate_and_truncate_physical_indices(toc, num_pages)`**: Pre-filters the candidate TOC to eliminate out-of-range page references before expensive LLM verification.

The retry mechanism `fix_incorrect_toc_with_retries` operates as a corrective loop, specifically targeting entries flagged during the initial verification pass.

## Practical Examples

### Basic Usage: Building a Verified Page Index

```python
from pageindex.page_index import page_index

# Initialize with a PDF document

doc = "tests/pdfs/2023-annual-report.pdf"

# Generate TOC with verified page numbers

toc = page_index(
    doc,
    model="gpt-4o-mini",
    toc_check_page_num=20,        # Check up to 20 items

    max_page_num_each_node=5,
    max_token_num_each_node=1500
)

# Result contains verified physical indices

print(toc)

# Output: [{'title': 'Executive Summary', 'physical_index': 3, ...}, ...]

```

### Advanced: Direct Verification API

```python
import asyncio
from pageindex.page_index import verify_toc

# Prepare document pages as (text, token_count) tuples

pages = [
    ("Introduction text...", 150),
    ("Executive Summary content...", 200),
    # ... additional pages

]

# Candidate TOC with unverified page numbers

candidate_toc = [
    {"title": "Executive Summary", "physical_index": 3},
    {"title": "Financial Highlights", "physical_index": 7},
]

# Run verification

accuracy, incorrect = asyncio.run(
    verify_toc(
        pages,
        candidate_toc,
        start_index=1,
        model="gpt-4o-mini",
        toc_check_page_num=10
    )
)

print(f"Accuracy: {accuracy*100:.1f}%")
print(f"Failed items: {incorrect}")

```

## Key Source Files

The verification pipeline is implemented across the following modules:

- **[`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)** – Core engine containing `verify_toc`, `check_title_appearance`, `validate_and_truncate_physical_indices`, and the retry logic `fix_incorrect_toc_with_retries`.
- **[`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py)** – Helper utilities including `ChatGPT_API_async` and `extract_json` for LLM communication and response parsing.
- **[`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py)** – Command-line interface demonstrating how to invoke the indexer with verification enabled.
- **[`README.md`](https://github.com/VectifyAI/PageIndex/blob/main/README.md)** – High-level documentation and usage examples for the library.

## Summary

The verification process in PageIndex ensures accurate page indices through a systematic four-stage pipeline:

- **Pre-validation** filters out-of-range page numbers before expensive LLM checks.
- **Fuzzy matching** uses concurrent LLM calls to verify that section titles actually appear on assigned pages.
- **Statistical sampling** balances thoroughness with efficiency by checking either all items or random subsets.
- **Iterative correction** automatically retries failed entries or falls back to alternative generation modes when accuracy falls below thresholds.

By cross-referencing every generated index against actual document content, PageIndex eliminates hallucinated page references and delivers high-confidence tables of contents.

## Frequently Asked Questions

### How does PageIndex verify that a section title actually appears on a specific page?

PageIndex uses the `check_title_appearance` function to build an LLM prompt that asks whether the section title **appears or starts** on the target page text. The system employs fuzzy matching that ignores whitespace variations, and the LLM returns a JSON response with `answer: "yes"` or `"no"` to indicate presence. This check runs concurrently for all sampled TOC items to maximize efficiency.

### What happens when the verification process detects incorrect page numbers?

When verification identifies mismatches, the system routes results through `fix_incorrect_toc_with_retries` if the overall accuracy exceeds 60%. This function re-runs the LLM on failing entries for up to three retry attempts to resolve errors. If accuracy falls below the threshold or retries fail, the processor falls back to an alternative generation mode—such as creating a TOC without page numbers—and repeats the entire verification pipeline.

### What accuracy threshold does PageIndex require to accept a table of contents?

PageIndex accepts a TOC immediately when verification returns **perfect accuracy (1.0)**. For results between **0.6 and 1.0**, the system attempts to fix incorrect entries through retries. Accuracy below 0.6 triggers a fallback to alternative processing modes, ensuring that only high-confidence indices reach the final output.

### How does PageIndex handle documents where no valid page indices are found?

During the initial verification phase, `verify_toc` checks for the highest non-`None` page number in the candidate list. If no valid indices exist, or if the highest index is suspiciously small (less than half the document length), the function aborts early and returns zero accuracy. This triggers the fallback mechanism in `meta_processor`, which switches to a mode that generates TOC entries without page numbers and attempts verification again.