How the Verification Process in PageIndex Ensures Accurate Page Indices
The verification process in PageIndex validates every generated page number by cross-referencing candidate table-of-contents entries against actual document content using LLM-based fuzzy matching, achieving high-confidence indices through iterative correction and accuracy thresholds.
The VectifyAI/PageIndex repository implements a robust verification pipeline that guarantees precise page-number assignments in automatically generated tables of contents. By combining candidate generation with rigorous content validation, the verification process in PageIndex eliminates hallucinated page references and ensures that each section title actually appears on its assigned page.
Overview of the PageIndex Verification Pipeline
The verification workflow operates as a four-stage pipeline implemented primarily in pageindex/page_index.py. This architecture separates candidate generation from validation, allowing the system to detect and correct errors before finalizing the table of contents.
The pipeline handles edge cases such as out-of-range page numbers, sections that span multiple pages, and documents where the initial TOC generation produces low-confidence results. By employing concurrent LLM checks and statistical sampling, the system balances thoroughness with computational efficiency.
Step-by-Step Verification Workflow
1. Generating the Candidate TOC
The process begins in meta_processor, which creates an initial table of contents with potential physical_index values (page numbers) using one of three strategies: process_toc_with_page_numbers, process_toc_no_page_numbers, or process_no_toc. These functions populate the candidate list that subsequent stages will validate.
2. Validating Page-Index Bounds
Before content verification begins, validate_and_truncate_physical_indices filters the candidate list to remove any physical_index entries that exceed the document length. This prevents out-of-range errors during the content-checking phase and ensures that verification only proceeds against valid page references.
3. Executing the Verification Check
The core verification logic resides in verify_toc, which implements a multi-step validation process:
Early Abort Conditions: The function first identifies the highest non-None page number in the candidate list. If no valid indices exist, or if the highest index is suspiciously small (less than half the document length), verification aborts early and returns zero accuracy.
Sampling Strategy: Depending on configuration, verify_toc either checks all TOC items or selects a random subset of size N. Each sampled item is annotated with its original list position (list_index) to maintain traceability.
Concurrent Content Checking: For each sampled item, check_title_appearance executes concurrently. This helper constructs a prompt asking the LLM to determine whether the section title appears or starts on the specified page text, using fuzzy matching that ignores whitespace variations. The LLM returns a JSON response with answer: "yes" or "no".
Accuracy Calculation: Results are aggregated to compute an accuracy score equal to correct_count / checked_count. Items receiving "no" responses are collected as incorrect_results for potential correction.
4. Handling Post-Verification Results
Back in meta_processor, the returned accuracy and incorrect_results drive subsequent actions:
- Perfect Accuracy (1.0): The TOC is accepted immediately without modification.
- High Accuracy (> 0.6): When accuracy exceeds 60% but incorrect items exist,
fix_incorrect_toc_with_retriesre-runs the LLM on failing entries, attempting up to three retries to resolve mismatches. - Low Accuracy: If accuracy falls below the threshold, the processor falls back to an alternative mode (such as generating a TOC without page numbers) and repeats the entire pipeline.
Code Implementation Details
The verification system relies on three critical functions defined in pageindex/page_index.py:
verify_toc(pages, toc, start_index, model, toc_check_page_num, ...): Orchestrates the entire verification workflow, implementing sampling logic and concurrency management.check_title_appearance(page_text, title, model): Constructs the LLM prompt for fuzzy title matching and parses the JSON response to determine page presence.validate_and_truncate_physical_indices(toc, num_pages): Pre-filters the candidate TOC to eliminate out-of-range page references before expensive LLM verification.
The retry mechanism fix_incorrect_toc_with_retries operates as a corrective loop, specifically targeting entries flagged during the initial verification pass.
Practical Examples
Basic Usage: Building a Verified Page Index
from pageindex.page_index import page_index
# Initialize with a PDF document
doc = "tests/pdfs/2023-annual-report.pdf"
# Generate TOC with verified page numbers
toc = page_index(
doc,
model="gpt-4o-mini",
toc_check_page_num=20, # Check up to 20 items
max_page_num_each_node=5,
max_token_num_each_node=1500
)
# Result contains verified physical indices
print(toc)
# Output: [{'title': 'Executive Summary', 'physical_index': 3, ...}, ...]
Advanced: Direct Verification API
import asyncio
from pageindex.page_index import verify_toc
# Prepare document pages as (text, token_count) tuples
pages = [
("Introduction text...", 150),
("Executive Summary content...", 200),
# ... additional pages
]
# Candidate TOC with unverified page numbers
candidate_toc = [
{"title": "Executive Summary", "physical_index": 3},
{"title": "Financial Highlights", "physical_index": 7},
]
# Run verification
accuracy, incorrect = asyncio.run(
verify_toc(
pages,
candidate_toc,
start_index=1,
model="gpt-4o-mini",
toc_check_page_num=10
)
)
print(f"Accuracy: {accuracy*100:.1f}%")
print(f"Failed items: {incorrect}")
Key Source Files
The verification pipeline is implemented across the following modules:
pageindex/page_index.py– Core engine containingverify_toc,check_title_appearance,validate_and_truncate_physical_indices, and the retry logicfix_incorrect_toc_with_retries.pageindex/utils.py– Helper utilities includingChatGPT_API_asyncandextract_jsonfor LLM communication and response parsing.run_pageindex.py– Command-line interface demonstrating how to invoke the indexer with verification enabled.README.md– High-level documentation and usage examples for the library.
Summary
The verification process in PageIndex ensures accurate page indices through a systematic four-stage pipeline:
- Pre-validation filters out-of-range page numbers before expensive LLM checks.
- Fuzzy matching uses concurrent LLM calls to verify that section titles actually appear on assigned pages.
- Statistical sampling balances thoroughness with efficiency by checking either all items or random subsets.
- Iterative correction automatically retries failed entries or falls back to alternative generation modes when accuracy falls below thresholds.
By cross-referencing every generated index against actual document content, PageIndex eliminates hallucinated page references and delivers high-confidence tables of contents.
Frequently Asked Questions
How does PageIndex verify that a section title actually appears on a specific page?
PageIndex uses the check_title_appearance function to build an LLM prompt that asks whether the section title appears or starts on the target page text. The system employs fuzzy matching that ignores whitespace variations, and the LLM returns a JSON response with answer: "yes" or "no" to indicate presence. This check runs concurrently for all sampled TOC items to maximize efficiency.
What happens when the verification process detects incorrect page numbers?
When verification identifies mismatches, the system routes results through fix_incorrect_toc_with_retries if the overall accuracy exceeds 60%. This function re-runs the LLM on failing entries for up to three retry attempts to resolve errors. If accuracy falls below the threshold or retries fail, the processor falls back to an alternative generation mode—such as creating a TOC without page numbers—and repeats the entire verification pipeline.
What accuracy threshold does PageIndex require to accept a table of contents?
PageIndex accepts a TOC immediately when verification returns perfect accuracy (1.0). For results between 0.6 and 1.0, the system attempts to fix incorrect entries through retries. Accuracy below 0.6 triggers a fallback to alternative processing modes, ensuring that only high-confidence indices reach the final output.
How does PageIndex handle documents where no valid page indices are found?
During the initial verification phase, verify_toc checks for the highest non-None page number in the candidate list. If no valid indices exist, or if the highest index is suspiciously small (less than half the document length), the function aborts early and returns zero accuracy. This triggers the fallback mechanism in meta_processor, which switches to a mode that generates TOC entries without page numbers and attempts verification again.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →