How the PageIndex Verification and Correction Pipeline Ensures Accurate PDF Table of Contents
PageIndex validates every table-of-contents entry by asking an LLM whether the section title actually appears on the predicted page, then iteratively repairs incorrect indices using surrounding correct entries as anchors until the TOC reaches 100% accuracy or triggers a fallback strategy.
PageIndex is an open-source Python library that automatically constructs hierarchical table-of-contents (TOC) structures from PDF documents. Because initial page predictions can be inaccurate, the library implements a verification and correction pipeline that validates each TOC entry against actual document content and repairs incorrect page indices using a retry-based anchor search strategy.
Overview of the Verification and Correction Pipeline
The pipeline operates in three distinct phases orchestrated by the meta_processor function in pageindex/page_index.py:
- Initial Validation – Truncate out-of-range page indices using
validate_and_truncate_physical_indices - Verification – Sample TOC entries and verify accuracy via
verify_toc - Correction – Repair failures using
fix_incorrect_toc_with_retriesif accuracy is between 0.6 and 1.0, or fall back to alternative strategies if accuracy is ≤0.6
The Verification Step (verify_toc)
Located in pageindex/page_index.py, the verify_toc async function validates whether predicted page numbers match actual content by querying an LLM.
How verify_toc Works
The function accepts a page_list (document pages), list_result (TOC entries), and optional sampling limit N. It performs these operations:
- Finds the last valid physical index to establish bounds
- Exits early with 0% accuracy if insufficient pages exist
- Attaches original list indices to sampled items
- Invokes
check_title_appearanceconcurrently for each sampled entry
async def verify_toc(page_list, list_result, start_index=1, N=None, model=None):
# ① Find the last valid physical_index
# ② If not enough pages → early exit (0% accuracy)
# ③ Choose which TOC items to sample (all or a random N)
# ④ Attach the original list index to each sampled item
# ⑤ Run `check_title_appearance` concurrently on every sampled item
# ⑥ Count "yes" answers → accuracy, collect the "no" results
return accuracy, incorrect_results
Sampling Strategy and Concurrency
When N is provided, verify_toc randomly samples N entries rather than checking the entire TOC, reducing LLM API costs while maintaining statistical validity. The function uses asyncio.gather to run check_title_appearance calls concurrently, maximizing throughput.
The Correction Step (fix_incorrect_toc)
When verification reveals inaccuracies, the system attempts repair via fix_incorrect_toc in pageindex/page_index.py.
Anchor-Based Page Window Search
For each incorrect entry, the algorithm:
- Locates the nearest preceding correct entry (
prev_correct) - Locates the nearest following correct entry (
next_correct) - Extracts the page window between these anchors
- Calls
single_toc_item_index_fixer(LLM) to identify the true start page within the window - Replaces the incorrect
physical_indexwith the corrected value
async def fix_incorrect_toc(toc_with_page_number, page_list, incorrect_results,
start_index=1, model=None, logger=None):
# ① Build a set of the bad list indices
# ② For each bad entry:
# - find prev_correct and next_correct anchors
# - gather pages between bounds
# - call single_toc_item_index_fixer to locate true start page
# - replace wrong physical_index with new one
return updated_toc, still_incorrect
Retry Logic with fix_incorrect_toc_with_retries
The wrapper fix_incorrect_toc_with_retries implements iterative refinement:
async def fix_incorrect_toc_with_retries(toc_with_page_number, page_list,
incorrect_results, start_index=1,
max_attempts=3, model=None, logger=None):
while current_incorrect:
current_toc, current_incorrect = await fix_incorrect_toc(...)
if attempts >= max_attempts:
break
return current_toc, current_incorrect
This retry mechanism allows up to three attempts by default, enabling cascading corrections where fixing one entry improves anchor context for subsequent fixes.
Orchestration in meta_processor
The meta_processor function in pageindex/page_index.py serves as the central controller:
- Initial TOC construction – Builds the raw TOC with predicted page numbers
- Validation – Calls
validate_and_truncate_physical_indicesto remove out-of-range indices - Verification – Invokes
verify_tocto calculate accuracy - Decision logic:
- 100% accuracy: Return TOC immediately
- >60% accuracy: Trigger
fix_incorrect_toc_with_retries - ≤60% accuracy: Fall back to alternative TOC-building strategies (e.g., without page numbers)
async def meta_processor(page_list, mode=None, toc_content=None,
toc_page_list=None, start_index=1, opt=None, logger=None):
# Build initial TOC
# Validate & truncate out-of-range indices
accuracy, incorrect_results = await verify_toc(...)
if accuracy == 1.0:
return toc_with_page_number
if accuracy > 0.6 and incorrect_results:
toc_with_page_number, _ = await fix_incorrect_toc_with_retries(...)
return toc_with_page_number
else:
# Fallback to alternative processing modes
...
Code Examples
Running the Full Pipeline with page_index
The simplest way to leverage the verification and correction pipeline is through the high-level page_index function:
from pageindex import page_index
# Path to a PDF (local file)
doc_path = "tests/pdfs/2023-annual-report.pdf"
# Run the full pipeline – it will verify and correct internally
result = page_index(
doc=doc_path,
model="gpt-4o-mini-2024-07-18",
toc_check_page_num=5, # how many pages after the TOC to scan for page numbers
max_page_num_each_node=30,
max_token_num_each_node=3000,
if_add_node_id="yes",
if_add_node_summary="yes",
if_add_doc_description="yes",
if_add_node_text="no"
)
print(result["structure"]) # hierarchical TOC with verified page numbers
This call internally reaches meta_processor, which triggers the verification (verify_toc) and correction (fix_incorrect_toc_with_retries) as described above.
Manual Verification with verify_toc
To inspect verification accuracy without running the full correction pipeline:
import asyncio
from pageindex.page_index import verify_toc, process_no_toc
async def demo():
# Assume `pages` is a list of (text, token_count) tuples already extracted from a PDF
pages = [...] # <list of page contents>
# Build a naive TOC without page numbers
raw_toc = process_no_toc(pages, start_index=1, model="gpt-4o-mini-2024-07-18")
# Verify it
accuracy, bad_items = await verify_toc(pages, raw_toc, start_index=1, model="gpt-4o-mini-2024-07-18")
print(f"Verification accuracy: {accuracy:.2%}")
if bad_items:
print("Incorrect entries:", bad_items)
asyncio.run(demo())
The verify_toc call returns the same accuracy and list of incorrect results that meta_processor later feeds into fix_incorrect_toc_with_retries.
Direct Correction with fix_incorrect_toc_with_retries
To programmatically repair a known bad TOC:
import asyncio
from pageindex.page_index import fix_incorrect_toc_with_retries
async def repair_demo(toc, pages):
# Suppose `toc` was built earlier and we already know it has errors
# First run verification to get the error list
from pageindex.page_index import verify_toc
accuracy, incorrect = await verify_toc(pages, toc, start_index=1, model="gpt-4o-mini-2024-07-18")
print(f"Before repair accuracy: {accuracy:.2%}")
# Apply the retry-based fixer
repaired_toc, still_bad = await fix_incorrect_toc_with_retries(
toc, pages, incorrect, start_index=1,
max_attempts=3, model="gpt-4o-mini-2024-07-18"
)
print("Repair finished. Remaining errors:", len(still_bad))
return repaired_toc
# asyncio.run(repair_demo(raw_toc, pages))
This snippet isolates the correction pipeline; it mirrors the logic inside meta_processor.
Key Files in the Verification and Correction Pipeline
| File | Role in the pipeline |
|---|---|
pageindex/page_index.py |
Core implementation containing verify_toc, fix_incorrect_toc, fix_incorrect_toc_with_retries, and the orchestration logic in meta_processor. |
pageindex/utils.py |
Helper functions for LLM calls (ChatGPT_API_async, extract_json), token counting, and JSON parsing that support the verification and correction stages. |
pageindex/config.yaml |
Default configuration including model names, accuracy thresholds (e.g., the 0.6 correction trigger), and retry limits that govern pipeline behavior. |
run_pageindex.py |
Command-line interface that invokes page_index, exercising the complete verification and correction workflow on user-provided PDFs. |
README.md |
Documentation describing the verification step and how to enable it via the toc_check_page_num parameter. |
Summary
- The verification and correction pipeline in PageIndex validates that each TOC entry actually begins on its predicted page using an LLM-based check (
verify_toc). - When accuracy falls below 100% but remains above the 0.6 threshold, the system triggers
fix_incorrect_toc_with_retries, which uses surrounding correct entries as anchors to locate true start pages. - The pipeline supports up to three retry attempts, allowing cascading corrections where fixing one entry improves anchor context for the next.
- If verification accuracy drops to 60% or below,
meta_processorabandons the current approach and falls back to alternative TOC-building strategies. - All core logic resides in
pageindex/page_index.py, with helper utilities inpageindex/utils.pyand configuration thresholds inpageindex/config.yaml.
Frequently Asked Questions
What triggers the correction pipeline in PageIndex?
The correction pipeline triggers when verify_toc returns an accuracy score between 0.6 and 1.0. According to the logic in meta_processor within pageindex/page_index.py, if accuracy is exactly 1.0, the TOC is returned immediately; if accuracy is greater than 0.6, the system invokes fix_incorrect_toc_with_retries to repair incorrect entries. Accuracy of 0.6 or below triggers a fallback strategy instead.
How does PageIndex verify that a TOC entry is on the correct page?
PageIndex uses the verify_toc function to validate entries by calling check_title_appearance (an LLM prompt) for each sampled TOC item. The function asks the model whether the section title actually appears on the predicted physical page, returning a binary yes/no response. It calculates overall accuracy by dividing the number of "yes" responses by the total sampled items, and returns both the accuracy score and a list of incorrect results for subsequent correction.
What is the anchor-based correction strategy in fix_incorrect_toc?
When fix_incorrect_toc processes an incorrect entry, it identifies the nearest preceding correct entry (prev_correct) and the nearest following correct entry (next_correct) to establish a bounded page window. It then calls single_toc_item_index_fixer (an LLM function) to search within that window and identify the true start page for the problematic title. This anchor-based approach constrains the search space to a narrow, contextually relevant range, minimizing hallucination and improving correction accuracy.
How many retry attempts does the correction pipeline allow?
The fix_incorrect_toc_with_retries wrapper allows a maximum of three attempts by default, controlled by the max_attempts parameter. During each iteration, the function calls fix_incorrect_toc to repair incorrect entries, then re-evaluates the remaining errors. The loop continues until either no incorrect entries remain or the attempt limit is reached, enabling cascading corrections where fixing one entry improves the anchor context for subsequent fixes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →