# How the PageIndex Verification and Correction Pipeline Ensures Accurate PDF Table of Contents

> Discover how the PageIndex verification and correction pipeline ensures accurate PDF tables of contents. Learn about its LLM-powered validation and iterative repair strategies.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: internals
- Published: 2026-02-15

---

**PageIndex validates every table-of-contents entry by asking an LLM whether the section title actually appears on the predicted page, then iteratively repairs incorrect indices using surrounding correct entries as anchors until the TOC reaches 100% accuracy or triggers a fallback strategy.**

PageIndex is an open-source Python library that automatically constructs hierarchical table-of-contents (TOC) structures from PDF documents. Because initial page predictions can be inaccurate, the library implements a **verification and correction pipeline** that validates each TOC entry against actual document content and repairs incorrect page indices using a retry-based anchor search strategy.

## Overview of the Verification and Correction Pipeline

The pipeline operates in three distinct phases orchestrated by the `meta_processor` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py):

1. **Initial Validation** – Truncate out-of-range page indices using `validate_and_truncate_physical_indices`
2. **Verification** – Sample TOC entries and verify accuracy via `verify_toc`
3. **Correction** – Repair failures using `fix_incorrect_toc_with_retries` if accuracy is between 0.6 and 1.0, or fall back to alternative strategies if accuracy is ≤0.6

## The Verification Step (`verify_toc`)

Located in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), the `verify_toc` async function validates whether predicted page numbers match actual content by querying an LLM.

### How `verify_toc` Works

The function accepts a `page_list` (document pages), `list_result` (TOC entries), and optional sampling limit `N`. It performs these operations:

- Finds the last valid physical index to establish bounds
- Exits early with 0% accuracy if insufficient pages exist
- Attaches original list indices to sampled items
- Invokes `check_title_appearance` concurrently for each sampled entry

```python
async def verify_toc(page_list, list_result, start_index=1, N=None, model=None):
    # ① Find the last valid physical_index

    # ② If not enough pages → early exit (0% accuracy)

    # ③ Choose which TOC items to sample (all or a random N)

    # ④ Attach the original list index to each sampled item

    # ⑤ Run `check_title_appearance` concurrently on every sampled item

    # ⑥ Count "yes" answers → accuracy, collect the "no" results

    return accuracy, incorrect_results

```

### Sampling Strategy and Concurrency

When `N` is provided, `verify_toc` randomly samples `N` entries rather than checking the entire TOC, reducing LLM API costs while maintaining statistical validity. The function uses `asyncio.gather` to run `check_title_appearance` calls concurrently, maximizing throughput.

## The Correction Step (`fix_incorrect_toc`)

When verification reveals inaccuracies, the system attempts repair via `fix_incorrect_toc` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py).

### Anchor-Based Page Window Search

For each incorrect entry, the algorithm:

1. Locates the nearest preceding correct entry (`prev_correct`)
2. Locates the nearest following correct entry (`next_correct`)
3. Extracts the page window between these anchors
4. Calls `single_toc_item_index_fixer` (LLM) to identify the true start page within the window
5. Replaces the incorrect `physical_index` with the corrected value

```python
async def fix_incorrect_toc(toc_with_page_number, page_list, incorrect_results,
                            start_index=1, model=None, logger=None):
    # ① Build a set of the bad list indices

    # ② For each bad entry:

    #      - find prev_correct and next_correct anchors

    #      - gather pages between bounds

    #      - call single_toc_item_index_fixer to locate true start page

    #      - replace wrong physical_index with new one

    return updated_toc, still_incorrect

```

### Retry Logic with `fix_incorrect_toc_with_retries`

The wrapper `fix_incorrect_toc_with_retries` implements iterative refinement:

```python
async def fix_incorrect_toc_with_retries(toc_with_page_number, page_list,
                                         incorrect_results, start_index=1,
                                         max_attempts=3, model=None, logger=None):
    while current_incorrect:
        current_toc, current_incorrect = await fix_incorrect_toc(...)
        if attempts >= max_attempts: 
            break
    return current_toc, current_incorrect

```

This retry mechanism allows up to three attempts by default, enabling cascading corrections where fixing one entry improves anchor context for subsequent fixes.

## Orchestration in `meta_processor`

The `meta_processor` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) serves as the central controller:

1. **Initial TOC construction** – Builds the raw TOC with predicted page numbers
2. **Validation** – Calls `validate_and_truncate_physical_indices` to remove out-of-range indices
3. **Verification** – Invokes `verify_toc` to calculate accuracy
4. **Decision logic**:
   - **100% accuracy**: Return TOC immediately
   - **>60% accuracy**: Trigger `fix_incorrect_toc_with_retries`
   - **≤60% accuracy**: Fall back to alternative TOC-building strategies (e.g., without page numbers)

```python
async def meta_processor(page_list, mode=None, toc_content=None,
                         toc_page_list=None, start_index=1, opt=None, logger=None):
    # Build initial TOC

    # Validate & truncate out-of-range indices

    accuracy, incorrect_results = await verify_toc(...)
    if accuracy == 1.0:
        return toc_with_page_number
    if accuracy > 0.6 and incorrect_results:
        toc_with_page_number, _ = await fix_incorrect_toc_with_retries(...)
        return toc_with_page_number
    else:
        # Fallback to alternative processing modes

        ...

```

## Code Examples

### Running the Full Pipeline with `page_index`

The simplest way to leverage the verification and correction pipeline is through the high-level `page_index` function:

```python
from pageindex import page_index

# Path to a PDF (local file)

doc_path = "tests/pdfs/2023-annual-report.pdf"

# Run the full pipeline – it will verify and correct internally

result = page_index(
    doc=doc_path,
    model="gpt-4o-mini-2024-07-18",
    toc_check_page_num=5,          # how many pages after the TOC to scan for page numbers

    max_page_num_each_node=30,
    max_token_num_each_node=3000,
    if_add_node_id="yes",
    if_add_node_summary="yes",
    if_add_doc_description="yes",
    if_add_node_text="no"
)

print(result["structure"])   # hierarchical TOC with verified page numbers

```

This call internally reaches `meta_processor`, which triggers the verification (`verify_toc`) and correction (`fix_incorrect_toc_with_retries`) as described above.

### Manual Verification with `verify_toc`

To inspect verification accuracy without running the full correction pipeline:

```python
import asyncio
from pageindex.page_index import verify_toc, process_no_toc

async def demo():
    # Assume `pages` is a list of (text, token_count) tuples already extracted from a PDF

    pages = [...]                         # <list of page contents>

    # Build a naive TOC without page numbers

    raw_toc = process_no_toc(pages, start_index=1, model="gpt-4o-mini-2024-07-18")
    # Verify it

    accuracy, bad_items = await verify_toc(pages, raw_toc, start_index=1, model="gpt-4o-mini-2024-07-18")
    print(f"Verification accuracy: {accuracy:.2%}")
    if bad_items:
        print("Incorrect entries:", bad_items)

asyncio.run(demo())

```

The `verify_toc` call returns the same accuracy and list of incorrect results that `meta_processor` later feeds into `fix_incorrect_toc_with_retries`.

### Direct Correction with `fix_incorrect_toc_with_retries`

To programmatically repair a known bad TOC:

```python
import asyncio
from pageindex.page_index import fix_incorrect_toc_with_retries

async def repair_demo(toc, pages):
    # Suppose `toc` was built earlier and we already know it has errors

    # First run verification to get the error list

    from pageindex.page_index import verify_toc
    accuracy, incorrect = await verify_toc(pages, toc, start_index=1, model="gpt-4o-mini-2024-07-18")
    print(f"Before repair accuracy: {accuracy:.2%}")

    # Apply the retry-based fixer

    repaired_toc, still_bad = await fix_incorrect_toc_with_retries(
        toc, pages, incorrect, start_index=1,
        max_attempts=3, model="gpt-4o-mini-2024-07-18"
    )
    print("Repair finished. Remaining errors:", len(still_bad))
    return repaired_toc

# asyncio.run(repair_demo(raw_toc, pages))

```

This snippet isolates the correction pipeline; it mirrors the logic inside `meta_processor`.

## Key Files in the Verification and Correction Pipeline

| File | Role in the pipeline |
|------|---------------------|
| [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) | Core implementation containing `verify_toc`, `fix_incorrect_toc`, `fix_incorrect_toc_with_retries`, and the orchestration logic in `meta_processor`. |
| [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) | Helper functions for LLM calls (`ChatGPT_API_async`, `extract_json`), token counting, and JSON parsing that support the verification and correction stages. |
| [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) | Default configuration including model names, accuracy thresholds (e.g., the 0.6 correction trigger), and retry limits that govern pipeline behavior. |
| [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) | Command-line interface that invokes `page_index`, exercising the complete verification and correction workflow on user-provided PDFs. |
| [`README.md`](https://github.com/VectifyAI/PageIndex/blob/main/README.md) | Documentation describing the verification step and how to enable it via the `toc_check_page_num` parameter. |

## Summary

- The **verification and correction pipeline** in PageIndex validates that each TOC entry actually begins on its predicted page using an LLM-based check (`verify_toc`).
- When accuracy falls below 100% but remains above the 0.6 threshold, the system triggers `fix_incorrect_toc_with_retries`, which uses surrounding correct entries as anchors to locate true start pages.
- The pipeline supports up to three retry attempts, allowing cascading corrections where fixing one entry improves anchor context for the next.
- If verification accuracy drops to 60% or below, `meta_processor` abandons the current approach and falls back to alternative TOC-building strategies.
- All core logic resides in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), with helper utilities in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) and configuration thresholds in [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml).

## Frequently Asked Questions

### What triggers the correction pipeline in PageIndex?

The correction pipeline triggers when `verify_toc` returns an accuracy score between 0.6 and 1.0. According to the logic in `meta_processor` within [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), if accuracy is exactly 1.0, the TOC is returned immediately; if accuracy is greater than 0.6, the system invokes `fix_incorrect_toc_with_retries` to repair incorrect entries. Accuracy of 0.6 or below triggers a fallback strategy instead.

### How does PageIndex verify that a TOC entry is on the correct page?

PageIndex uses the `verify_toc` function to validate entries by calling `check_title_appearance` (an LLM prompt) for each sampled TOC item. The function asks the model whether the section title actually appears on the predicted physical page, returning a binary yes/no response. It calculates overall accuracy by dividing the number of "yes" responses by the total sampled items, and returns both the accuracy score and a list of incorrect results for subsequent correction.

### What is the anchor-based correction strategy in `fix_incorrect_toc`?

When `fix_incorrect_toc` processes an incorrect entry, it identifies the nearest preceding correct entry (`prev_correct`) and the nearest following correct entry (`next_correct`) to establish a bounded page window. It then calls `single_toc_item_index_fixer` (an LLM function) to search within that window and identify the true start page for the problematic title. This anchor-based approach constrains the search space to a narrow, contextually relevant range, minimizing hallucination and improving correction accuracy.

### How many retry attempts does the correction pipeline allow?

The `fix_incorrect_toc_with_retries` wrapper allows a maximum of three attempts by default, controlled by the `max_attempts` parameter. During each iteration, the function calls `fix_incorrect_toc` to repair incorrect entries, then re-evaluates the remaining errors. The loop continues until either no incorrect entries remain or the attempt limit is reached, enabling cascading corrections where fixing one entry improves the anchor context for subsequent fixes.