What Happens When `verify_toc` Fails with Low Accuracy in PageIndex

When verify_toc returns an accuracy score of 0.6 or lower in PageIndex, the system discards the current table-of-contents extraction and falls back to progressively simpler extraction strategies, eventually raising an exception if all methods fail.

PageIndex is an open-source document processing pipeline from VectifyAI that extracts hierarchical structure from PDFs. A critical quality gate in this pipeline is the verify_toc function, which validates whether extracted TOC titles actually appear in the document pages. Understanding how the system handles low accuracy scores is essential for debugging extraction failures and tuning the pipeline.

How verify_toc Calculates Accuracy in PageIndex

The verify_toc function in pageindex/page_index.py (lines 91-144) serves as the quality assurance layer for TOC extraction. It works by sampling a subset (or all) of the extracted TOC items and verifying their presence in the corresponding pages.

The function calls check_title_appearance concurrently across the sampled items, comparing the extracted title text against the actual page content. It then computes an accuracy value defined as:


accuracy = correct_items / checked_items

The function returns a tuple containing this accuracy score and a list of incorrect results:

(accuracy, incorrect_results)  # e.g., (0.42, [{...}, ...])

The Three Accuracy Tiers in PageIndex TOC Verification

The meta_processor function in pageindex/page_index.py (lines 71-89) implements a three-tier decision tree based on the accuracy score returned by verify_toc.

Perfect Accuracy (1.0): Immediate Acceptance

When verify_toc returns an accuracy of exactly 1.0 and the incorrect_results list is empty, the system considers the TOC extraction perfect. The pipeline immediately returns the toc_with_page_number without any further processing or correction attempts.

Medium Accuracy (>0.6): Automatic Repair with Retries

For accuracy scores greater than 0.6 but less than 1.0, PageIndex attempts to repair the incorrect entries automatically. The system calls fix_incorrect_toc_with_retries, which internally invokes fix_incorrect_toc up to three times to rewrite the problematic entries.

If the repair succeeds, the corrected TOC is returned; if not, the system proceeds with the partially corrected version.

Low Accuracy (≤0.6): Fallback Strategy Chain

When verify_toc reports an accuracy of 0.6 or lower, the system assumes the current TOC generation approach is fundamentally unreliable. Instead of attempting repairs on likely garbage data, PageIndex triggers a fallback chain that progressively simplifies the extraction strategy.

Low Accuracy Fallback Chain: From Page Numbers to Pure Content

The low-accuracy handling logic in meta_processor implements a degradation strategy that moves from complex TOC-guided extraction to simple content-based hierarchy detection.

The fallback progression follows this pattern:

  1. process_toc_with_page_numbers → If this mode fails with low accuracy, fall back to...
  2. process_toc_no_page_numbers → If this also fails with low accuracy, fall back to...
  3. process_no_toc → Pure content-based hierarchy extraction without TOC guidance.

The implementation in pageindex/page_index.py (lines 71-89) handles this via recursive calls to meta_processor with different mode parameters:

if accuracy == 1.0 and len(incorrect_results) == 0:
    return toc_with_page_number                # perfect – finish

if accuracy > 0.6 and len(incorrect_results) > 0:
    # moderate accuracy – try to repair

    toc_with_page_number, incorrect_results = await fix_incorrect_toc_with_retries(...)
    return toc_with_page_number
else:
    # low accuracy (≤ 0.6) – fallback strategies

    if mode == 'process_toc_with_page_numbers':
        return await meta_processor(..., mode='process_toc_no_page_numbers', ...)
    elif mode == 'process_toc_no_page_numbers':
        return await meta_processor(..., mode='process_no_toc', ...)
    else:
        raise Exception('Processing failed')

If the system reaches process_no_toc and still encounters low accuracy (or if this mode is already active when low accuracy is detected), it raises an exception indicating that processing has failed completely.

Code Example: Handling Low Accuracy in Your Pipeline

When integrating PageIndex into your document processing workflow, you can observe the accuracy scores and implement custom fallback logic:

from pageindex.page_index import verify_toc, meta_processor

# Assume page_list and initial TOC have been extracted

accuracy, incorrect_items = await verify_toc(
    page_list, 
    toc_items, 
    model=llm_client
)

if accuracy <= 0.6:
    print(f"Low accuracy detected ({accuracy:.2f}). Triggering fallback...")
    
    # The meta_processor will automatically handle the fallback chain

    final_toc = await meta_processor(
        page_list,
        mode='process_toc_with_page_numbers',  # Start with full TOC mode

        toc_content=raw_toc_text,
        toc_page_list=raw_toc_pages,
        start_index=1,
        opt=processing_options,
        logger=app_logger
    )
else:
    # Proceed with standard processing or repair

    final_toc = await meta_processor(...)

This pattern ensures that your application gracefully handles documents with unreliable table-of-contents metadata by automatically degrading to more robust extraction methods.

Summary

  • verify_toc validates extracted TOC titles against actual page content, returning an accuracy score and list of incorrect items.
  • Perfect accuracy (1.0) results in immediate acceptance of the TOC without modifications.
  • Medium accuracy (>0.6) triggers automatic repair via fix_incorrect_toc_with_retries, which attempts up to three correction cycles.
  • Low accuracy (≤0.6) initiates a fallback chain in meta_processor, progressing from process_toc_with_page_numbers → process_toc_no_page_numbers → process_no_toc.
  • If all fallback strategies fail, the system raises an exception indicating processing failure.

Frequently Asked Questions

What is the accuracy threshold for low accuracy in PageIndex?

PageIndex defines low accuracy as 0.6 (60%) or below. When verify_toc returns a score at or below this threshold, the system assumes the TOC extraction is unreliable and triggers fallback strategies rather than attempting repairs.

How many retry attempts does PageIndex make for medium accuracy entries?

For medium accuracy scores (greater than 0.6 but less than 1.0), PageIndex attempts automatic repair through fix_incorrect_toc_with_retries. This routine calls fix_incorrect_toc up to three times to correct problematic entries before returning the result.

What are the fallback modes when verify_toc fails with low accuracy?

When low accuracy is detected, meta_processor implements a three-stage degradation: first attempting process_toc_with_page_numbers, then falling back to process_toc_no_page_numbers, and finally resorting to process_no_toc (pure content-based hierarchy extraction). If the final mode fails, the system raises an exception.

Where is the low accuracy handling logic implemented in the PageIndex source code?

The primary logic for handling low accuracy resides in pageindex/page_index.py within the meta_processor function (lines 71-89). The verify_toc function (lines 91-144) in the same file computes the accuracy scores that trigger this handling. Helper functions for title verification are located in pageindex/utils.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →