How Asynchronous Processing in PageIndex Enables Concurrent Title Verification

PageIndex leverages Python's asyncio framework to parallelize LLM-driven verification of table-of-contents entries, dramatically reducing wall-clock time by submitting multiple HTTP requests to the OpenAI API simultaneously rather than processing them sequentially.

When extracting structured data from large PDF documents, verifying that each table-of-contents (TOC) entry actually appears at the start of its claimed page can become a significant bottleneck. The VectifyAI/PageIndex repository solves this scalability challenge by implementing asynchronous processing in PageIndex to run concurrent title verification against large language models (LLMs), enabling efficient processing of documents containing hundreds of sections.

The Challenge: Verifying Hundreds of TOC Entries

PageIndex extracts a structured TOC from PDFs, mapping each entry to a physical page index via the physical_index field. Before trusting these mappings, the system must validate that the section title text actually appears at the beginning of the corresponding page content. Performing these checks sequentially for documents with extensive hierarchies would create prohibitive latency due to network round-trips to the LLM API, making concurrent title verification essential for performance.

Core Async Architecture

The concurrent verification logic centers on two primary async functions in pageindex/page_index.py, both utilizing asyncio.gather to parallelize LLM calls.

Building Concurrent Tasks with check_title_appearance_in_start_concurrent

Located at lines 74-99 in pageindex/page_index.py, this function constructs a list of coroutine tasks—one for every TOC entry that possesses a valid physical_index. Each task invokes check_title_appearance_in_start, which queries the LLM to determine whether the section title begins the supplied page text.

Rather than awaiting each coroutine individually, the function aggregates them into a task list and executes await asyncio.gather(*tasks, return_exceptions=True). This submits all LLM requests simultaneously, letting the network and remote service handle concurrency while the local Python process awaits completion. Results are then mapped back to the original TOC structure via the appear_start field.


# pageindex/page_index.py – concurrent start-check (lines 74-99)

async def check_title_appearance_in_start_concurrent(structure, page_list, model=None, logger=None):
    if logger:
        logger.info("Checking title appearance in start concurrently")
    
    # Skip items without physical_index

    for item in structure:
        if item.get('physical_index') is None:
            item['appear_start'] = 'no'

    tasks = []
    valid_items = []
    for item in structure:
        if item.get('physical_index') is not None:
            page_text = page_list[item['physical_index'] - 1][0]
            tasks.append(check_title_appearance_in_start(
                item['title'], page_text, model=model, logger=logger
            ))
            valid_items.append(item)

    results = await asyncio.gather(*tasks, return_exceptions=True)
    for item, result in zip(valid_items, results):
        if isinstance(result, Exception):
            if logger:
                logger.error(f"Error checking start for {item['title']}: {result}")
            item['appear_start'] = 'no'
        else:
            item['appear_start'] = result

    return structure

Orchestrating Verification with verify_toc

The verify_toc function (lines 892-931 in pageindex/page_index.py) manages higher-level TOC validation. It first determines whether to sample a subset of entries (controlled by parameter N) or verify the entire TOC. For each selected entry with a valid physical_index, it creates a task calling check_title_appearance.

Using asyncio.gather, it runs these checks in parallel, then aggregates results to compute an accuracy percentage and a list of incorrect entries. This metric determines whether the pipeline invokes corrective routines like fix_incorrect_toc_with_retries.


# pageindex/page_index.py – TOC verification (lines 892-931)

async def verify_toc(page_list, list_result, start_index=1, N=None, model=None):
    print('start verify_toc')
    
    # Determine which items to sample

    if N is None:
        sample_indices = range(0, len(list_result))
    else:
        sample_indices = random.sample(range(0, len(list_result)), N)

    indexed_sample_list = []
    for idx in sample_indices:
        item = list_result[idx]
        if item.get('physical_index') is not None:
            item_with_index = item.copy()
            item_with_index['list_index'] = idx
            indexed_sample_list.append(item_with_index)

    # Run all checks in parallel

    tasks = [
        check_title_appearance(item, page_list, start_index, model)
        for item in indexed_sample_list
    ]
    results = await asyncio.gather(*tasks)

    # Summarise outcome

    correct_count = 0
    incorrect_results = []
    for result in results:
        if result['answer'] == 'yes':
            correct_count += 1
        else:
            incorrect_results.append(result)

    accuracy = correct_count / len(results) if results else 0
    print(f"accuracy: {accuracy*100:.2f}%")
    return accuracy, incorrect_results

Under the Hood: Async LLM Calls

The actual HTTP requests to the OpenAI API are wrapped in ChatGPT_API_async within pageindex/utils.py (lines 89-100). This low-level async wrapper accepts prompts and model configurations (defaulting to gpt-4o-2024-11-20 as specified in pageindex/config.yaml), then returns responses that the verification functions parse for "yes" or "no" answers regarding title appearance.

Because the LLM calls dominate the runtime, the asynchronous design lets the program submit many HTTP requests simultaneously, letting the network and the remote service handle the concurrency while the local Python process simply awaits their completion.

Complete Working Example

The following example demonstrates how to invoke the concurrent verification pipeline manually:

import asyncio
from pageindex.page_index import (
    check_title_appearance_in_start_concurrent,
    verify_toc,
    page_index
)
from pageindex.utils import get_page_tokens, structure_to_list

# Step 1: Extract TOC from PDF (synchronous entry point)

toc_tree = page_index("my-report.pdf")

# Step 2: Flatten the hierarchical TOC structure

flat_toc = structure_to_list(toc_tree)

# Step 3: Get token-aware page list for verification

page_list = get_page_tokens("my-report.pdf")

# Step 4: Run concurrent start-check on all entries

async def run_concurrent_verification():
    await check_title_appearance_in_start_concurrent(
        flat_toc, 
        page_list,
        model="gpt-4o-2024-11-20"
    )

# Step 5: Verify a random sample of 20 entries

async def run_sample_verification():
    accuracy, incorrect = await verify_toc(
        page_list, 
        flat_toc,
        start_index=1, 
        N=20,
        model="gpt-4o-2024-11-20"
    )
    print(f"Sample accuracy: {accuracy:.2%}")
    print(f"Incorrect entries: {incorrect}")

# Execute async workflows

asyncio.run(run_concurrent_verification())
asyncio.run(run_sample_verification())

Summary

  • Asynchronous processing in PageIndex leverages Python's asyncio framework to parallelize LLM-driven verification of TOC entries, eliminating the bottleneck of sequential network requests.
  • The check_title_appearance_in_start_concurrent function in pageindex/page_index.py (lines 74-99) builds coroutine tasks for all valid entries and executes them simultaneously using asyncio.gather with return_exceptions=True for fault tolerance.
  • The verify_toc function (lines 892-931) orchestrates sampling and parallel validation, returning accuracy metrics that determine whether corrective passes like fix_incorrect_toc_with_retries are invoked.
  • Low-level async LLM calls are handled by ChatGPT_API_async in pageindex/utils.py (lines 89-100), configured via pageindex/config.yaml to use models like gpt-4o-2024-11-20 by default.

Frequently Asked Questions

Why does PageIndex use asyncio instead of threading for concurrent verification?

PageIndex uses asyncio rather than threading because the bottleneck is I/O-bound network latency when calling the OpenAI API, not CPU processing. Asyncio efficiently manages thousands of concurrent network connections within a single thread, avoiding the overhead and complexity of thread synchronization while maximizing throughput for HTTP requests.

How does PageIndex handle failures in individual verification tasks?

The check_title_appearance_in_start_concurrent function passes return_exceptions=True to asyncio.gather, which captures exceptions rather than halting the entire batch. When iterating through results, the code checks isinstance(result, Exception) and logs the error while defaulting the entry's appear_start field to 'no', allowing the pipeline to continue processing remaining entries.

What LLM model does PageIndex use for asynchronous title verification?

According to the source code configuration in pageindex/config.yaml, PageIndex defaults to gpt-4o-2024-11-20 for verification tasks. The ChatGPT_API_async function in pageindex/utils.py accepts a model parameter that propagates through check_title_appearance_in_start_concurrent and verify_toc, allowing users to override the default model when invoking these functions.

Can I adjust the number of concurrent verification tasks in PageIndex?

While the code does not expose a semaphore or explicit concurrency limit in the high-level API, the degree of parallelism is effectively controlled by the batch size passed to asyncio.gather. The verify_toc function allows you to limit the sample size via the N parameter, which indirectly controls how many concurrent LLM calls are made. For full TOC verification, all entries with valid physical_index values are gathered into a single batch, relying on Python's async event loop and the underlying HTTP connection pooling to manage resource utilization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →