# How Asynchronous Processing in PageIndex Enables Concurrent Title Verification

> Learn how PageIndex uses Python asyncio for asynchronous processing to achieve concurrent title verification, speeding up OpenAI API calls and reducing processing time.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: internals
- Published: 2026-02-16

---

**PageIndex leverages Python's `asyncio` framework to parallelize LLM-driven verification of table-of-contents entries, dramatically reducing wall-clock time by submitting multiple HTTP requests to the OpenAI API simultaneously rather than processing them sequentially.**

When extracting structured data from large PDF documents, verifying that each table-of-contents (TOC) entry actually appears at the start of its claimed page can become a significant bottleneck. The VectifyAI/PageIndex repository solves this scalability challenge by implementing **asynchronous processing in PageIndex** to run concurrent title verification against large language models (LLMs), enabling efficient processing of documents containing hundreds of sections.

## The Challenge: Verifying Hundreds of TOC Entries

PageIndex extracts a structured TOC from PDFs, mapping each entry to a physical page index via the `physical_index` field. Before trusting these mappings, the system must validate that the section title text actually appears at the beginning of the corresponding page content. Performing these checks sequentially for documents with extensive hierarchies would create prohibitive latency due to network round-trips to the LLM API, making **concurrent title verification** essential for performance.

## Core Async Architecture

The concurrent verification logic centers on two primary async functions in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), both utilizing `asyncio.gather` to parallelize LLM calls.

### Building Concurrent Tasks with `check_title_appearance_in_start_concurrent`

Located at lines 74-99 in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), this function constructs a list of coroutine tasks—one for every TOC entry that possesses a valid `physical_index`. Each task invokes `check_title_appearance_in_start`, which queries the LLM to determine whether the section title begins the supplied page text.

Rather than awaiting each coroutine individually, the function aggregates them into a task list and executes `await asyncio.gather(*tasks, return_exceptions=True)`. This submits all LLM requests simultaneously, letting the network and remote service handle concurrency while the local Python process awaits completion. Results are then mapped back to the original TOC structure via the `appear_start` field.

```python

# pageindex/page_index.py – concurrent start-check (lines 74-99)

async def check_title_appearance_in_start_concurrent(structure, page_list, model=None, logger=None):
    if logger:
        logger.info("Checking title appearance in start concurrently")
    
    # Skip items without physical_index

    for item in structure:
        if item.get('physical_index') is None:
            item['appear_start'] = 'no'

    tasks = []
    valid_items = []
    for item in structure:
        if item.get('physical_index') is not None:
            page_text = page_list[item['physical_index'] - 1][0]
            tasks.append(check_title_appearance_in_start(
                item['title'], page_text, model=model, logger=logger
            ))
            valid_items.append(item)

    results = await asyncio.gather(*tasks, return_exceptions=True)
    for item, result in zip(valid_items, results):
        if isinstance(result, Exception):
            if logger:
                logger.error(f"Error checking start for {item['title']}: {result}")
            item['appear_start'] = 'no'
        else:
            item['appear_start'] = result

    return structure

```

### Orchestrating Verification with `verify_toc`

The `verify_toc` function (lines 892-931 in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)) manages higher-level TOC validation. It first determines whether to sample a subset of entries (controlled by parameter `N`) or verify the entire TOC. For each selected entry with a valid `physical_index`, it creates a task calling `check_title_appearance`.

Using `asyncio.gather`, it runs these checks in parallel, then aggregates results to compute an accuracy percentage and a list of incorrect entries. This metric determines whether the pipeline invokes corrective routines like `fix_incorrect_toc_with_retries`.

```python

# pageindex/page_index.py – TOC verification (lines 892-931)

async def verify_toc(page_list, list_result, start_index=1, N=None, model=None):
    print('start verify_toc')
    
    # Determine which items to sample

    if N is None:
        sample_indices = range(0, len(list_result))
    else:
        sample_indices = random.sample(range(0, len(list_result)), N)

    indexed_sample_list = []
    for idx in sample_indices:
        item = list_result[idx]
        if item.get('physical_index') is not None:
            item_with_index = item.copy()
            item_with_index['list_index'] = idx
            indexed_sample_list.append(item_with_index)

    # Run all checks in parallel

    tasks = [
        check_title_appearance(item, page_list, start_index, model)
        for item in indexed_sample_list
    ]
    results = await asyncio.gather(*tasks)

    # Summarise outcome

    correct_count = 0
    incorrect_results = []
    for result in results:
        if result['answer'] == 'yes':
            correct_count += 1
        else:
            incorrect_results.append(result)

    accuracy = correct_count / len(results) if results else 0
    print(f"accuracy: {accuracy*100:.2f}%")
    return accuracy, incorrect_results

```

## Under the Hood: Async LLM Calls

The actual HTTP requests to the OpenAI API are wrapped in `ChatGPT_API_async` within [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 89-100). This low-level async wrapper accepts prompts and model configurations (defaulting to `gpt-4o-2024-11-20` as specified in [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml)), then returns responses that the verification functions parse for "yes" or "no" answers regarding title appearance.

Because the LLM calls dominate the runtime, the asynchronous design lets the program submit many HTTP requests simultaneously, letting the network and the remote service handle the concurrency while the local Python process simply awaits their completion.

## Complete Working Example

The following example demonstrates how to invoke the concurrent verification pipeline manually:

```python
import asyncio
from pageindex.page_index import (
    check_title_appearance_in_start_concurrent,
    verify_toc,
    page_index
)
from pageindex.utils import get_page_tokens, structure_to_list

# Step 1: Extract TOC from PDF (synchronous entry point)

toc_tree = page_index("my-report.pdf")

# Step 2: Flatten the hierarchical TOC structure

flat_toc = structure_to_list(toc_tree)

# Step 3: Get token-aware page list for verification

page_list = get_page_tokens("my-report.pdf")

# Step 4: Run concurrent start-check on all entries

async def run_concurrent_verification():
    await check_title_appearance_in_start_concurrent(
        flat_toc, 
        page_list,
        model="gpt-4o-2024-11-20"
    )

# Step 5: Verify a random sample of 20 entries

async def run_sample_verification():
    accuracy, incorrect = await verify_toc(
        page_list, 
        flat_toc,
        start_index=1, 
        N=20,
        model="gpt-4o-2024-11-20"
    )
    print(f"Sample accuracy: {accuracy:.2%}")
    print(f"Incorrect entries: {incorrect}")

# Execute async workflows

asyncio.run(run_concurrent_verification())
asyncio.run(run_sample_verification())

```

## Summary

- **Asynchronous processing in PageIndex** leverages Python's `asyncio` framework to parallelize LLM-driven verification of TOC entries, eliminating the bottleneck of sequential network requests.
- The `check_title_appearance_in_start_concurrent` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) (lines 74-99) builds coroutine tasks for all valid entries and executes them simultaneously using `asyncio.gather` with `return_exceptions=True` for fault tolerance.
- The `verify_toc` function (lines 892-931) orchestrates sampling and parallel validation, returning accuracy metrics that determine whether corrective passes like `fix_incorrect_toc_with_retries` are invoked.
- Low-level async LLM calls are handled by `ChatGPT_API_async` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 89-100), configured via [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) to use models like `gpt-4o-2024-11-20` by default.

## Frequently Asked Questions

### Why does PageIndex use asyncio instead of threading for concurrent verification?

PageIndex uses `asyncio` rather than threading because the bottleneck is I/O-bound network latency when calling the OpenAI API, not CPU processing. Asyncio efficiently manages thousands of concurrent network connections within a single thread, avoiding the overhead and complexity of thread synchronization while maximizing throughput for HTTP requests.

### How does PageIndex handle failures in individual verification tasks?

The `check_title_appearance_in_start_concurrent` function passes `return_exceptions=True` to `asyncio.gather`, which captures exceptions rather than halting the entire batch. When iterating through results, the code checks `isinstance(result, Exception)` and logs the error while defaulting the entry's `appear_start` field to `'no'`, allowing the pipeline to continue processing remaining entries.

### What LLM model does PageIndex use for asynchronous title verification?

According to the source code configuration in [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml), PageIndex defaults to `gpt-4o-2024-11-20` for verification tasks. The `ChatGPT_API_async` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) accepts a `model` parameter that propagates through `check_title_appearance_in_start_concurrent` and `verify_toc`, allowing users to override the default model when invoking these functions.

### Can I adjust the number of concurrent verification tasks in PageIndex?

While the code does not expose a semaphore or explicit concurrency limit in the high-level API, the degree of parallelism is effectively controlled by the batch size passed to `asyncio.gather`. The `verify_toc` function allows you to limit the sample size via the `N` parameter, which indirectly controls how many concurrent LLM calls are made. For full TOC verification, all entries with valid `physical_index` values are gathered into a single batch, relying on Python's async event loop and the underlying HTTP connection pooling to manage resource utilization.