# How PageIndex Detects and Processes Tables of Contents from Documents

> Learn how PageIndex automatically detects and processes document tables of contents, converting them into hierarchical JSON with precise page indices using an LLM pipeline.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: how-to-guide
- Published: 2026-02-16

---

**PageIndex automatically locates, extracts, and converts document tables of contents into hierarchical JSON trees with accurate physical page indices using a multi-stage LLM-driven pipeline.**

PageIndex is an open-source Python library that analyzes PDF documents to build structured outlines. The system detects and processes tables of contents through a sophisticated orchestration of page scanning, LLM-based extraction, and hierarchical tree construction, ultimately delivering a navigable tree with precise physical page references.

## Stage 1: Scanning Pages for TOC Detection

The detection process begins in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) with the `find_toc_pages` function (lines 332-359). This function iterates through the document page list, feeding each page’s raw text to `toc_detector_single_page` (lines 4-23).

The detector prompts an LLM with the specific question *“Is there a table of contents in this text?”* and returns a binary yes/no response. The `find_toc_pages` scan automatically terminates once the system encounters a run of non-TOC pages, preventing unnecessary processing of the document body.

## Stage 2: Extracting Raw TOC Content

When TOC pages are identified, the `toc_extractor` function (lines 19-30) concatenates the contents of all detected pages. This function submits the combined text to an LLM with instructions to extract the full table of contents while explicitly filtering out non-TOC elements such as abstracts, figure lists, and appendix listings.

## Stage 3: Detecting and Handling Page Numbers

The system checks for the presence of explicit page numbers using `detect_page_index` (lines 99-117). Based on this detection, PageIndex branches into two distinct processing paths to handle both scenarios when it detects and processes tables of contents.

### Processing TOCs Without Page Numbers

If the extracted TOC lacks explicit page numbers, PageIndex invokes `add_page_number_to_toc` (lines 534-584). This function iterates through the PDF pages and queries the LLM to verify whether a specific section title starts on a given page slice. When confirmed, it inserts the `physical_index` for that section.

The utility function `convert_physical_index_to_int` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 445-466) then converts these physical index markers into integer values for consistency.

### Processing TOCs With Existing Page Numbers

When the TOC contains explicit page markers (e.g., `<physical_index_12>`), `toc_index_extractor` (lines 40-67) maps these markers to their corresponding entries. However, these indices often require alignment with the actual document pagination.

The `calculate_page_offset` function (lines 86-105) determines the most common offset between provisional indices and true page numbers. Subsequently, `add_page_offset_to_toc_json` (lines 108-114) applies this offset to all entries, ensuring accurate physical references.

## Stage 4: Verification and Auto-Repair

PageIndex implements a robust verification mechanism through `verify_toc` (lines 891-945). This function samples a subset of generated entries and executes `check_title_appearance`, which performs an LLM-based fuzzy match against the original PDF pages to confirm title presence.

The system computes an accuracy score based on these verifications. If the score falls below the configured threshold, `fix_incorrect_toc` (lines 332-426) re-queries the LLM for the correct physical indices of each failing section, iteratively updating the TOC until accuracy requirements are met.

## Stage 5: Building the Hierarchical Tree

The final transformation occurs in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) through `post_processing` (lines 560-580). This function converts the flat list of TOC entries—each containing `structure`, `title`, and `physical_index`—into a nested tree of section nodes.

The resulting tree includes calculated `start_index` and `end_index` ranges for each node, enabling precise document navigation. Depending on user configuration, the system optionally enriches nodes with unique IDs, text snippets, and AI-generated summaries.

## Orchestration and Public API

The entire workflow is coordinated through several high-level functions in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py). The `check_toc` function (lines 889-905) determines whether a valid TOC exists in the document. Based on this determination, `meta_processor` (lines 951-999) selects the appropriate processing path: `process_toc_with_page_numbers`, `process_toc_no_page_numbers`, or `process_no_toc`.

Finally, `tree_parser` (lines 1021-1055) drives the complete pipeline and returns the finished structure. For end-users, the public API entry point is `page_index` exposed in [`pageindex/__init__.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/__init__.py), which initializes a `ConfigLoader` and invokes `page_index_main` to begin processing.

## Practical Implementation Examples

### High-Level API Usage

The simplest way to detect and process tables of contents is through the `page_index` function:

```python
from pageindex import page_index

# `my_doc.pdf` can be a path or a BytesIO object

result = page_index(
    doc="my_doc.pdf",
    model="gpt-4o-2024-11-20",
    toc_check_page_num=30,          # how many pages to scan for a TOC

    max_page_num_each_node=20,
    max_token_num_each_node=5000,
    if_add_node_id="yes",
    if_add_node_summary="yes",
    if_add_doc_description="yes",
    if_add_node_text="yes",
)

# `result['structure']` is the hierarchical TOC with page numbers

print(result["structure"])

```

This high-level call chains through `page_index_main` → `tree_parser` → the complete TOC detection and processing pipeline.

### Manual Pipeline Invocation

For granular control over specific stages:

```python
from pageindex.page_index import find_toc_pages, toc_extractor, toc_transformer
from pageindex.utils import get_page_tokens

pages = get_page_tokens("my_doc.pdf")                 # [(text, token_len), ...]

toc_pages = find_toc_pages(start_page_index=0, page_list=pages, opt=opt)  # opt contains model, toc_check_page_num, …

raw_toc = toc_extractor(pages, toc_pages, opt.model) # raw text of the TOC

structured = toc_transformer(raw_toc, opt.model)     # JSON list of sections (no page numbers yet)

```

From here you can call `process_toc_no_page_numbers` or `process_toc_with_page_numbers` depending on whether `detect_page_index` reported page numbers in the extracted content.

## Key Source Files and Architecture

| File | Primary responsibilities | Important symbols |
|------|--------------------------|-------------------|
| **[`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)** | Orchestrates TOC detection, extraction, transformation, verification, and tree construction. Holds all LLM‑prompt functions. | `find_toc_pages`, `toc_detector_single_page`, `toc_extractor`, `detect_page_index`, `toc_transformer`, `add_page_number_to_toc`, `toc_index_extractor`, `process_toc_*`, `verify_toc`, `meta_processor`, `tree_parser` |
| **[`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py)** | Low‑level PDF handling, token counting, JSON utilities, tree helpers, logging, and configuration loading. | `get_page_tokens`, `convert_physical_index_to_int`, `post_processing`, `list_to_tree`, `ConfigLoader`, `JsonLogger` |
| **[`pageindex/__init__.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/__init__.py)** | Exposes the public `page_index` function. | `page_index` |
| **[`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml)** | Default configuration values (model name, thresholds, etc.) used by `ConfigLoader`. | — |

These files together implement the end‑to‑end detection and processing of tables of contents in the **PageIndex** repository.

## Summary

- **PageIndex** detects tables of contents by scanning pages with `find_toc_pages` and using an LLM to classify content via `toc_detector_single_page`.
- Raw TOC text is extracted by `toc_extractor` and structured into JSON by `toc_transformer`, creating hierarchical entries with section titles and structural indices.
- The system branches based on `detect_page_index`: TOCs without page numbers are processed by `add_page_number_to_toc`, while those with existing numbers undergo offset correction via `calculate_page_offset` and `add_page_offset_to_toc_json`.
- **Verification** occurs through `verify_toc` and `check_title_appearance`, with automatic repair handled by `fix_incorrect_toc` if accuracy falls below thresholds.
- Final tree construction happens in `post_processing` within [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py), converting flat lists into nested structures with `start_index` and `end_index` ranges.
- The entire workflow is orchestrated by `meta_processor` and `tree_parser`, exposed through the public `page_index` API.

## Frequently Asked Questions

### How does PageIndex determine if a page contains a table of contents?

PageIndex uses the `toc_detector_single_page` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) to analyze individual pages. This function prompts an LLM with the specific question *“Is there a table of contents in this text?”* and returns a binary yes/no response. The `find_toc_pages` function orchestrates this scan across the document, stopping when it encounters a run of non-TOC pages to prevent unnecessary processing of the document body.

### What happens if the extracted table of contents does not contain page numbers?

If `detect_page_index` determines that the TOC lacks explicit page numbers, PageIndex invokes `add_page_number_to_toc` (lines 534-584). This function iterates through the PDF pages and queries the LLM to verify whether a specific section title starts on a given page slice. When confirmed, it inserts the `physical_index` for that section. The utility function `convert_physical_index_to_int` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 445-466) then converts these markers into integer values for consistency.

### How does PageIndex verify the accuracy of the generated table of contents?

PageIndex implements a robust verification mechanism through `verify_toc` (lines 891-945). This function samples a subset of generated entries and executes `check_title_appearance`, which performs an LLM-based fuzzy match against the original PDF pages to confirm title presence. The system computes an accuracy score based on these verifications. If the score falls below the configured threshold, `fix_incorrect_toc` (lines 332-426) re-queries the LLM for the correct physical indices of each failing section, iteratively updating the TOC until accuracy requirements are met.

### Can I manually invoke specific stages of the TOC detection pipeline?

Yes, PageIndex exposes individual pipeline stages for granular control. You can manually invoke `find_toc_pages` to detect TOC locations, `toc_extractor` to pull raw text, and `toc_transformer` to generate structured JSON. After these steps, you can branch to either `process_toc_no_page_numbers` or `process_toc_with_page_numbers` depending on the output of `detect_page_index`. These functions are available in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) and can be imported alongside utilities like `get_page_tokens` from [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py).