How PageIndex Detects and Processes Tables of Contents from Documents

PageIndex automatically locates, extracts, and converts document tables of contents into hierarchical JSON trees with accurate physical page indices using a multi-stage LLM-driven pipeline.

PageIndex is an open-source Python library that analyzes PDF documents to build structured outlines. The system detects and processes tables of contents through a sophisticated orchestration of page scanning, LLM-based extraction, and hierarchical tree construction, ultimately delivering a navigable tree with precise physical page references.

Stage 1: Scanning Pages for TOC Detection

The detection process begins in pageindex/page_index.py with the find_toc_pages function (lines 332-359). This function iterates through the document page list, feeding each page’s raw text to toc_detector_single_page (lines 4-23).

The detector prompts an LLM with the specific question “Is there a table of contents in this text?” and returns a binary yes/no response. The find_toc_pages scan automatically terminates once the system encounters a run of non-TOC pages, preventing unnecessary processing of the document body.

Stage 2: Extracting Raw TOC Content

When TOC pages are identified, the toc_extractor function (lines 19-30) concatenates the contents of all detected pages. This function submits the combined text to an LLM with instructions to extract the full table of contents while explicitly filtering out non-TOC elements such as abstracts, figure lists, and appendix listings.

Stage 3: Detecting and Handling Page Numbers

The system checks for the presence of explicit page numbers using detect_page_index (lines 99-117). Based on this detection, PageIndex branches into two distinct processing paths to handle both scenarios when it detects and processes tables of contents.

Processing TOCs Without Page Numbers

If the extracted TOC lacks explicit page numbers, PageIndex invokes add_page_number_to_toc (lines 534-584). This function iterates through the PDF pages and queries the LLM to verify whether a specific section title starts on a given page slice. When confirmed, it inserts the physical_index for that section.

The utility function convert_physical_index_to_int in pageindex/utils.py (lines 445-466) then converts these physical index markers into integer values for consistency.

Processing TOCs With Existing Page Numbers

When the TOC contains explicit page markers (e.g., <physical_index_12>), toc_index_extractor (lines 40-67) maps these markers to their corresponding entries. However, these indices often require alignment with the actual document pagination.

The calculate_page_offset function (lines 86-105) determines the most common offset between provisional indices and true page numbers. Subsequently, add_page_offset_to_toc_json (lines 108-114) applies this offset to all entries, ensuring accurate physical references.

Stage 4: Verification and Auto-Repair

PageIndex implements a robust verification mechanism through verify_toc (lines 891-945). This function samples a subset of generated entries and executes check_title_appearance, which performs an LLM-based fuzzy match against the original PDF pages to confirm title presence.

The system computes an accuracy score based on these verifications. If the score falls below the configured threshold, fix_incorrect_toc (lines 332-426) re-queries the LLM for the correct physical indices of each failing section, iteratively updating the TOC until accuracy requirements are met.

Stage 5: Building the Hierarchical Tree

The final transformation occurs in pageindex/utils.py through post_processing (lines 560-580). This function converts the flat list of TOC entries—each containing structure, title, and physical_index—into a nested tree of section nodes.

The resulting tree includes calculated start_index and end_index ranges for each node, enabling precise document navigation. Depending on user configuration, the system optionally enriches nodes with unique IDs, text snippets, and AI-generated summaries.

Orchestration and Public API

The entire workflow is coordinated through several high-level functions in pageindex/page_index.py. The check_toc function (lines 889-905) determines whether a valid TOC exists in the document. Based on this determination, meta_processor (lines 951-999) selects the appropriate processing path: process_toc_with_page_numbers, process_toc_no_page_numbers, or process_no_toc.

Finally, tree_parser (lines 1021-1055) drives the complete pipeline and returns the finished structure. For end-users, the public API entry point is page_index exposed in pageindex/__init__.py, which initializes a ConfigLoader and invokes page_index_main to begin processing.

Practical Implementation Examples

High-Level API Usage

The simplest way to detect and process tables of contents is through the page_index function:

from pageindex import page_index

# `my_doc.pdf` can be a path or a BytesIO object

result = page_index(
    doc="my_doc.pdf",
    model="gpt-4o-2024-11-20",
    toc_check_page_num=30,          # how many pages to scan for a TOC

    max_page_num_each_node=20,
    max_token_num_each_node=5000,
    if_add_node_id="yes",
    if_add_node_summary="yes",
    if_add_doc_description="yes",
    if_add_node_text="yes",
)

# `result['structure']` is the hierarchical TOC with page numbers

print(result["structure"])

This high-level call chains through page_index_main → tree_parser → the complete TOC detection and processing pipeline.

Manual Pipeline Invocation

For granular control over specific stages:

from pageindex.page_index import find_toc_pages, toc_extractor, toc_transformer
from pageindex.utils import get_page_tokens

pages = get_page_tokens("my_doc.pdf")                 # [(text, token_len), ...]

toc_pages = find_toc_pages(start_page_index=0, page_list=pages, opt=opt)  # opt contains model, toc_check_page_num, …

raw_toc = toc_extractor(pages, toc_pages, opt.model) # raw text of the TOC

structured = toc_transformer(raw_toc, opt.model)     # JSON list of sections (no page numbers yet)

From here you can call process_toc_no_page_numbers or process_toc_with_page_numbers depending on whether detect_page_index reported page numbers in the extracted content.

Key Source Files and Architecture

File Primary responsibilities Important symbols
pageindex/page_index.py Orchestrates TOC detection, extraction, transformation, verification, and tree construction. Holds all LLM‑prompt functions. find_toc_pages, toc_detector_single_page, toc_extractor, detect_page_index, toc_transformer, add_page_number_to_toc, toc_index_extractor, process_toc_*, verify_toc, meta_processor, tree_parser
pageindex/utils.py Low‑level PDF handling, token counting, JSON utilities, tree helpers, logging, and configuration loading. get_page_tokens, convert_physical_index_to_int, post_processing, list_to_tree, ConfigLoader, JsonLogger
pageindex/__init__.py Exposes the public page_index function. page_index
pageindex/config.yaml Default configuration values (model name, thresholds, etc.) used by ConfigLoader. —

These files together implement the end‑to‑end detection and processing of tables of contents in the PageIndex repository.

Summary

  • PageIndex detects tables of contents by scanning pages with find_toc_pages and using an LLM to classify content via toc_detector_single_page.
  • Raw TOC text is extracted by toc_extractor and structured into JSON by toc_transformer, creating hierarchical entries with section titles and structural indices.
  • The system branches based on detect_page_index: TOCs without page numbers are processed by add_page_number_to_toc, while those with existing numbers undergo offset correction via calculate_page_offset and add_page_offset_to_toc_json.
  • Verification occurs through verify_toc and check_title_appearance, with automatic repair handled by fix_incorrect_toc if accuracy falls below thresholds.
  • Final tree construction happens in post_processing within utils.py, converting flat lists into nested structures with start_index and end_index ranges.
  • The entire workflow is orchestrated by meta_processor and tree_parser, exposed through the public page_index API.

Frequently Asked Questions

How does PageIndex determine if a page contains a table of contents?

PageIndex uses the toc_detector_single_page function in pageindex/page_index.py to analyze individual pages. This function prompts an LLM with the specific question “Is there a table of contents in this text?” and returns a binary yes/no response. The find_toc_pages function orchestrates this scan across the document, stopping when it encounters a run of non-TOC pages to prevent unnecessary processing of the document body.

What happens if the extracted table of contents does not contain page numbers?

If detect_page_index determines that the TOC lacks explicit page numbers, PageIndex invokes add_page_number_to_toc (lines 534-584). This function iterates through the PDF pages and queries the LLM to verify whether a specific section title starts on a given page slice. When confirmed, it inserts the physical_index for that section. The utility function convert_physical_index_to_int in pageindex/utils.py (lines 445-466) then converts these markers into integer values for consistency.

How does PageIndex verify the accuracy of the generated table of contents?

PageIndex implements a robust verification mechanism through verify_toc (lines 891-945). This function samples a subset of generated entries and executes check_title_appearance, which performs an LLM-based fuzzy match against the original PDF pages to confirm title presence. The system computes an accuracy score based on these verifications. If the score falls below the configured threshold, fix_incorrect_toc (lines 332-426) re-queries the LLM for the correct physical indices of each failing section, iteratively updating the TOC until accuracy requirements are met.

Can I manually invoke specific stages of the TOC detection pipeline?

Yes, PageIndex exposes individual pipeline stages for granular control. You can manually invoke find_toc_pages to detect TOC locations, toc_extractor to pull raw text, and toc_transformer to generate structured JSON. After these steps, you can branch to either process_toc_no_page_numbers or process_toc_with_page_numbers depending on the output of detect_page_index. These functions are available in pageindex/page_index.py and can be imported alongside utilities like get_page_tokens from pageindex/utils.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →