How PageIndex Handles Three Table of Contents Processing Modes

PageIndex automatically selects one of three table of contents processing modes—process_toc_with_page_numbers, process_toc_no_page_numbers, or process_no_toc—based on whether the PDF contains a detectable TOC and whether that TOC includes page numbers.

The VectifyAI/PageIndex library adapts its extraction strategy to the structural quality of your PDF documents. Understanding these three table of contents processing modes helps you predict how the system will index your documents and troubleshoot extraction accuracy.

The Three Table of Contents Processing Modes

PageIndex implements distinct pipelines in pageindex/page_index.py to handle varying levels of TOC completeness. Each mode corresponds to a specific function ranging from lines 68 to 144.

Mode 1: process_toc_with_page_numbers

This mode activates when detect_page_index identifies that the TOC already contains explicit page numbers (e.g., "1 Introduction … 3"). According to the source code in pageindex/page_index.py (lines 114-144), the pipeline executes five sequential steps:

  1. Transform the raw TOC into structured JSON using toc_transformer
  2. Strip existing page numbers via remove_page_number to isolate titles
  3. Extract physical indices from a small window of pages following the TOC using toc_index_extractor
  4. Align extracted indices with original page numbers through extract_matching_page_pairs, calculate_page_offset, and add_page_offset_to_toc_json
  5. Fill missing entries using process_none_page_numbers

Mode 2: process_toc_no_page_numbers

When detect_page_index returns "no" but a TOC page exists without page references, PageIndex falls back to process_toc_no_page_numbers (lines 89-109). This mode handles TOCs that list chapter titles but omit pagination:

  1. Convert the raw TOC to JSON using toc_transformer
  2. Iterate over the entire document grouped into token-size chunks
  3. Assign page numbers via LLM calls through add_page_number_to_toc, which returns <physical_index_X> strings
  4. Convert these string tokens to integers using convert_physical_index_to_int

Mode 3: process_no_toc

If check_toc fails to detect any TOC page, the system triggers process_no_toc (lines 68-87) to generate a hierarchical structure from scratch:

  1. Split the entire PDF into token-size groups
  2. Initialize the TOC by prompting the LLM to create a hierarchy for the first group via generate_toc_init
  3. Continue the tree for subsequent groups using generate_toc_continue
  4. Convert the generated <physical_index_X> strings to integers

How Mode Selection Works

The orchestration logic resides in meta_processor at lines 951-967 of pageindex/page_index.py. After check_toc completes its initial assessment, the function executes a conditional branch:

if mode == 'process_toc_with_page_numbers':
    toc_with_page_number = process_toc_with_page_numbers(...)
elif mode == 'process_toc_no_page_numbers':
    toc_with_page_number = process_toc_no_page_numbers(...)
else:  # process_no_toc

    toc_with_page_number = process_no_toc(...)

This automatic selection requires no user intervention. The system evaluates the PDF structure and selects the appropriate table of contents processing mode to maximize extraction accuracy.

Practical Code Examples

Basic Usage with Automatic Mode Selection

The standard API handles mode selection transparently. The page_index function in pageindex/page_index.py loads config.yaml, constructs the opt object, and invokes page_index_main, which triggers tree_parser and subsequently check_toc:

from pageindex.page_index import page_index

# Path to a PDF (replace with your own file)

pdf_path = "tests/pdfs/2023-annual-report.pdf"

# Run the full pipeline with default configuration (model, token limits, etc.)

result = page_index(pdf_path)

print("Document name:", result["doc_name"])
print("Extracted TOC:")
for item in result["structure"]:
    print(f"{item['structure']}  {item['title']}  (page {item['physical_index']})")

Forcing a Specific Mode via Configuration

While the public API does not expose the mode parameter directly, you can influence the selection by adjusting TOC detection parameters in config.yaml. To force process_no_toc by preventing TOC detection:

from pageindex.page_index import page_index_main, ConfigLoader
from pageindex.utils import JsonLogger

# Load a custom config that disables TOC detection (set a very low toc_check_page_num)

custom_opt = ConfigLoader().load({
    "toc_check_page_num": 0,          # forces find_toc_pages to return []

    "model": "gpt-4o-2024-11-20"
})

# Run the low‑level entry point, bypassing the automatic check

with open("tests/pdfs/2023-annual-report.pdf", "rb") as f:
    result = page_index_main(f, opt=custom_opt)

print(result["structure"])

Setting toc_check_page_num to 0 ensures find_toc_pages returns an empty list, causing meta_processor to execute the process_no_toc branch.

Inspecting the Selected Mode

Enable Python logging to observe which table of contents processing mode the system selects:

import logging
from pageindex.page_index import page_index

logging.basicConfig(level=logging.INFO)

pdf_path = "tests/pdfs/four-lectures.pdf"
result = page_index(pdf_path)

# The logger prints something like:

#   INFO:root:process_toc_no_page_numbers

#   INFO:root:accuracy: 98.75%

The meta_processor function explicitly prints the mode variable at execution start, allowing you to verify the automatic selection logic.

Key Implementation Files

Understanding the file structure helps navigate the three table of contents processing modes:

File Role Relevant Excerpts
pageindex/page_index.py Core orchestration – defines the three processing functions, TOC detection, and the meta_processor that selects the mode. process_no_toc (lines 68-87), process_toc_no_page_numbers (lines 89-109), process_toc_with_page_numbers (lines 114-144), meta_processor (lines 951-967)
pageindex/utils.py Helper functions for token counting, LLM calls, and low-level PDF → token conversion used by all three modes. get_page_tokens, ChatGPT_API, count_tokens
pageindex/config.yaml Default configuration (model, page-check limits, node-size thresholds). model, toc_check_page_num, max_page_num_each_node
run_pageindex.py CLI entry point – parses arguments and forwards to page_index. Shows how the library is normally executed. Argument handling for --toc-check-pages, --model
README.md High-level description of the library and sample commands. Usage examples, installation steps

These files collectively implement the adaptive TOC extraction strategy that gracefully degrades from full-featured page-indexed TOC → page-less TOC → LLM-generated TOC.

Summary

  • Automatic mode selection occurs in meta_processor (lines 951-967 of pageindex/page_index.py) based on TOC detection results, requiring no user configuration.
  • process_toc_with_page_numbers handles documents with explicit page numbers in the TOC, using alignment algorithms to map physical indices to logical page numbers.
  • process_toc_no_page_numbers processes TOCs lacking pagination by querying an LLM across document chunks to assign physical indices.
  • process_no_toc generates a complete hierarchical TOC from scratch when no TOC page exists, using iterative LLM calls to build the document structure.
  • Configuration control allows advanced users to force specific modes by adjusting toc_check_page_num in config.yaml to manipulate detection behavior.

Frequently Asked Questions

How does PageIndex decide which table of contents processing mode to use?

PageIndex automatically determines the appropriate mode through the check_toc and detect_page_index functions in pageindex/page_index.py. If the system detects a TOC containing page numbers, it selects process_toc_with_page_numbers. If it finds a TOC without page numbers, it chooses process_toc_no_page_numbers. If no TOC is detected at all, it defaults to process_no_toc. This decision happens inside meta_processor without requiring user intervention.

Can I force PageIndex to use a specific processing mode?

While the public API does not expose a direct mode parameter, you can influence the selection by modifying configuration values that affect TOC detection. For example, setting toc_check_page_num to 0 in your configuration forces find_toc_pages to return an empty list, which causes meta_processor to execute the process_no_toc branch. Similarly, adjusting detection thresholds can bias the system toward treating detected TOCs as having or lacking page numbers.

What happens when PageIndex cannot find any table of contents in the document?

When check_toc fails to identify a TOC page, PageIndex invokes process_no_toc (lines 68-87 in pageindex/page_index.py). This function splits the entire PDF into token-size groups and uses an LLM to generate a hierarchical TOC from scratch. It first initializes the structure with generate_toc_init for the first group, then continues building the tree with generate_toc_continue for subsequent groups, effectively creating a document outline where none existed.

Which processing mode provides the highest accuracy for page numbering?

The process_toc_with_page_numbers mode typically yields the highest accuracy because it uses algorithmic alignment rather than pure LLM inference. This mode extracts physical indices from the document content using toc_index_extractor, then calculates the offset between these physical indices and the logical page numbers found in the original TOC via calculate_page_offset. By aligning these two data sources and filling gaps with process_none_page_numbers, it achieves precise pagination even when the PDF contains complex formatting or offset page numbering.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →