How PageIndex Handles Three Table of Contents Processing Modes
PageIndex automatically selects one of three table of contents processing modes—process_toc_with_page_numbers, process_toc_no_page_numbers, or process_no_toc—based on whether the PDF contains a detectable TOC and whether that TOC includes page numbers.
The VectifyAI/PageIndex library adapts its extraction strategy to the structural quality of your PDF documents. Understanding these three table of contents processing modes helps you predict how the system will index your documents and troubleshoot extraction accuracy.
The Three Table of Contents Processing Modes
PageIndex implements distinct pipelines in pageindex/page_index.py to handle varying levels of TOC completeness. Each mode corresponds to a specific function ranging from lines 68 to 144.
Mode 1: process_toc_with_page_numbers
This mode activates when detect_page_index identifies that the TOC already contains explicit page numbers (e.g., "1 Introduction … 3"). According to the source code in pageindex/page_index.py (lines 114-144), the pipeline executes five sequential steps:
- Transform the raw TOC into structured JSON using
toc_transformer - Strip existing page numbers via
remove_page_numberto isolate titles - Extract physical indices from a small window of pages following the TOC using
toc_index_extractor - Align extracted indices with original page numbers through
extract_matching_page_pairs,calculate_page_offset, andadd_page_offset_to_toc_json - Fill missing entries using
process_none_page_numbers
Mode 2: process_toc_no_page_numbers
When detect_page_index returns "no" but a TOC page exists without page references, PageIndex falls back to process_toc_no_page_numbers (lines 89-109). This mode handles TOCs that list chapter titles but omit pagination:
- Convert the raw TOC to JSON using
toc_transformer - Iterate over the entire document grouped into token-size chunks
- Assign page numbers via LLM calls through
add_page_number_to_toc, which returns<physical_index_X>strings - Convert these string tokens to integers using
convert_physical_index_to_int
Mode 3: process_no_toc
If check_toc fails to detect any TOC page, the system triggers process_no_toc (lines 68-87) to generate a hierarchical structure from scratch:
- Split the entire PDF into token-size groups
- Initialize the TOC by prompting the LLM to create a hierarchy for the first group via
generate_toc_init - Continue the tree for subsequent groups using
generate_toc_continue - Convert the generated
<physical_index_X>strings to integers
How Mode Selection Works
The orchestration logic resides in meta_processor at lines 951-967 of pageindex/page_index.py. After check_toc completes its initial assessment, the function executes a conditional branch:
if mode == 'process_toc_with_page_numbers':
toc_with_page_number = process_toc_with_page_numbers(...)
elif mode == 'process_toc_no_page_numbers':
toc_with_page_number = process_toc_no_page_numbers(...)
else: # process_no_toc
toc_with_page_number = process_no_toc(...)
This automatic selection requires no user intervention. The system evaluates the PDF structure and selects the appropriate table of contents processing mode to maximize extraction accuracy.
Practical Code Examples
Basic Usage with Automatic Mode Selection
The standard API handles mode selection transparently. The page_index function in pageindex/page_index.py loads config.yaml, constructs the opt object, and invokes page_index_main, which triggers tree_parser and subsequently check_toc:
from pageindex.page_index import page_index
# Path to a PDF (replace with your own file)
pdf_path = "tests/pdfs/2023-annual-report.pdf"
# Run the full pipeline with default configuration (model, token limits, etc.)
result = page_index(pdf_path)
print("Document name:", result["doc_name"])
print("Extracted TOC:")
for item in result["structure"]:
print(f"{item['structure']} {item['title']} (page {item['physical_index']})")
Forcing a Specific Mode via Configuration
While the public API does not expose the mode parameter directly, you can influence the selection by adjusting TOC detection parameters in config.yaml. To force process_no_toc by preventing TOC detection:
from pageindex.page_index import page_index_main, ConfigLoader
from pageindex.utils import JsonLogger
# Load a custom config that disables TOC detection (set a very low toc_check_page_num)
custom_opt = ConfigLoader().load({
"toc_check_page_num": 0, # forces find_toc_pages to return []
"model": "gpt-4o-2024-11-20"
})
# Run the low‑level entry point, bypassing the automatic check
with open("tests/pdfs/2023-annual-report.pdf", "rb") as f:
result = page_index_main(f, opt=custom_opt)
print(result["structure"])
Setting toc_check_page_num to 0 ensures find_toc_pages returns an empty list, causing meta_processor to execute the process_no_toc branch.
Inspecting the Selected Mode
Enable Python logging to observe which table of contents processing mode the system selects:
import logging
from pageindex.page_index import page_index
logging.basicConfig(level=logging.INFO)
pdf_path = "tests/pdfs/four-lectures.pdf"
result = page_index(pdf_path)
# The logger prints something like:
# INFO:root:process_toc_no_page_numbers
# INFO:root:accuracy: 98.75%
The meta_processor function explicitly prints the mode variable at execution start, allowing you to verify the automatic selection logic.
Key Implementation Files
Understanding the file structure helps navigate the three table of contents processing modes:
| File | Role | Relevant Excerpts |
|---|---|---|
pageindex/page_index.py |
Core orchestration – defines the three processing functions, TOC detection, and the meta_processor that selects the mode. |
process_no_toc (lines 68-87), process_toc_no_page_numbers (lines 89-109), process_toc_with_page_numbers (lines 114-144), meta_processor (lines 951-967) |
pageindex/utils.py |
Helper functions for token counting, LLM calls, and low-level PDF → token conversion used by all three modes. | get_page_tokens, ChatGPT_API, count_tokens |
pageindex/config.yaml |
Default configuration (model, page-check limits, node-size thresholds). | model, toc_check_page_num, max_page_num_each_node |
run_pageindex.py |
CLI entry point – parses arguments and forwards to page_index. Shows how the library is normally executed. |
Argument handling for --toc-check-pages, --model |
README.md |
High-level description of the library and sample commands. | Usage examples, installation steps |
These files collectively implement the adaptive TOC extraction strategy that gracefully degrades from full-featured page-indexed TOC → page-less TOC → LLM-generated TOC.
Summary
- Automatic mode selection occurs in
meta_processor(lines 951-967 ofpageindex/page_index.py) based on TOC detection results, requiring no user configuration. process_toc_with_page_numbershandles documents with explicit page numbers in the TOC, using alignment algorithms to map physical indices to logical page numbers.process_toc_no_page_numbersprocesses TOCs lacking pagination by querying an LLM across document chunks to assign physical indices.process_no_tocgenerates a complete hierarchical TOC from scratch when no TOC page exists, using iterative LLM calls to build the document structure.- Configuration control allows advanced users to force specific modes by adjusting
toc_check_page_numinconfig.yamlto manipulate detection behavior.
Frequently Asked Questions
How does PageIndex decide which table of contents processing mode to use?
PageIndex automatically determines the appropriate mode through the check_toc and detect_page_index functions in pageindex/page_index.py. If the system detects a TOC containing page numbers, it selects process_toc_with_page_numbers. If it finds a TOC without page numbers, it chooses process_toc_no_page_numbers. If no TOC is detected at all, it defaults to process_no_toc. This decision happens inside meta_processor without requiring user intervention.
Can I force PageIndex to use a specific processing mode?
While the public API does not expose a direct mode parameter, you can influence the selection by modifying configuration values that affect TOC detection. For example, setting toc_check_page_num to 0 in your configuration forces find_toc_pages to return an empty list, which causes meta_processor to execute the process_no_toc branch. Similarly, adjusting detection thresholds can bias the system toward treating detected TOCs as having or lacking page numbers.
What happens when PageIndex cannot find any table of contents in the document?
When check_toc fails to identify a TOC page, PageIndex invokes process_no_toc (lines 68-87 in pageindex/page_index.py). This function splits the entire PDF into token-size groups and uses an LLM to generate a hierarchical TOC from scratch. It first initializes the structure with generate_toc_init for the first group, then continues building the tree with generate_toc_continue for subsequent groups, effectively creating a document outline where none existed.
Which processing mode provides the highest accuracy for page numbering?
The process_toc_with_page_numbers mode typically yields the highest accuracy because it uses algorithmic alignment rather than pure LLM inference. This mode extracts physical indices from the document content using toc_index_extractor, then calculates the offset between these physical indices and the logical page numbers found in the original TOC via calculate_page_offset. By aligning these two data sources and filling gaps with process_none_page_numbers, it achieves precise pagination even when the PDF contains complex formatting or offset page numbering.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →