# How PageIndex Handles Three Table of Contents Processing Modes

> Discover how PageIndex intelligently selects one of three table of contents processing modes based on your PDF's structure and page number inclusion. Optimize your document analysis now.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: how-to-guide
- Published: 2026-02-16

---

**PageIndex automatically selects one of three table of contents processing modes—`process_toc_with_page_numbers`, `process_toc_no_page_numbers`, or `process_no_toc`—based on whether the PDF contains a detectable TOC and whether that TOC includes page numbers.**

The VectifyAI/PageIndex library adapts its extraction strategy to the structural quality of your PDF documents. Understanding these three table of contents processing modes helps you predict how the system will index your documents and troubleshoot extraction accuracy.

## The Three Table of Contents Processing Modes

PageIndex implements distinct pipelines in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) to handle varying levels of TOC completeness. Each mode corresponds to a specific function ranging from lines 68 to 144.

### Mode 1: process_toc_with_page_numbers

This mode activates when `detect_page_index` identifies that the TOC already contains explicit page numbers (e.g., "1 Introduction … 3"). According to the source code in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) (lines 114-144), the pipeline executes five sequential steps:

1. **Transform** the raw TOC into structured JSON using `toc_transformer`
2. **Strip** existing page numbers via `remove_page_number` to isolate titles
3. **Extract** physical indices from a small window of pages following the TOC using `toc_index_extractor`
4. **Align** extracted indices with original page numbers through `extract_matching_page_pairs`, `calculate_page_offset`, and `add_page_offset_to_toc_json`
5. **Fill** missing entries using `process_none_page_numbers`

### Mode 2: process_toc_no_page_numbers

When `detect_page_index` returns "no" but a TOC page exists without page references, PageIndex falls back to `process_toc_no_page_numbers` (lines 89-109). This mode handles TOCs that list chapter titles but omit pagination:

1. **Convert** the raw TOC to JSON using `toc_transformer`
2. **Iterate** over the entire document grouped into token-size chunks
3. **Assign** page numbers via LLM calls through `add_page_number_to_toc`, which returns `<physical_index_X>` strings
4. **Convert** these string tokens to integers using `convert_physical_index_to_int`

### Mode 3: process_no_toc

If `check_toc` fails to detect any TOC page, the system triggers `process_no_toc` (lines 68-87) to generate a hierarchical structure from scratch:

1. **Split** the entire PDF into token-size groups
2. **Initialize** the TOC by prompting the LLM to create a hierarchy for the first group via `generate_toc_init`
3. **Continue** the tree for subsequent groups using `generate_toc_continue`
4. **Convert** the generated `<physical_index_X>` strings to integers

## How Mode Selection Works

The orchestration logic resides in `meta_processor` at lines 951-967 of [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py). After `check_toc` completes its initial assessment, the function executes a conditional branch:

```python
if mode == 'process_toc_with_page_numbers':
    toc_with_page_number = process_toc_with_page_numbers(...)
elif mode == 'process_toc_no_page_numbers':
    toc_with_page_number = process_toc_no_page_numbers(...)
else:  # process_no_toc

    toc_with_page_number = process_no_toc(...)

```

This automatic selection requires no user intervention. The system evaluates the PDF structure and selects the appropriate table of contents processing mode to maximize extraction accuracy.

## Practical Code Examples

### Basic Usage with Automatic Mode Selection

The standard API handles mode selection transparently. The `page_index` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) loads [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml), constructs the `opt` object, and invokes `page_index_main`, which triggers `tree_parser` and subsequently `check_toc`:

```python
from pageindex.page_index import page_index

# Path to a PDF (replace with your own file)

pdf_path = "tests/pdfs/2023-annual-report.pdf"

# Run the full pipeline with default configuration (model, token limits, etc.)

result = page_index(pdf_path)

print("Document name:", result["doc_name"])
print("Extracted TOC:")
for item in result["structure"]:
    print(f"{item['structure']}  {item['title']}  (page {item['physical_index']})")

```

### Forcing a Specific Mode via Configuration

While the public API does not expose the mode parameter directly, you can influence the selection by adjusting TOC detection parameters in [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml). To force `process_no_toc` by preventing TOC detection:

```python
from pageindex.page_index import page_index_main, ConfigLoader
from pageindex.utils import JsonLogger

# Load a custom config that disables TOC detection (set a very low toc_check_page_num)

custom_opt = ConfigLoader().load({
    "toc_check_page_num": 0,          # forces find_toc_pages to return []

    "model": "gpt-4o-2024-11-20"
})

# Run the low‑level entry point, bypassing the automatic check

with open("tests/pdfs/2023-annual-report.pdf", "rb") as f:
    result = page_index_main(f, opt=custom_opt)

print(result["structure"])

```

Setting `toc_check_page_num` to `0` ensures `find_toc_pages` returns an empty list, causing `meta_processor` to execute the `process_no_toc` branch.

### Inspecting the Selected Mode

Enable Python logging to observe which table of contents processing mode the system selects:

```python
import logging
from pageindex.page_index import page_index

logging.basicConfig(level=logging.INFO)

pdf_path = "tests/pdfs/four-lectures.pdf"
result = page_index(pdf_path)

# The logger prints something like:

#   INFO:root:process_toc_no_page_numbers

#   INFO:root:accuracy: 98.75%

```

The `meta_processor` function explicitly prints the mode variable at execution start, allowing you to verify the automatic selection logic.

## Key Implementation Files

Understanding the file structure helps navigate the three table of contents processing modes:

| File | Role | Relevant Excerpts |
|------|------|-------------------|
| **[`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)** | Core orchestration – defines the three processing functions, TOC detection, and the `meta_processor` that selects the mode. | `process_no_toc` (lines 68-87), `process_toc_no_page_numbers` (lines 89-109), `process_toc_with_page_numbers` (lines 114-144), `meta_processor` (lines 951-967) |
| **[`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py)** | Helper functions for token counting, LLM calls, and low-level PDF → token conversion used by all three modes. | `get_page_tokens`, `ChatGPT_API`, `count_tokens` |
| **[`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml)** | Default configuration (model, page-check limits, node-size thresholds). | `model`, `toc_check_page_num`, `max_page_num_each_node` |
| **[`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py)** | CLI entry point – parses arguments and forwards to `page_index`. Shows how the library is normally executed. | Argument handling for `--toc-check-pages`, `--model` |
| **[`README.md`](https://github.com/VectifyAI/PageIndex/blob/main/README.md)** | High-level description of the library and sample commands. | Usage examples, installation steps |

These files collectively implement the adaptive TOC extraction strategy that gracefully degrades from **full-featured page-indexed TOC** → **page-less TOC** → **LLM-generated TOC**.

## Summary

- **Automatic mode selection** occurs in `meta_processor` (lines 951-967 of [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)) based on TOC detection results, requiring no user configuration.
- **`process_toc_with_page_numbers`** handles documents with explicit page numbers in the TOC, using alignment algorithms to map physical indices to logical page numbers.
- **`process_toc_no_page_numbers`** processes TOCs lacking pagination by querying an LLM across document chunks to assign physical indices.
- **`process_no_toc`** generates a complete hierarchical TOC from scratch when no TOC page exists, using iterative LLM calls to build the document structure.
- **Configuration control** allows advanced users to force specific modes by adjusting `toc_check_page_num` in [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) to manipulate detection behavior.

## Frequently Asked Questions

### How does PageIndex decide which table of contents processing mode to use?

PageIndex automatically determines the appropriate mode through the `check_toc` and `detect_page_index` functions in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py). If the system detects a TOC containing page numbers, it selects `process_toc_with_page_numbers`. If it finds a TOC without page numbers, it chooses `process_toc_no_page_numbers`. If no TOC is detected at all, it defaults to `process_no_toc`. This decision happens inside `meta_processor` without requiring user intervention.

### Can I force PageIndex to use a specific processing mode?

While the public API does not expose a direct `mode` parameter, you can influence the selection by modifying configuration values that affect TOC detection. For example, setting `toc_check_page_num` to `0` in your configuration forces `find_toc_pages` to return an empty list, which causes `meta_processor` to execute the `process_no_toc` branch. Similarly, adjusting detection thresholds can bias the system toward treating detected TOCs as having or lacking page numbers.

### What happens when PageIndex cannot find any table of contents in the document?

When `check_toc` fails to identify a TOC page, PageIndex invokes `process_no_toc` (lines 68-87 in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)). This function splits the entire PDF into token-size groups and uses an LLM to generate a hierarchical TOC from scratch. It first initializes the structure with `generate_toc_init` for the first group, then continues building the tree with `generate_toc_continue` for subsequent groups, effectively creating a document outline where none existed.

### Which processing mode provides the highest accuracy for page numbering?

The `process_toc_with_page_numbers` mode typically yields the highest accuracy because it uses algorithmic alignment rather than pure LLM inference. This mode extracts physical indices from the document content using `toc_index_extractor`, then calculates the offset between these physical indices and the logical page numbers found in the original TOC via `calculate_page_offset`. By aligning these two data sources and filling gaps with `process_none_page_numbers`, it achieves precise pagination even when the PDF contains complex formatting or offset page numbering.