# What Happens When a Document Processed by PageIndex Has No Table of Contents?

> Discover how PageIndex handles documents lacking a table of contents. Learn about its automatic fallback to LLM-based hierarchical extraction for synthetic TOC generation.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: deep-dive
- Published: 2026-02-16

---

**When PageIndex encounters a PDF without a recognizable table of contents, it automatically falls back to LLM-based hierarchical extraction, generating a synthetic TOC from the document's content structure.**

PageIndex is an open-source Python library designed to extract structured hierarchical indices from PDF documents. When a document processed by PageIndex has no table of contents, the library doesn't fail—it seamlessly transitions to an intelligent fallback mode that analyzes the document's semantic structure to infer headings and page mappings.

## How PageIndex Detects Missing Tables of Contents

The detection logic begins in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) within the `check_toc` function. This routine scans the first *N* pages of the document—controlled by the `opt.toc_check_page_num` configuration parameter—using the `toc_detector_single_page` function to identify potential TOC pages.

If the scan returns an empty list for `toc_page_list`, the result dictionary contains `toc_content: None`, explicitly signaling that no table of contents was found:

```python

# From pageindex/page_index.py, lines 33-42

if not toc_page_list:
    return {
        "toc_page_list": [],
        "toc_content": None,
        "toc_page_range": None
    }

```

This `None` value serves as the trigger for the fallback extraction pipeline.

## Automatic Fallback to Hierarchical Extraction

When `toc_content` is `None`, the `tree_parser` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) (lines 25-38) routes execution to the `process_no_toc` branch instead of parsing an existing TOC structure:

```python
else:
    toc_with_page_number = await meta_processor(
        page_list,
        mode='process_no_toc',
        start_index=1,
        opt=opt,
        logger=logger)

```

This branch invokes the **automatic hierarchical extraction** system, which treats the entire document as unstructured content requiring semantic analysis.

### The process_no_toc Pipeline

The `process_no_toc` function (lines 68-87 in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)) implements a multi-stage pipeline:

1. **Document chunking**: The PDF is split into manageable segments for LLM processing
2. **Initial TOC generation**: The `generate_toc_init` prompt analyzes the first chunk to identify top-level headings
3. **Continuation processing**: The `generate_toc_continue` prompt processes subsequent chunks to find subsections and maintain hierarchy
4. **Physical index mapping**: Each inferred heading is assigned a `physical_index` representing its approximate page location

### LLM-Based TOC Generation

Unlike traditional PDF parsers that rely on embedded metadata, PageIndex uses **language model prompts** to recognize document structure. The `generate_toc_init` and `generate_toc_continue` prompts instruct the model to identify headings based on formatting patterns, font changes, and semantic content rather than PDF bookmarks.

The function returns a list of section entries formatted as:

```python
[
    {
        "title": "Introduction",
        "physical_index": 1,
        "level": 1
    },
    {
        "title": "Methodology",
        "physical_index": 5,
        "level": 1
    }
]

```

## Verification and Post-Processing

Even when generating a synthetic TOC, PageIndex maintains strict quality controls. The output from `process_no_toc` flows through the same verification pipeline as extracted TOCs:

- **`verify_toc`**: Validates that inferred headings actually exist in the document content
- **`fix_incorrect_toc`**: Corrects hierarchical inconsistencies or missing level assignments
- **`post_processing`**: Converts the flat list into a nested tree structure

As implemented in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) (lines 45-48), items that cannot be anchored to a specific page are filtered out:

```python

# Post-processing filters unanchored items

structure = [item for item in structure if item.get("physical_index") is not None]

```

This ensures the final output contains only valid, navigable entries with concrete page mappings.

## Practical Code Examples

### Processing a PDF Without an Embedded TOC

The default behavior handles missing TOCs transparently. When you call `page_index` on a document without bookmarks, the library automatically selects the `process_no_toc` path:

```python
from pageindex import page_index

# Process a PDF that lacks a table of contents

result = page_index("document_without_toc.pdf")

# The structure contains LLM-generated headings

print(f"Document: {result['doc_name']}")
for section in result['structure']:
    print(f"- {section['title']} (page {section['physical_index']})")

```

### Forcing the No-TOC Fallback Mode

You can disable TOC detection entirely by setting `toc_check_page_num` to `0`. This forces PageIndex to always use the LLM-based extraction method, even if the PDF contains embedded bookmarks:

```python
from pageindex import page_index, ConfigLoader

# Configuration to skip TOC detection

user_config = {
    "toc_check_page_num": 0,  # Disable scanning for existing TOCs

    "model": "gpt-4o-mini-2024-07-18"
}
opt = ConfigLoader().load(user_config)

# Process with forced hierarchical extraction

result = page_index("annual_report.pdf", opt=opt)

```

### Debugging TOC Detection Decisions

Enable logging to observe whether PageIndex found an existing TOC or fell back to automatic extraction:

```python
import logging
logging.basicConfig(level=logging.INFO)

from pageindex import page_index

result = page_index("ambiguous_document.pdf")

# Check logs for:

# - "no toc found" (triggers process_no_toc)

# - "toc found" (uses existing bookmarks)

```

## Key Source Files and Functions

Understanding the no-TOC fallback requires familiarity with these core components:

| File | Critical Functions | Purpose |
|------|-------------------|---------|
| [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) | `check_toc` (lines 33-42), `tree_parser` (lines 25-38), `process_no_toc` (lines 68-87) | Implements TOC detection logic and the LLM-based fallback extraction pipeline. |
| [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) | `ChatGPT_API`, `ChatGPT_API_with_finish_reason` | Provides LLM interaction wrappers for the `generate_toc_init` and `generate_toc_continue` prompts. |
| [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) | `toc_check_page_num`, model parameters | Controls how many pages are inspected for existing TOCs before falling back to synthesis. |

These files together implement the graceful degradation path that allows PageIndex to handle documents lacking a table of contents, guaranteeing that an index is always produced.

## Summary

When a document processed by PageIndex has no table of contents, the library executes a graceful fallback strategy:

- **Detection**: The `check_toc` function scans the first *N* pages (configurable via `toc_check_page_num`) and returns `toc_content: None` when no TOC is found.

- **Fallback**: The `tree_parser` routes to `process_no_toc`, which uses LLM prompts (`generate_toc_init` and `generate_toc_continue`) to infer document structure from content.

- **Validation**: Synthetic TOCs undergo the same verification pipeline (`verify_toc`, `fix_incorrect_toc`, `post_processing`) as extracted ones, filtering out unanchored entries.

- **Output**: The final result contains a complete hierarchical index with `physical_index` mappings, even for documents that originally lacked any navigational metadata.

## Frequently Asked Questions

### Does PageIndex fail if a PDF has no table of contents?

No. PageIndex is designed to handle documents without embedded TOCs gracefully. When the `check_toc` function returns `toc_content: None`, the system automatically invokes `process_no_toc` to generate a synthetic table of contents using the configured language model. The process never raises an error solely due to missing TOC metadata.

### How accurate is the automatically generated table of contents?

Accuracy depends on the language model's ability to recognize heading patterns, font changes, and semantic structure within the document. The system uses specialized prompts (`generate_toc_init` for initial sections and `generate_toc_continue` for subsequent content) to maintain hierarchical consistency. Additionally, the `verify_toc` and `fix_incorrect_toc` functions post-process the results to correct inconsistencies and filter out entries that cannot be mapped to specific pages.

### Can I force PageIndex to ignore an existing table of contents?

Yes. You can disable TOC detection by setting the `toc_check_page_num` configuration parameter to `0` in your `ConfigLoader` options. This forces the `check_toc` function to skip scanning entirely, causing `tree_parser` to always execute the `process_no_toc` branch. This is useful when embedded TOCs are malformed or when you want consistent LLM-based extraction across all documents.

### What configuration options control the TOC detection behavior?

The primary configuration is `toc_check_page_num`, which specifies how many pages at the beginning of the document to scan for a table of contents (default is typically a small number like 3-5). Additional relevant configurations include the `model` parameter (specifying which LLM to use for `process_no_toc`) and various prompt templates (`generate_toc_init`, `generate_toc_continue`) that can be customized to improve extraction accuracy for specific document types.