What Happens When a Document Processed by PageIndex Has No Table of Contents?

When PageIndex encounters a PDF without a recognizable table of contents, it automatically falls back to LLM-based hierarchical extraction, generating a synthetic TOC from the document's content structure.

PageIndex is an open-source Python library designed to extract structured hierarchical indices from PDF documents. When a document processed by PageIndex has no table of contents, the library doesn't fail—it seamlessly transitions to an intelligent fallback mode that analyzes the document's semantic structure to infer headings and page mappings.

How PageIndex Detects Missing Tables of Contents

The detection logic begins in pageindex/page_index.py within the check_toc function. This routine scans the first N pages of the document—controlled by the opt.toc_check_page_num configuration parameter—using the toc_detector_single_page function to identify potential TOC pages.

If the scan returns an empty list for toc_page_list, the result dictionary contains toc_content: None, explicitly signaling that no table of contents was found:


# From pageindex/page_index.py, lines 33-42

if not toc_page_list:
    return {
        "toc_page_list": [],
        "toc_content": None,
        "toc_page_range": None
    }

This None value serves as the trigger for the fallback extraction pipeline.

Automatic Fallback to Hierarchical Extraction

When toc_content is None, the tree_parser function in pageindex/page_index.py (lines 25-38) routes execution to the process_no_toc branch instead of parsing an existing TOC structure:

else:
    toc_with_page_number = await meta_processor(
        page_list,
        mode='process_no_toc',
        start_index=1,
        opt=opt,
        logger=logger)

This branch invokes the automatic hierarchical extraction system, which treats the entire document as unstructured content requiring semantic analysis.

The process_no_toc Pipeline

The process_no_toc function (lines 68-87 in pageindex/page_index.py) implements a multi-stage pipeline:

  1. Document chunking: The PDF is split into manageable segments for LLM processing
  2. Initial TOC generation: The generate_toc_init prompt analyzes the first chunk to identify top-level headings
  3. Continuation processing: The generate_toc_continue prompt processes subsequent chunks to find subsections and maintain hierarchy
  4. Physical index mapping: Each inferred heading is assigned a physical_index representing its approximate page location

LLM-Based TOC Generation

Unlike traditional PDF parsers that rely on embedded metadata, PageIndex uses language model prompts to recognize document structure. The generate_toc_init and generate_toc_continue prompts instruct the model to identify headings based on formatting patterns, font changes, and semantic content rather than PDF bookmarks.

The function returns a list of section entries formatted as:

[
    {
        "title": "Introduction",
        "physical_index": 1,
        "level": 1
    },
    {
        "title": "Methodology",
        "physical_index": 5,
        "level": 1
    }
]

Verification and Post-Processing

Even when generating a synthetic TOC, PageIndex maintains strict quality controls. The output from process_no_toc flows through the same verification pipeline as extracted TOCs:

  • verify_toc: Validates that inferred headings actually exist in the document content
  • fix_incorrect_toc: Corrects hierarchical inconsistencies or missing level assignments
  • post_processing: Converts the flat list into a nested tree structure

As implemented in pageindex/page_index.py (lines 45-48), items that cannot be anchored to a specific page are filtered out:


# Post-processing filters unanchored items

structure = [item for item in structure if item.get("physical_index") is not None]

This ensures the final output contains only valid, navigable entries with concrete page mappings.

Practical Code Examples

Processing a PDF Without an Embedded TOC

The default behavior handles missing TOCs transparently. When you call page_index on a document without bookmarks, the library automatically selects the process_no_toc path:

from pageindex import page_index

# Process a PDF that lacks a table of contents

result = page_index("document_without_toc.pdf")

# The structure contains LLM-generated headings

print(f"Document: {result['doc_name']}")
for section in result['structure']:
    print(f"- {section['title']} (page {section['physical_index']})")

Forcing the No-TOC Fallback Mode

You can disable TOC detection entirely by setting toc_check_page_num to 0. This forces PageIndex to always use the LLM-based extraction method, even if the PDF contains embedded bookmarks:

from pageindex import page_index, ConfigLoader

# Configuration to skip TOC detection

user_config = {
    "toc_check_page_num": 0,  # Disable scanning for existing TOCs

    "model": "gpt-4o-mini-2024-07-18"
}
opt = ConfigLoader().load(user_config)

# Process with forced hierarchical extraction

result = page_index("annual_report.pdf", opt=opt)

Debugging TOC Detection Decisions

Enable logging to observe whether PageIndex found an existing TOC or fell back to automatic extraction:

import logging
logging.basicConfig(level=logging.INFO)

from pageindex import page_index

result = page_index("ambiguous_document.pdf")

# Check logs for:

# - "no toc found" (triggers process_no_toc)

# - "toc found" (uses existing bookmarks)

Key Source Files and Functions

Understanding the no-TOC fallback requires familiarity with these core components:

File Critical Functions Purpose
pageindex/page_index.py check_toc (lines 33-42), tree_parser (lines 25-38), process_no_toc (lines 68-87) Implements TOC detection logic and the LLM-based fallback extraction pipeline.
pageindex/utils.py ChatGPT_API, ChatGPT_API_with_finish_reason Provides LLM interaction wrappers for the generate_toc_init and generate_toc_continue prompts.
pageindex/config.yaml toc_check_page_num, model parameters Controls how many pages are inspected for existing TOCs before falling back to synthesis.

These files together implement the graceful degradation path that allows PageIndex to handle documents lacking a table of contents, guaranteeing that an index is always produced.

Summary

When a document processed by PageIndex has no table of contents, the library executes a graceful fallback strategy:

  • Detection: The check_toc function scans the first N pages (configurable via toc_check_page_num) and returns toc_content: None when no TOC is found.

  • Fallback: The tree_parser routes to process_no_toc, which uses LLM prompts (generate_toc_init and generate_toc_continue) to infer document structure from content.

  • Validation: Synthetic TOCs undergo the same verification pipeline (verify_toc, fix_incorrect_toc, post_processing) as extracted ones, filtering out unanchored entries.

  • Output: The final result contains a complete hierarchical index with physical_index mappings, even for documents that originally lacked any navigational metadata.

Frequently Asked Questions

Does PageIndex fail if a PDF has no table of contents?

No. PageIndex is designed to handle documents without embedded TOCs gracefully. When the check_toc function returns toc_content: None, the system automatically invokes process_no_toc to generate a synthetic table of contents using the configured language model. The process never raises an error solely due to missing TOC metadata.

How accurate is the automatically generated table of contents?

Accuracy depends on the language model's ability to recognize heading patterns, font changes, and semantic structure within the document. The system uses specialized prompts (generate_toc_init for initial sections and generate_toc_continue for subsequent content) to maintain hierarchical consistency. Additionally, the verify_toc and fix_incorrect_toc functions post-process the results to correct inconsistencies and filter out entries that cannot be mapped to specific pages.

Can I force PageIndex to ignore an existing table of contents?

Yes. You can disable TOC detection by setting the toc_check_page_num configuration parameter to 0 in your ConfigLoader options. This forces the check_toc function to skip scanning entirely, causing tree_parser to always execute the process_no_toc branch. This is useful when embedded TOCs are malformed or when you want consistent LLM-based extraction across all documents.

What configuration options control the TOC detection behavior?

The primary configuration is toc_check_page_num, which specifies how many pages at the beginning of the document to scan for a table of contents (default is typically a small number like 3-5). Additional relevant configurations include the model parameter (specifying which LLM to use for process_no_toc) and various prompt templates (generate_toc_init, generate_toc_continue) that can be customized to improve extraction accuracy for specific document types.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →