What Happens When a Document Processed by PageIndex Has No Table of Contents?
When PageIndex encounters a PDF without a recognizable table of contents, it automatically falls back to LLM-based hierarchical extraction, generating a synthetic TOC from the document's content structure.
PageIndex is an open-source Python library designed to extract structured hierarchical indices from PDF documents. When a document processed by PageIndex has no table of contents, the library doesn't fail—it seamlessly transitions to an intelligent fallback mode that analyzes the document's semantic structure to infer headings and page mappings.
How PageIndex Detects Missing Tables of Contents
The detection logic begins in pageindex/page_index.py within the check_toc function. This routine scans the first N pages of the document—controlled by the opt.toc_check_page_num configuration parameter—using the toc_detector_single_page function to identify potential TOC pages.
If the scan returns an empty list for toc_page_list, the result dictionary contains toc_content: None, explicitly signaling that no table of contents was found:
# From pageindex/page_index.py, lines 33-42
if not toc_page_list:
return {
"toc_page_list": [],
"toc_content": None,
"toc_page_range": None
}
This None value serves as the trigger for the fallback extraction pipeline.
Automatic Fallback to Hierarchical Extraction
When toc_content is None, the tree_parser function in pageindex/page_index.py (lines 25-38) routes execution to the process_no_toc branch instead of parsing an existing TOC structure:
else:
toc_with_page_number = await meta_processor(
page_list,
mode='process_no_toc',
start_index=1,
opt=opt,
logger=logger)
This branch invokes the automatic hierarchical extraction system, which treats the entire document as unstructured content requiring semantic analysis.
The process_no_toc Pipeline
The process_no_toc function (lines 68-87 in pageindex/page_index.py) implements a multi-stage pipeline:
- Document chunking: The PDF is split into manageable segments for LLM processing
- Initial TOC generation: The
generate_toc_initprompt analyzes the first chunk to identify top-level headings - Continuation processing: The
generate_toc_continueprompt processes subsequent chunks to find subsections and maintain hierarchy - Physical index mapping: Each inferred heading is assigned a
physical_indexrepresenting its approximate page location
LLM-Based TOC Generation
Unlike traditional PDF parsers that rely on embedded metadata, PageIndex uses language model prompts to recognize document structure. The generate_toc_init and generate_toc_continue prompts instruct the model to identify headings based on formatting patterns, font changes, and semantic content rather than PDF bookmarks.
The function returns a list of section entries formatted as:
[
{
"title": "Introduction",
"physical_index": 1,
"level": 1
},
{
"title": "Methodology",
"physical_index": 5,
"level": 1
}
]
Verification and Post-Processing
Even when generating a synthetic TOC, PageIndex maintains strict quality controls. The output from process_no_toc flows through the same verification pipeline as extracted TOCs:
verify_toc: Validates that inferred headings actually exist in the document contentfix_incorrect_toc: Corrects hierarchical inconsistencies or missing level assignmentspost_processing: Converts the flat list into a nested tree structure
As implemented in pageindex/page_index.py (lines 45-48), items that cannot be anchored to a specific page are filtered out:
# Post-processing filters unanchored items
structure = [item for item in structure if item.get("physical_index") is not None]
This ensures the final output contains only valid, navigable entries with concrete page mappings.
Practical Code Examples
Processing a PDF Without an Embedded TOC
The default behavior handles missing TOCs transparently. When you call page_index on a document without bookmarks, the library automatically selects the process_no_toc path:
from pageindex import page_index
# Process a PDF that lacks a table of contents
result = page_index("document_without_toc.pdf")
# The structure contains LLM-generated headings
print(f"Document: {result['doc_name']}")
for section in result['structure']:
print(f"- {section['title']} (page {section['physical_index']})")
Forcing the No-TOC Fallback Mode
You can disable TOC detection entirely by setting toc_check_page_num to 0. This forces PageIndex to always use the LLM-based extraction method, even if the PDF contains embedded bookmarks:
from pageindex import page_index, ConfigLoader
# Configuration to skip TOC detection
user_config = {
"toc_check_page_num": 0, # Disable scanning for existing TOCs
"model": "gpt-4o-mini-2024-07-18"
}
opt = ConfigLoader().load(user_config)
# Process with forced hierarchical extraction
result = page_index("annual_report.pdf", opt=opt)
Debugging TOC Detection Decisions
Enable logging to observe whether PageIndex found an existing TOC or fell back to automatic extraction:
import logging
logging.basicConfig(level=logging.INFO)
from pageindex import page_index
result = page_index("ambiguous_document.pdf")
# Check logs for:
# - "no toc found" (triggers process_no_toc)
# - "toc found" (uses existing bookmarks)
Key Source Files and Functions
Understanding the no-TOC fallback requires familiarity with these core components:
| File | Critical Functions | Purpose |
|---|---|---|
pageindex/page_index.py |
check_toc (lines 33-42), tree_parser (lines 25-38), process_no_toc (lines 68-87) |
Implements TOC detection logic and the LLM-based fallback extraction pipeline. |
pageindex/utils.py |
ChatGPT_API, ChatGPT_API_with_finish_reason |
Provides LLM interaction wrappers for the generate_toc_init and generate_toc_continue prompts. |
pageindex/config.yaml |
toc_check_page_num, model parameters |
Controls how many pages are inspected for existing TOCs before falling back to synthesis. |
These files together implement the graceful degradation path that allows PageIndex to handle documents lacking a table of contents, guaranteeing that an index is always produced.
Summary
When a document processed by PageIndex has no table of contents, the library executes a graceful fallback strategy:
-
Detection: The
check_tocfunction scans the first N pages (configurable viatoc_check_page_num) and returnstoc_content: Nonewhen no TOC is found. -
Fallback: The
tree_parserroutes toprocess_no_toc, which uses LLM prompts (generate_toc_initandgenerate_toc_continue) to infer document structure from content. -
Validation: Synthetic TOCs undergo the same verification pipeline (
verify_toc,fix_incorrect_toc,post_processing) as extracted ones, filtering out unanchored entries. -
Output: The final result contains a complete hierarchical index with
physical_indexmappings, even for documents that originally lacked any navigational metadata.
Frequently Asked Questions
Does PageIndex fail if a PDF has no table of contents?
No. PageIndex is designed to handle documents without embedded TOCs gracefully. When the check_toc function returns toc_content: None, the system automatically invokes process_no_toc to generate a synthetic table of contents using the configured language model. The process never raises an error solely due to missing TOC metadata.
How accurate is the automatically generated table of contents?
Accuracy depends on the language model's ability to recognize heading patterns, font changes, and semantic structure within the document. The system uses specialized prompts (generate_toc_init for initial sections and generate_toc_continue for subsequent content) to maintain hierarchical consistency. Additionally, the verify_toc and fix_incorrect_toc functions post-process the results to correct inconsistencies and filter out entries that cannot be mapped to specific pages.
Can I force PageIndex to ignore an existing table of contents?
Yes. You can disable TOC detection by setting the toc_check_page_num configuration parameter to 0 in your ConfigLoader options. This forces the check_toc function to skip scanning entirely, causing tree_parser to always execute the process_no_toc branch. This is useful when embedded TOCs are malformed or when you want consistent LLM-based extraction across all documents.
What configuration options control the TOC detection behavior?
The primary configuration is toc_check_page_num, which specifies how many pages at the beginning of the document to scan for a table of contents (default is typically a small number like 3-5). Additional relevant configurations include the model parameter (specifying which LLM to use for process_no_toc) and various prompt templates (generate_toc_init, generate_toc_continue) that can be customized to improve extraction accuracy for specific document types.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →