How PageIndex Handles Very Large Documents That Exceed LLM Context Limits

PageIndex handles documents exceeding LLM context limits by tokenizing PDFs into bounded chunks, recursively splitting oversized sections, and verifying the generated table of contents before final assembly.

Processing massive PDFs with modern LLMs presents a fundamental challenge: context windows are finite, but documents can span thousands of pages. The VectifyAI/PageIndex repository solves this by implementing a sophisticated chunking and recursive refinement pipeline that ensures every LLM request stays within token limits while preserving document structure.

Tokenization and Chunking Strategy

The foundation of PageIndex's approach is converting raw PDF pages into token-counted units that respect the model's context window.

Converting PDF Pages to Token Counts

In pageindex/utils.py (lines 13-23), the get_page_tokens function reads each page of the PDF, extracts raw text, and counts tokens using the model-specific encoding. This produces a list of tuples containing (page_text, token_length), enabling precise measurement of the document's token footprint.

Grouping Pages into Context-Safe Chunks

Once token counts are established, page_list_to_group_text in pageindex/page_index.py (lines 418-452) merges consecutive pages until the cumulative token count reaches the max-tokens per request threshold (default 20,000).

For documents exceeding this threshold, the function creates multiple overlapping subsets with a default overlap of one page. This overlap ensures that sections straddling chunk boundaries remain detectable in subsequent requests, preventing content loss at segmentation points.

Building and Validating the Table of Contents

After chunking, PageIndex constructs a hierarchical table of contents (TOC) through generation, page number assignment, and iterative verification.

Initial TOC Generation

The system first attempts to extract a physical TOC using toc_extractor. If none exists, the pipeline invokes process_no_toc (lines 68-84 in pageindex/page_index.py), which calls generate_toc_init and generate_toc_continue to prompt the LLM for a hierarchical structure based on the chunked content.

Page Number Assignment

When the generated TOC lacks page numbers, add_page_number_to_toc (lines 53-84) iterates over each chunk and queries the LLM to identify where each section begins, populating the physical_index field for every entry.

Verification and Correction

The verify_toc function (lines 92-145) samples generated entries and asks the LLM to confirm that each physical_index correctly points to its associated section. It returns an accuracy score and a list of incorrect entries.

For erroneous indices, fix_incorrect_toc_with_retries (lines 52-87) invokes fix_incorrect_toc repeatedly, narrowing the search range between valid neighboring indices to locate the correct start page. This process stops after a configurable number of attempts (default 3).

Recursive Processing of Oversized Sections

Even after initial chunking, some sections may still exceed safe processing limits. PageIndex addresses this through recursive decomposition.

The Large Node Detection Logic

In pageindex/page_index.py (lines 991-1019), process_large_node_recursively detects nodes that violate size constraints—specifically those spanning more than opt.max_page_num_each_node (default 10 pages) and exceeding opt.max_token_num_each_node (default 20,000 tokens).

When such a node is detected, the function extracts the sub-document corresponding to that node and runs a mini-page-index on that slice. The original node is then replaced with the resulting deeper hierarchy. This recursion continues until every node satisfies both the page-count and token-count constraints.

Configuration Parameters for Size Limits

The pipeline respects several configurable thresholds defined in config.yaml:

  • max_tokens: Caps individual LLM requests (default 20,000)
  • overlap_page: Ensures content continuity across chunks (default 1)
  • max_page_num_each_node: Maximum pages allowed per TOC node before recursive splitting (default 10)
  • max_token_num_each_node: Maximum tokens allowed per node (default 20,000)
  • toc_check_page_num: Pages examined when searching for a physical TOC

Adjusting these values allows the same pipeline to handle even larger files without code modifications.

Final Assembly and Output

Once all chunks are processed and large nodes recursively split, PageIndex assembles the final hierarchical structure.

Tree Construction

The post_processing function in pageindex/utils.py (lines 60-74) converts the flat list of sections into a nested tree using list_to_tree. If the document content begins after page 1, the system automatically prepends a Preface node to maintain structural integrity.

Optional Enrichments

Following tree construction, optional post-processing steps in utils.py (lines 58-85) can add node IDs, raw text content, summaries, and single-sentence document descriptions. These enrichments occur after the TOC is built, meaning they do not affect the token limits enforced during the initial processing stages.

Practical Configuration and Code Examples

Basic Usage for Any PDF Size

The default configuration automatically handles documents exceeding context limits:

from pageindex import page_index

result = page_index("path/to/your/large_document.pdf")
print(result["doc_name"])
print(result["structure"])   # nested dict/list representing the TOC tree

This invocation executes the complete pipeline—tokenization, chunking, TOC generation, verification, and recursive splitting—without requiring manual intervention.

Tuning Limits for Extra-Large Books

For documents with unusually long chapters or sections, increase the thresholds:

from pageindex import page_index

# Increase token budget per request to 40,000 and allow nodes up to 20 pages

result = page_index(
    "big_book.pdf",
    max_token_num_each_node=40000,
    max_page_num_each_node=20,
    toc_check_page_num=30   # examine more pages when looking for a physical TOC

)

print(result["structure"])

These parameters override the defaults from config.yaml, reducing the depth of recursive splitting required for large sections.

Debugging Chunk Boundaries

To inspect how PageIndex segments your document:

from pageindex.utils import get_page_tokens, page_list_to_group_text

pages = get_page_tokens("huge.pdf")

# Simulate the chunking step

chunks = page_list_to_group_text(
    page_contents=[p[0] for p in pages],
    token_lengths=[p[1] for p in pages],
    max_tokens=20000,
    overlap_page=1
)

print(f"Document split into {len(chunks)} chunks")
for i, chunk in enumerate(chunks, 1):
    print(f"Chunk {i} length: {len(chunk)} characters")

This reproduces the exact grouping logic used internally by process_no_toc and process_toc_no_page_numbers.

Summary

  • Token-bounded chunking: PageIndex converts PDF pages into token-counted units and groups them into chunks under 20,000 tokens using page_list_to_group_text in pageindex/page_index.py.
  • Overlap protection: A default one-page overlap between chunks ensures content spanning boundaries remains detectable.
  • Hierarchical verification: The pipeline generates a TOC, assigns page numbers via add_page_number_to_toc, verifies accuracy with verify_toc, and corrects errors through fix_incorrect_toc_with_retries.
  • Recursive decomposition: process_large_node_recursively (lines 991-1019) splits oversized sections into sub-hierarchies until all nodes fit within configured limits.
  • Configurable limits: All thresholds—including max_tokens, max_page_num_each_node, and max_token_num_each_node—are adjustable via config.yaml or function parameters.

Frequently Asked Questions

How does PageIndex prevent content loss when splitting documents into chunks?

PageIndex uses an overlapping chunking strategy implemented in page_list_to_group_text (pageindex/page_index.py, lines 418-452). By default, each chunk overlaps with the previous one by one page, ensuring that sections spanning chunk boundaries remain detectable in subsequent processing steps. This overlap guarantees that no structural information is lost at segmentation points.

What happens if a single chapter exceeds the token limit after initial chunking?

When a node exceeds both the page count (max_page_num_each_node, default 10) and token count (max_token_num_each_node, default 20,000) thresholds, process_large_node_recursively (pageindex/page_index.py, lines 991-1019) extracts that section and runs a mini-page-index on it. This recursive process continues until all resulting sub-nodes fit within the configured limits, creating a deeper hierarchy for lengthy chapters without breaking the context window.

Can I adjust the token limits to work with different LLM context windows?

Yes. All token and page limits are configurable through config.yaml or directly via function parameters. You can modify max_tokens (default 20,000) to match your specific model's context window, or adjust max_page_num_each_node and max_token_num_each_node to control when recursive splitting occurs. The pipeline automatically respects these thresholds without requiring code modifications.

How does PageIndex verify that generated page numbers are accurate?

The pipeline implements a verification loop using verify_toc (pageindex/page_index.py, lines 92-145), which samples generated entries and asks the LLM to confirm that each physical_index correctly points to its associated section. If discrepancies are found, fix_incorrect_toc_with_retries (lines 52-87) narrows the search range between valid neighboring indices and attempts correction up to three times by default, ensuring high accuracy before final tree assembly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →