How PageIndex Handles Very Large Documents That Exceed LLM Context Limits
PageIndex handles documents exceeding LLM context limits by tokenizing PDFs into bounded chunks, recursively splitting oversized sections, and verifying the generated table of contents before final assembly.
Processing massive PDFs with modern LLMs presents a fundamental challenge: context windows are finite, but documents can span thousands of pages. The VectifyAI/PageIndex repository solves this by implementing a sophisticated chunking and recursive refinement pipeline that ensures every LLM request stays within token limits while preserving document structure.
Tokenization and Chunking Strategy
The foundation of PageIndex's approach is converting raw PDF pages into token-counted units that respect the model's context window.
Converting PDF Pages to Token Counts
In pageindex/utils.py (lines 13-23), the get_page_tokens function reads each page of the PDF, extracts raw text, and counts tokens using the model-specific encoding. This produces a list of tuples containing (page_text, token_length), enabling precise measurement of the document's token footprint.
Grouping Pages into Context-Safe Chunks
Once token counts are established, page_list_to_group_text in pageindex/page_index.py (lines 418-452) merges consecutive pages until the cumulative token count reaches the max-tokens per request threshold (default 20,000).
For documents exceeding this threshold, the function creates multiple overlapping subsets with a default overlap of one page. This overlap ensures that sections straddling chunk boundaries remain detectable in subsequent requests, preventing content loss at segmentation points.
Building and Validating the Table of Contents
After chunking, PageIndex constructs a hierarchical table of contents (TOC) through generation, page number assignment, and iterative verification.
Initial TOC Generation
The system first attempts to extract a physical TOC using toc_extractor. If none exists, the pipeline invokes process_no_toc (lines 68-84 in pageindex/page_index.py), which calls generate_toc_init and generate_toc_continue to prompt the LLM for a hierarchical structure based on the chunked content.
Page Number Assignment
When the generated TOC lacks page numbers, add_page_number_to_toc (lines 53-84) iterates over each chunk and queries the LLM to identify where each section begins, populating the physical_index field for every entry.
Verification and Correction
The verify_toc function (lines 92-145) samples generated entries and asks the LLM to confirm that each physical_index correctly points to its associated section. It returns an accuracy score and a list of incorrect entries.
For erroneous indices, fix_incorrect_toc_with_retries (lines 52-87) invokes fix_incorrect_toc repeatedly, narrowing the search range between valid neighboring indices to locate the correct start page. This process stops after a configurable number of attempts (default 3).
Recursive Processing of Oversized Sections
Even after initial chunking, some sections may still exceed safe processing limits. PageIndex addresses this through recursive decomposition.
The Large Node Detection Logic
In pageindex/page_index.py (lines 991-1019), process_large_node_recursively detects nodes that violate size constraints—specifically those spanning more than opt.max_page_num_each_node (default 10 pages) and exceeding opt.max_token_num_each_node (default 20,000 tokens).
When such a node is detected, the function extracts the sub-document corresponding to that node and runs a mini-page-index on that slice. The original node is then replaced with the resulting deeper hierarchy. This recursion continues until every node satisfies both the page-count and token-count constraints.
Configuration Parameters for Size Limits
The pipeline respects several configurable thresholds defined in config.yaml:
- max_tokens: Caps individual LLM requests (default 20,000)
- overlap_page: Ensures content continuity across chunks (default 1)
- max_page_num_each_node: Maximum pages allowed per TOC node before recursive splitting (default 10)
- max_token_num_each_node: Maximum tokens allowed per node (default 20,000)
- toc_check_page_num: Pages examined when searching for a physical TOC
Adjusting these values allows the same pipeline to handle even larger files without code modifications.
Final Assembly and Output
Once all chunks are processed and large nodes recursively split, PageIndex assembles the final hierarchical structure.
Tree Construction
The post_processing function in pageindex/utils.py (lines 60-74) converts the flat list of sections into a nested tree using list_to_tree. If the document content begins after page 1, the system automatically prepends a Preface node to maintain structural integrity.
Optional Enrichments
Following tree construction, optional post-processing steps in utils.py (lines 58-85) can add node IDs, raw text content, summaries, and single-sentence document descriptions. These enrichments occur after the TOC is built, meaning they do not affect the token limits enforced during the initial processing stages.
Practical Configuration and Code Examples
Basic Usage for Any PDF Size
The default configuration automatically handles documents exceeding context limits:
from pageindex import page_index
result = page_index("path/to/your/large_document.pdf")
print(result["doc_name"])
print(result["structure"]) # nested dict/list representing the TOC tree
This invocation executes the complete pipeline—tokenization, chunking, TOC generation, verification, and recursive splitting—without requiring manual intervention.
Tuning Limits for Extra-Large Books
For documents with unusually long chapters or sections, increase the thresholds:
from pageindex import page_index
# Increase token budget per request to 40,000 and allow nodes up to 20 pages
result = page_index(
"big_book.pdf",
max_token_num_each_node=40000,
max_page_num_each_node=20,
toc_check_page_num=30 # examine more pages when looking for a physical TOC
)
print(result["structure"])
These parameters override the defaults from config.yaml, reducing the depth of recursive splitting required for large sections.
Debugging Chunk Boundaries
To inspect how PageIndex segments your document:
from pageindex.utils import get_page_tokens, page_list_to_group_text
pages = get_page_tokens("huge.pdf")
# Simulate the chunking step
chunks = page_list_to_group_text(
page_contents=[p[0] for p in pages],
token_lengths=[p[1] for p in pages],
max_tokens=20000,
overlap_page=1
)
print(f"Document split into {len(chunks)} chunks")
for i, chunk in enumerate(chunks, 1):
print(f"Chunk {i} length: {len(chunk)} characters")
This reproduces the exact grouping logic used internally by process_no_toc and process_toc_no_page_numbers.
Summary
- Token-bounded chunking: PageIndex converts PDF pages into token-counted units and groups them into chunks under 20,000 tokens using
page_list_to_group_textinpageindex/page_index.py. - Overlap protection: A default one-page overlap between chunks ensures content spanning boundaries remains detectable.
- Hierarchical verification: The pipeline generates a TOC, assigns page numbers via
add_page_number_to_toc, verifies accuracy withverify_toc, and corrects errors throughfix_incorrect_toc_with_retries. - Recursive decomposition:
process_large_node_recursively(lines 991-1019) splits oversized sections into sub-hierarchies until all nodes fit within configured limits. - Configurable limits: All thresholds—including
max_tokens,max_page_num_each_node, andmax_token_num_each_node—are adjustable viaconfig.yamlor function parameters.
Frequently Asked Questions
How does PageIndex prevent content loss when splitting documents into chunks?
PageIndex uses an overlapping chunking strategy implemented in page_list_to_group_text (pageindex/page_index.py, lines 418-452). By default, each chunk overlaps with the previous one by one page, ensuring that sections spanning chunk boundaries remain detectable in subsequent processing steps. This overlap guarantees that no structural information is lost at segmentation points.
What happens if a single chapter exceeds the token limit after initial chunking?
When a node exceeds both the page count (max_page_num_each_node, default 10) and token count (max_token_num_each_node, default 20,000) thresholds, process_large_node_recursively (pageindex/page_index.py, lines 991-1019) extracts that section and runs a mini-page-index on it. This recursive process continues until all resulting sub-nodes fit within the configured limits, creating a deeper hierarchy for lengthy chapters without breaking the context window.
Can I adjust the token limits to work with different LLM context windows?
Yes. All token and page limits are configurable through config.yaml or directly via function parameters. You can modify max_tokens (default 20,000) to match your specific model's context window, or adjust max_page_num_each_node and max_token_num_each_node to control when recursive splitting occurs. The pipeline automatically respects these thresholds without requiring code modifications.
How does PageIndex verify that generated page numbers are accurate?
The pipeline implements a verification loop using verify_toc (pageindex/page_index.py, lines 92-145), which samples generated entries and asks the LLM to confirm that each physical_index correctly points to its associated section. If discrepancies are found, fix_incorrect_toc_with_retries (lines 52-87) narrows the search range between valid neighboring indices and attempts correction up to three times by default, ensuring high accuracy before final tree assembly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →