# How PageIndex Handles Very Large Documents That Exceed LLM Context Limits

> Discover how PageIndex tackles large documents beyond LLM context limits by expertly chunking, splitting, and verifying content for seamless processing. Learn more now.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: how-to-guide
- Published: 2026-02-16

---

**PageIndex handles documents exceeding LLM context limits by tokenizing PDFs into bounded chunks, recursively splitting oversized sections, and verifying the generated table of contents before final assembly.**

Processing massive PDFs with modern LLMs presents a fundamental challenge: context windows are finite, but documents can span thousands of pages. The VectifyAI/PageIndex repository solves this by implementing a sophisticated chunking and recursive refinement pipeline that ensures every LLM request stays within token limits while preserving document structure.

## Tokenization and Chunking Strategy

The foundation of PageIndex's approach is converting raw PDF pages into token-counted units that respect the model's context window.

### Converting PDF Pages to Token Counts

In [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 13-23), the `get_page_tokens` function reads each page of the PDF, extracts raw text, and counts tokens using the model-specific encoding. This produces a list of tuples containing `(page_text, token_length)`, enabling precise measurement of the document's token footprint.

### Grouping Pages into Context-Safe Chunks

Once token counts are established, `page_list_to_group_text` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) (lines 418-452) merges consecutive pages until the cumulative token count reaches the **max-tokens per request** threshold (default 20,000).

For documents exceeding this threshold, the function creates multiple overlapping subsets with a default overlap of one page. This overlap ensures that sections straddling chunk boundaries remain detectable in subsequent requests, preventing content loss at segmentation points.

## Building and Validating the Table of Contents

After chunking, PageIndex constructs a hierarchical table of contents (TOC) through generation, page number assignment, and iterative verification.

### Initial TOC Generation

The system first attempts to extract a physical TOC using `toc_extractor`. If none exists, the pipeline invokes `process_no_toc` (lines 68-84 in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)), which calls `generate_toc_init` and `generate_toc_continue` to prompt the LLM for a hierarchical structure based on the chunked content.

### Page Number Assignment

When the generated TOC lacks page numbers, `add_page_number_to_toc` (lines 53-84) iterates over each chunk and queries the LLM to identify where each section begins, populating the `physical_index` field for every entry.

### Verification and Correction

The `verify_toc` function (lines 92-145) samples generated entries and asks the LLM to confirm that each `physical_index` correctly points to its associated section. It returns an accuracy score and a list of incorrect entries.

For erroneous indices, `fix_incorrect_toc_with_retries` (lines 52-87) invokes `fix_incorrect_toc` repeatedly, narrowing the search range between valid neighboring indices to locate the correct start page. This process stops after a configurable number of attempts (default 3).

## Recursive Processing of Oversized Sections

Even after initial chunking, some sections may still exceed safe processing limits. PageIndex addresses this through recursive decomposition.

### The Large Node Detection Logic

In [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) (lines 991-1019), `process_large_node_recursively` detects nodes that violate size constraints—specifically those spanning more than `opt.max_page_num_each_node` (default 10 pages) **and** exceeding `opt.max_token_num_each_node` (default 20,000 tokens).

When such a node is detected, the function extracts the sub-document corresponding to that node and runs a **mini-page-index** on that slice. The original node is then replaced with the resulting deeper hierarchy. This recursion continues until every node satisfies both the page-count and token-count constraints.

### Configuration Parameters for Size Limits

The pipeline respects several configurable thresholds defined in [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml):

- **max_tokens**: Caps individual LLM requests (default 20,000)
- **overlap_page**: Ensures content continuity across chunks (default 1)
- **max_page_num_each_node**: Maximum pages allowed per TOC node before recursive splitting (default 10)
- **max_token_num_each_node**: Maximum tokens allowed per node (default 20,000)
- **toc_check_page_num**: Pages examined when searching for a physical TOC

Adjusting these values allows the same pipeline to handle even larger files without code modifications.

## Final Assembly and Output

Once all chunks are processed and large nodes recursively split, PageIndex assembles the final hierarchical structure.

### Tree Construction

The `post_processing` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 60-74) converts the flat list of sections into a nested tree using `list_to_tree`. If the document content begins after page 1, the system automatically prepends a *Preface* node to maintain structural integrity.

### Optional Enrichments

Following tree construction, optional post-processing steps in [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py) (lines 58-85) can add node IDs, raw text content, summaries, and single-sentence document descriptions. These enrichments occur **after** the TOC is built, meaning they do not affect the token limits enforced during the initial processing stages.

## Practical Configuration and Code Examples

### Basic Usage for Any PDF Size

The default configuration automatically handles documents exceeding context limits:

```python
from pageindex import page_index

result = page_index("path/to/your/large_document.pdf")
print(result["doc_name"])
print(result["structure"])   # nested dict/list representing the TOC tree

```

This invocation executes the complete pipeline—tokenization, chunking, TOC generation, verification, and recursive splitting—without requiring manual intervention.

### Tuning Limits for Extra-Large Books

For documents with unusually long chapters or sections, increase the thresholds:

```python
from pageindex import page_index

# Increase token budget per request to 40,000 and allow nodes up to 20 pages

result = page_index(
    "big_book.pdf",
    max_token_num_each_node=40000,
    max_page_num_each_node=20,
    toc_check_page_num=30   # examine more pages when looking for a physical TOC

)

print(result["structure"])

```

These parameters override the defaults from [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml), reducing the depth of recursive splitting required for large sections.

### Debugging Chunk Boundaries

To inspect how PageIndex segments your document:

```python
from pageindex.utils import get_page_tokens, page_list_to_group_text

pages = get_page_tokens("huge.pdf")

# Simulate the chunking step

chunks = page_list_to_group_text(
    page_contents=[p[0] for p in pages],
    token_lengths=[p[1] for p in pages],
    max_tokens=20000,
    overlap_page=1
)

print(f"Document split into {len(chunks)} chunks")
for i, chunk in enumerate(chunks, 1):
    print(f"Chunk {i} length: {len(chunk)} characters")

```

This reproduces the exact grouping logic used internally by `process_no_toc` and `process_toc_no_page_numbers`.

## Summary

- **Token-bounded chunking**: PageIndex converts PDF pages into token-counted units and groups them into chunks under 20,000 tokens using `page_list_to_group_text` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py).
- **Overlap protection**: A default one-page overlap between chunks ensures content spanning boundaries remains detectable.
- **Hierarchical verification**: The pipeline generates a TOC, assigns page numbers via `add_page_number_to_toc`, verifies accuracy with `verify_toc`, and corrects errors through `fix_incorrect_toc_with_retries`.
- **Recursive decomposition**: `process_large_node_recursively` (lines 991-1019) splits oversized sections into sub-hierarchies until all nodes fit within configured limits.
- **Configurable limits**: All thresholds—including `max_tokens`, `max_page_num_each_node`, and `max_token_num_each_node`—are adjustable via [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) or function parameters.

## Frequently Asked Questions

### How does PageIndex prevent content loss when splitting documents into chunks?

PageIndex uses an overlapping chunking strategy implemented in `page_list_to_group_text` ([`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), lines 418-452). By default, each chunk overlaps with the previous one by one page, ensuring that sections spanning chunk boundaries remain detectable in subsequent processing steps. This overlap guarantees that no structural information is lost at segmentation points.

### What happens if a single chapter exceeds the token limit after initial chunking?

When a node exceeds both the page count (`max_page_num_each_node`, default 10) and token count (`max_token_num_each_node`, default 20,000) thresholds, `process_large_node_recursively` ([`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), lines 991-1019) extracts that section and runs a mini-page-index on it. This recursive process continues until all resulting sub-nodes fit within the configured limits, creating a deeper hierarchy for lengthy chapters without breaking the context window.

### Can I adjust the token limits to work with different LLM context windows?

Yes. All token and page limits are configurable through [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) or directly via function parameters. You can modify `max_tokens` (default 20,000) to match your specific model's context window, or adjust `max_page_num_each_node` and `max_token_num_each_node` to control when recursive splitting occurs. The pipeline automatically respects these thresholds without requiring code modifications.

### How does PageIndex verify that generated page numbers are accurate?

The pipeline implements a verification loop using `verify_toc` ([`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), lines 92-145), which samples generated entries and asks the LLM to confirm that each `physical_index` correctly points to its associated section. If discrepancies are found, `fix_incorrect_toc_with_retries` (lines 52-87) narrows the search range between valid neighboring indices and attempts correction up to three times by default, ensuring high accuracy before final tree assembly.