# PageIndex Document Size Limits and Scaling Strategies: A Complete Guide

> Discover PageIndex's document size limits and scaling strategies. Learn how PageIndex handles large documents and optimize your LLM processing.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: how-to-guide
- Published: 2026-02-16

---

**PageIndex enforces hard limits of 10 pages and 20,000 tokens per node by default, automatically splitting oversized segments recursively to stay within LLM context windows while processing arbitrarily large documents.**

PageIndex, developed by VectifyAI, is designed to index very long professional documents without exceeding the context limitations of underlying language models. Understanding the document size limits and scaling mechanisms is essential for optimizing performance when processing multi-hundred-page PDFs. This guide examines the specific constraints implemented in the codebase and explains how to configure them for your use case.

## Understanding PageIndex Document Size Limits

PageIndex operates on a **node-based architecture**, where each node represents a contiguous range of pages that will be processed as a single unit by the LLM. To prevent context window overflows, the system enforces strict ceilings on both page count and token volume per node.

### Hard Limits Per Node

Two configurable hard limits govern individual node size:

- **Maximum pages per node**: Default of `10` pages, controlled via the CLI flag `--max-pages-per-node` (mapped to `opt.max_page_num_each_node` in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py))【1†L15-L19】 and enforced during node creation in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)【2†L996-L998】.

- **Maximum tokens per node**: Default of `20,000` tokens, set via `--max-tokens-per-node` (stored as `opt.max_token_num_each_node`)【1†L19-L21】 and checked in the same node-splitting logic【2†L996-L998】.

When either threshold is exceeded, PageIndex triggers an automatic splitting mechanism rather than attempting to process the oversized node.

### Global Token Safety Guard

Beyond the per-node limits, a **global safety ceiling** of `110,000` tokens prevents any single node from approaching the typical GPT-4o context window boundary. This guard is implemented in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) within the `check_token_limit` function【3†L533-L538】, providing a hard backstop even if per-node limits are configured generously.

## How PageIndex Handles Oversized Documents

Rather than failing on large inputs, PageIndex employs recursive decomposition strategies that break documents into LLM-compliant chunks while preserving semantic structure.

### Automatic Node Splitting

When a node exceeds the page or token limits, the system does not truncate content. Instead, it partitions the node into smaller child nodes covering subsets of the original page range. This logic resides in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), specifically within the node processing loop【2†L996-L998】.

### Recursive Processing

The function `process_large_node_recursively` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)【2†L991-L1010】 orchestrates the hierarchical decomposition. This asynchronous function:

1. Evaluates the current node against size limits
2. If limits are exceeded, spawns child nodes covering smaller page ranges
3. Recursively processes each child via `asyncio.gather` for parallel execution
4. Continues until all leaf nodes satisfy the configured constraints

This recursive approach enables PageIndex to index documents with hundreds of pages by transforming one massive LLM call into many smaller, parallelizable requests.

## Scaling PageIndex Processing

Several mechanisms allow users to optimize throughput and resource utilization when processing large document collections.

### Adjusting Per-Node Limits

Users can tune the trade-off between node granularity and API call volume by modifying the default limits. Lower values create more, smaller nodes (higher parallelism, more LLM calls), while higher values reduce API overhead but increase per-request latency.

Configuration options include:

- **CLI flags**: `--max-pages-per-node` and `--max-tokens-per-node` in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py)【1†L15-L21】
- **Python API**: `max_page_num_each_node` and `max_token_num_each_node` parameters in the `page_index` function exposed via [`pageindex/__init__.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/__init__.py)

```bash
python3 run_pageindex.py \
    --pdf_path my_report.pdf \
    --max-pages-per-node 20 \
    --max-tokens-per-node 40000

```

### Parallelism and Async Processing

The core engine leverages Python's `asyncio` for concurrent node processing. When `process_large_node_recursively` splits a node, it uses `asyncio.gather` to process child nodes simultaneously【2†L991-L1010】.

Additionally, I/O-bound operations (PDF tokenization, HTTP requests to LLM APIs) utilize `ThreadPoolExecutor` to prevent blocking the async event loop. This dual-layer concurrency ensures CPU and network resources are fully utilized during large batch jobs.

### Chunk-Based Tokenization

Before LLM processing, the helper function `page_list_to_group_text` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) groups pages into token-aware blocks【2†L18-L45】. This function:

- Defaults to `max_tokens=20000` per block
- Uses an arithmetic-based heuristic to estimate token counts and prevent overflow
- Implements `overlap_page=1` to maintain contextual continuity across block boundaries

This preprocessing ensures that input to the recursive processor is already optimally sized, reducing the frequency of expensive splitting operations.

### Configuration-Driven Defaults

Default scaling parameters are centralized in [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml)【4†L1-L5】, allowing repository-wide policies (such as stricter token ceilings for cost-sensitive deployments) without code changes. The `ConfigLoader` class in [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py) ingests these values, which are then overridden by CLI flags or Python API arguments at runtime.

## Code Examples

### Using the CLI with Custom Limits

Process a 250-page PDF with expanded node sizes to reduce API call volume:

```bash
python3 run_pageindex.py \
    --pdf_path ./tests/pdfs/2023-annual-report.pdf \
    --max-pages-per-node 15 \
    --max-tokens-per-node 30000

```

### Calling page_index Programmatically

Override limits directly when using the Python API:

```python
from pageindex import page_index

result = page_index(
    doc="big_document.pdf",
    max_page_num_each_node=20,          # allow up to 20 pages per node

    max_token_num_each_node=40000,      # allow up to 40k tokens per node

    if_add_node_summary="yes",          # optional extra output

)

print(result["doc_name"])
print(result["structure"])   # hierarchical tree respecting the limits

```

### Inspecting Effective Limits at Runtime

Verify configuration values during execution:

```python
import json
from pageindex.utils import ConfigLoader

loader = ConfigLoader()
opt = loader.load()
print(f"Pages per node: {opt.max_page_num_each_node}")
print(f"Tokens per node: {opt.max_token_num_each_node}")

```

### Advanced Parallel Processing

Manually invoke the recursive processor for custom workflows:

```python
import asyncio
from pageindex.page_index import process_large_node_recursively, get_page_tokens

async def index_subtrees(pdf_path):
    page_list = get_page_tokens(pdf_path)
    
    top_node = {"title": "ROOT", "start_index": 1,
                "end_index": len(page_list), "nodes": []}
    
    await process_large_node_recursively(top_node, page_list)
    return top_node

tree = asyncio.run(index_subtrees("huge.pdf"))

```

## Summary

- **PageIndex enforces dual hard limits**: 10 pages and 20,000 tokens per node by default, with a global safety ceiling of 110,000 tokens to protect against context window overflow.
- **Automatic recursive splitting**: The `process_large_node_recursively` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) breaks oversized nodes into smaller child nodes until all segments comply with configured limits.
- **Scalable architecture**: Async processing via `asyncio.gather`, `ThreadPoolExecutor` for I/O operations, and chunk-based tokenization enable efficient handling of multi-hundred-page documents.
- **Flexible configuration**: Adjust per-node thresholds via CLI flags (`--max-pages-per-node`, `--max-tokens-per-node`), Python API parameters, or the [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) file to balance granularity against API call volume.

## Frequently Asked Questions

### What happens if a PDF exceeds the default page or token limits?

PageIndex does not reject large documents. Instead, it automatically invokes the `process_large_node_recursively` function to split the document into smaller child nodes. This recursion continues until every node satisfies the configured page and token limits, ensuring the LLM never receives a request exceeding its context window.

### Can I process documents larger than 110,000 tokens?

Yes. The 110,000-token limit in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) is a safety guard for individual nodes, not a document-wide restriction. By recursively splitting content into smaller segments (each under the limit), PageIndex can index documents with millions of tokens. Simply ensure your per-node limits are set appropriately via `--max-tokens-per-node` or the [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) file.

### How do I reduce the number of LLM API calls when processing large documents?

Increase the `--max-pages-per-node` and `--max-tokens-per-node` values to allow larger segments per API request. For example, raising the token limit from 20,000 to 40,000 halves the number of nodes (and thus API calls) for a given document. However, ensure these values remain well below your LLM's context window to avoid truncation errors.

### Where are the default size limits configured?

Default values reside in [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml)【4†L1-L5】 and are loaded at runtime by the `ConfigLoader` class in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py). These defaults are overridden by CLI arguments in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) or by passing parameters directly to the `page_index` function in [`pageindex/__init__.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/__init__.py).