PageIndex Document Size Limits and Scaling Strategies: A Complete Guide

PageIndex enforces hard limits of 10 pages and 20,000 tokens per node by default, automatically splitting oversized segments recursively to stay within LLM context windows while processing arbitrarily large documents.

PageIndex, developed by VectifyAI, is designed to index very long professional documents without exceeding the context limitations of underlying language models. Understanding the document size limits and scaling mechanisms is essential for optimizing performance when processing multi-hundred-page PDFs. This guide examines the specific constraints implemented in the codebase and explains how to configure them for your use case.

Understanding PageIndex Document Size Limits

PageIndex operates on a node-based architecture, where each node represents a contiguous range of pages that will be processed as a single unit by the LLM. To prevent context window overflows, the system enforces strict ceilings on both page count and token volume per node.

Hard Limits Per Node

Two configurable hard limits govern individual node size:

  • Maximum pages per node: Default of 10 pages, controlled via the CLI flag --max-pages-per-node (mapped to opt.max_page_num_each_node in run_pageindex.py)【1†L15-L19】 and enforced during node creation in pageindex/page_index.py【2†L996-L998】.

  • Maximum tokens per node: Default of 20,000 tokens, set via --max-tokens-per-node (stored as opt.max_token_num_each_node)【1†L19-L21】 and checked in the same node-splitting logic【2†L996-L998】.

When either threshold is exceeded, PageIndex triggers an automatic splitting mechanism rather than attempting to process the oversized node.

Global Token Safety Guard

Beyond the per-node limits, a global safety ceiling of 110,000 tokens prevents any single node from approaching the typical GPT-4o context window boundary. This guard is implemented in pageindex/utils.py within the check_token_limit function【3†L533-L538】, providing a hard backstop even if per-node limits are configured generously.

How PageIndex Handles Oversized Documents

Rather than failing on large inputs, PageIndex employs recursive decomposition strategies that break documents into LLM-compliant chunks while preserving semantic structure.

Automatic Node Splitting

When a node exceeds the page or token limits, the system does not truncate content. Instead, it partitions the node into smaller child nodes covering subsets of the original page range. This logic resides in pageindex/page_index.py, specifically within the node processing loop【2†L996-L998】.

Recursive Processing

The function process_large_node_recursively in pageindex/page_index.py【2†L991-L1010】 orchestrates the hierarchical decomposition. This asynchronous function:

  1. Evaluates the current node against size limits
  2. If limits are exceeded, spawns child nodes covering smaller page ranges
  3. Recursively processes each child via asyncio.gather for parallel execution
  4. Continues until all leaf nodes satisfy the configured constraints

This recursive approach enables PageIndex to index documents with hundreds of pages by transforming one massive LLM call into many smaller, parallelizable requests.

Scaling PageIndex Processing

Several mechanisms allow users to optimize throughput and resource utilization when processing large document collections.

Adjusting Per-Node Limits

Users can tune the trade-off between node granularity and API call volume by modifying the default limits. Lower values create more, smaller nodes (higher parallelism, more LLM calls), while higher values reduce API overhead but increase per-request latency.

Configuration options include:

  • CLI flags: --max-pages-per-node and --max-tokens-per-node in run_pageindex.py【1†L15-L21】
  • Python API: max_page_num_each_node and max_token_num_each_node parameters in the page_index function exposed via pageindex/__init__.py
python3 run_pageindex.py \
    --pdf_path my_report.pdf \
    --max-pages-per-node 20 \
    --max-tokens-per-node 40000

Parallelism and Async Processing

The core engine leverages Python's asyncio for concurrent node processing. When process_large_node_recursively splits a node, it uses asyncio.gather to process child nodes simultaneously【2†L991-L1010】.

Additionally, I/O-bound operations (PDF tokenization, HTTP requests to LLM APIs) utilize ThreadPoolExecutor to prevent blocking the async event loop. This dual-layer concurrency ensures CPU and network resources are fully utilized during large batch jobs.

Chunk-Based Tokenization

Before LLM processing, the helper function page_list_to_group_text in pageindex/page_index.py groups pages into token-aware blocks【2†L18-L45】. This function:

  • Defaults to max_tokens=20000 per block
  • Uses an arithmetic-based heuristic to estimate token counts and prevent overflow
  • Implements overlap_page=1 to maintain contextual continuity across block boundaries

This preprocessing ensures that input to the recursive processor is already optimally sized, reducing the frequency of expensive splitting operations.

Configuration-Driven Defaults

Default scaling parameters are centralized in pageindex/config.yaml【4†L1-L5】, allowing repository-wide policies (such as stricter token ceilings for cost-sensitive deployments) without code changes. The ConfigLoader class in utils.py ingests these values, which are then overridden by CLI flags or Python API arguments at runtime.

Code Examples

Using the CLI with Custom Limits

Process a 250-page PDF with expanded node sizes to reduce API call volume:

python3 run_pageindex.py \
    --pdf_path ./tests/pdfs/2023-annual-report.pdf \
    --max-pages-per-node 15 \
    --max-tokens-per-node 30000

Calling page_index Programmatically

Override limits directly when using the Python API:

from pageindex import page_index

result = page_index(
    doc="big_document.pdf",
    max_page_num_each_node=20,          # allow up to 20 pages per node

    max_token_num_each_node=40000,      # allow up to 40k tokens per node

    if_add_node_summary="yes",          # optional extra output

)

print(result["doc_name"])
print(result["structure"])   # hierarchical tree respecting the limits

Inspecting Effective Limits at Runtime

Verify configuration values during execution:

import json
from pageindex.utils import ConfigLoader

loader = ConfigLoader()
opt = loader.load()
print(f"Pages per node: {opt.max_page_num_each_node}")
print(f"Tokens per node: {opt.max_token_num_each_node}")

Advanced Parallel Processing

Manually invoke the recursive processor for custom workflows:

import asyncio
from pageindex.page_index import process_large_node_recursively, get_page_tokens

async def index_subtrees(pdf_path):
    page_list = get_page_tokens(pdf_path)
    
    top_node = {"title": "ROOT", "start_index": 1,
                "end_index": len(page_list), "nodes": []}
    
    await process_large_node_recursively(top_node, page_list)
    return top_node

tree = asyncio.run(index_subtrees("huge.pdf"))

Summary

  • PageIndex enforces dual hard limits: 10 pages and 20,000 tokens per node by default, with a global safety ceiling of 110,000 tokens to protect against context window overflow.
  • Automatic recursive splitting: The process_large_node_recursively function in pageindex/page_index.py breaks oversized nodes into smaller child nodes until all segments comply with configured limits.
  • Scalable architecture: Async processing via asyncio.gather, ThreadPoolExecutor for I/O operations, and chunk-based tokenization enable efficient handling of multi-hundred-page documents.
  • Flexible configuration: Adjust per-node thresholds via CLI flags (--max-pages-per-node, --max-tokens-per-node), Python API parameters, or the config.yaml file to balance granularity against API call volume.

Frequently Asked Questions

What happens if a PDF exceeds the default page or token limits?

PageIndex does not reject large documents. Instead, it automatically invokes the process_large_node_recursively function to split the document into smaller child nodes. This recursion continues until every node satisfies the configured page and token limits, ensuring the LLM never receives a request exceeding its context window.

Can I process documents larger than 110,000 tokens?

Yes. The 110,000-token limit in pageindex/utils.py is a safety guard for individual nodes, not a document-wide restriction. By recursively splitting content into smaller segments (each under the limit), PageIndex can index documents with millions of tokens. Simply ensure your per-node limits are set appropriately via --max-tokens-per-node or the config.yaml file.

How do I reduce the number of LLM API calls when processing large documents?

Increase the --max-pages-per-node and --max-tokens-per-node values to allow larger segments per API request. For example, raising the token limit from 20,000 to 40,000 halves the number of nodes (and thus API calls) for a given document. However, ensure these values remain well below your LLM's context window to avoid truncation errors.

Where are the default size limits configured?

Default values reside in pageindex/config.yaml【4†L1-L5】 and are loaded at runtime by the ConfigLoader class in pageindex/utils.py. These defaults are overridden by CLI arguments in run_pageindex.py or by passing parameters directly to the page_index function in pageindex/__init__.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →