Understanding max_page_num_each_node and max_token_num_each_node in PageIndex

The max_page_num_each_node and max_token_num_each_node parameters act as safety guards that limit how much content each node in the document tree can contain, preventing LLM token limit errors while preserving logical document hierarchy.

When processing large PDF documents with VectifyAI/PageIndex, you need to balance comprehensive indexing against LLM API constraints. The max_page_num_each_node and max_token_num_each_node configuration parameters control exactly how the library segments your documents into manageable chunks. These settings ensure that no single node in the generated tree structure grows too large for downstream processing tasks like summarization and title verification.

What Do These PageIndex Configuration Parameters Control?

Parameter Default Value Purpose
max_page_num_each_node 10 Maximum number of PDF pages allowed per node
max_token_num_each_node 20000 Maximum number of LLM tokens allowed per node's aggregated content

Both parameters work as hard limits during the recursive tree construction phase. When a node exceeds either threshold, the system triggers a split operation to maintain the constraints.

How the Code Implements These Limits

Loading Configuration Defaults

The ConfigLoader class in pageindex/utils.py handles parameter initialization. It merges default values from pageindex/config.yaml with any user-provided overrides:


# From pageindex/utils.py lines 81-112

def load(self):
    self._default_dict = self._load_yaml(self.default_config_path)
    # Merge with user_dict for overrides

By default, the system reads max_page_num_each_node: 10 and max_token_num_each_node: 20000 from the YAML configuration.

Detecting Oversized Nodes

During tree construction, the process_large_node_recursively function in pageindex/page_index.py checks both limits before proceeding with node processing:


# From pageindex/page_index.py lines 996-1000

if node['end_index'] - node['start_index'] > opt.max_page_num_each_node \
    and token_num >= opt.max_token_num_each_node:
    # Trigger further splitting

This condition ensures that nodes exceeding both the page count and token count thresholds undergo additional segmentation.

Splitting Nodes Recursively

When the limits are exceeded, the function re-processes the node's pages with fresh table-of-contents extraction, then builds sub-nodes until every child satisfies the constraints. This recursive splitting prevents any single node from growing large enough to cause LLM request failures during summary generation or title verification.

CLI Overrides

Users can override these defaults via command-line arguments defined in run_pageindex.py:

python run_pageindex.py --pdf_path document.pdf \
    --max-pages-per-node 15 \
    --max-tokens-per-node 25000

The argument definitions appear in run_pageindex.py at lines 15-20, mapping CLI flags to the configuration parameters.

Why These Limits Matter for LLM Processing

  • API Token Caps: Most LLM APIs enforce maximum token limits per request. If a node exceeds max_token_num_each_node, summarization calls truncate or fail entirely.
  • Cost Optimization: Smaller nodes reduce token consumption, lowering API costs while maintaining processing speed.
  • Hierarchical Fidelity: Forcing splits on large sections preserves the document's logical outline structure, ensuring the generated tree accurately reflects the original PDF's organization.

Practical Usage Examples

Python API Configuration

from pageindex import page_index

# Use default limits (10 pages, 20,000 tokens per node)

structure = page_index("technical_manual.pdf")

# Override for dense academic papers with small fonts

structure = page_index(
    "research_paper.pdf",
    max_page_num_each_node=20,
    max_token_num_each_node=30000
)

Command-Line Configuration


# Process a lengthy legal document with expanded limits

python run_pageindex.py \
    --pdf_path contract.pdf \
    --max-pages-per-node 25 \
    --max-tokens-per-node 35000

Key Source Files to Explore

File Relevance
pageindex/config.yaml Default values for max_page_num_each_node and max_token_num_each_node
pageindex/page_index.py Core logic at lines 996-1000 where limits are enforced during node splitting
run_pageindex.py CLI argument definitions (lines 15-20) for runtime overrides
pageindex/utils.py ConfigLoader implementation (lines 81-112) handling configuration merging

These components together define how PageIndex balances document size, LLM token constraints, and hierarchical fidelity through the max_page_num_each_node and max_token_num_each_node settings.

Summary

  • The max_page_num_each_node and max_token_num_each_node parameters prevent individual nodes from growing too large for LLM processing.
  • Defaults are set in config.yaml (10 pages and 20,000 tokens) and enforced in page_index.py during recursive tree construction.
  • Users can override these limits via Python API arguments or CLI flags --max-pages-per-node and --max-tokens-per-node.
  • Proper tuning ensures API requests stay within token limits while preserving the document's logical hierarchy.

Frequently Asked Questions

What happens if a node exceeds both max_page_num_each_node and max_token_num_each_node?

When both thresholds are exceeded, the process_large_node_recursively function in pageindex/page_index.py triggers a recursive split. The system re-processes the node's pages with fresh table-of-contents extraction and creates sub-nodes until every child satisfies both constraints.

Can I set max_page_num_each_node and max_token_num_each_node to unlimited values?

While you can set high values via the CLI or Python API, doing so risks LLM API failures. Most LLM providers enforce hard token limits (typically 4K-128K depending on the model). Exceeding max_token_num_each_node beyond your LLM's context window will cause summarization and verification calls to fail or truncate.

How do max_page_num_each_node and max_token_num_each_node interact with the document hierarchy?

These parameters act as safety guards rather than primary structuring logic. The system first attempts to build nodes based on the document's table of contents and logical sections. Only when a natural section exceeds either limit does the recursive splitting activate, ensuring the hierarchy remains as close to the original document structure as possible while respecting technical constraints.

Where are the default values for these parameters defined?

Default values are defined in pageindex/config.yaml where max_page_num_each_node defaults to 10 and max_token_num_each_node defaults to 20000. The ConfigLoader class in pageindex/utils.py (lines 81-112) loads these defaults and merges them with any user-provided configuration overrides.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →