# Understanding max_page_num_each_node and max_token_num_each_node in PageIndex

> Learn how PageIndex configures max_page_num_each_node and max_token_num_each_node to prevent LLM token limit errors and maintain document hierarchy. Optimize your content management.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: deep-dive
- Published: 2026-02-16

---

**The `max_page_num_each_node` and `max_token_num_each_node` parameters act as safety guards that limit how much content each node in the document tree can contain, preventing LLM token limit errors while preserving logical document hierarchy.**

When processing large PDF documents with VectifyAI/PageIndex, you need to balance comprehensive indexing against LLM API constraints. The `max_page_num_each_node` and `max_token_num_each_node` configuration parameters control exactly how the library segments your documents into manageable chunks. These settings ensure that no single node in the generated tree structure grows too large for downstream processing tasks like summarization and title verification.

## What Do These PageIndex Configuration Parameters Control?

| Parameter | Default Value | Purpose |
|-----------|--------------|---------|
| **`max_page_num_each_node`** | `10` | Maximum number of PDF pages allowed per node |
| **`max_token_num_each_node`** | `20000` | Maximum number of LLM tokens allowed per node's aggregated content |

Both parameters work as **hard limits** during the recursive tree construction phase. When a node exceeds either threshold, the system triggers a split operation to maintain the constraints.

## How the Code Implements These Limits

### Loading Configuration Defaults

The `ConfigLoader` class in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) handles parameter initialization. It merges default values from [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) with any user-provided overrides:

```python

# From pageindex/utils.py lines 81-112

def load(self):
    self._default_dict = self._load_yaml(self.default_config_path)
    # Merge with user_dict for overrides

```

By default, the system reads `max_page_num_each_node: 10` and `max_token_num_each_node: 20000` from the YAML configuration.

### Detecting Oversized Nodes

During tree construction, the `process_large_node_recursively` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) checks both limits before proceeding with node processing:

```python

# From pageindex/page_index.py lines 996-1000

if node['end_index'] - node['start_index'] > opt.max_page_num_each_node \
    and token_num >= opt.max_token_num_each_node:
    # Trigger further splitting

```

This condition ensures that nodes exceeding **both** the page count and token count thresholds undergo additional segmentation.

### Splitting Nodes Recursively

When the limits are exceeded, the function re-processes the node's pages with fresh table-of-contents extraction, then builds sub-nodes until every child satisfies the constraints. This recursive splitting prevents any single node from growing large enough to cause LLM request failures during summary generation or title verification.

### CLI Overrides

Users can override these defaults via command-line arguments defined in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py):

```bash
python run_pageindex.py --pdf_path document.pdf \
    --max-pages-per-node 15 \
    --max-tokens-per-node 25000

```

The argument definitions appear in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) at lines 15-20, mapping CLI flags to the configuration parameters.

## Why These Limits Matter for LLM Processing

- **API Token Caps**: Most LLM APIs enforce maximum token limits per request. If a node exceeds `max_token_num_each_node`, summarization calls truncate or fail entirely.
- **Cost Optimization**: Smaller nodes reduce token consumption, lowering API costs while maintaining processing speed.
- **Hierarchical Fidelity**: Forcing splits on large sections preserves the document's logical outline structure, ensuring the generated tree accurately reflects the original PDF's organization.

## Practical Usage Examples

### Python API Configuration

```python
from pageindex import page_index

# Use default limits (10 pages, 20,000 tokens per node)

structure = page_index("technical_manual.pdf")

# Override for dense academic papers with small fonts

structure = page_index(
    "research_paper.pdf",
    max_page_num_each_node=20,
    max_token_num_each_node=30000
)

```

### Command-Line Configuration

```bash

# Process a lengthy legal document with expanded limits

python run_pageindex.py \
    --pdf_path contract.pdf \
    --max-pages-per-node 25 \
    --max-tokens-per-node 35000

```

## Key Source Files to Explore

| File | Relevance |
|------|-----------|
| [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) | Default values for `max_page_num_each_node` and `max_token_num_each_node` |
| [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) | Core logic at lines 996-1000 where limits are enforced during node splitting |
| [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) | CLI argument definitions (lines 15-20) for runtime overrides |
| [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) | `ConfigLoader` implementation (lines 81-112) handling configuration merging |

These components together define how **PageIndex** balances document size, LLM token constraints, and hierarchical fidelity through the `max_page_num_each_node` and `max_token_num_each_node` settings.

## Summary

- The `max_page_num_each_node` and `max_token_num_each_node` parameters prevent individual nodes from growing too large for LLM processing.
- Defaults are set in [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) (10 pages and 20,000 tokens) and enforced in [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py) during recursive tree construction.
- Users can override these limits via Python API arguments or CLI flags `--max-pages-per-node` and `--max-tokens-per-node`.
- Proper tuning ensures API requests stay within token limits while preserving the document's logical hierarchy.

## Frequently Asked Questions

### What happens if a node exceeds both max_page_num_each_node and max_token_num_each_node?

When both thresholds are exceeded, the `process_large_node_recursively` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) triggers a recursive split. The system re-processes the node's pages with fresh table-of-contents extraction and creates sub-nodes until every child satisfies both constraints.

### Can I set max_page_num_each_node and max_token_num_each_node to unlimited values?

While you can set high values via the CLI or Python API, doing so risks LLM API failures. Most LLM providers enforce hard token limits (typically 4K-128K depending on the model). Exceeding `max_token_num_each_node` beyond your LLM's context window will cause summarization and verification calls to fail or truncate.

### How do max_page_num_each_node and max_token_num_each_node interact with the document hierarchy?

These parameters act as safety guards rather than primary structuring logic. The system first attempts to build nodes based on the document's table of contents and logical sections. Only when a natural section exceeds either limit does the recursive splitting activate, ensuring the hierarchy remains as close to the original document structure as possible while respecting technical constraints.

### Where are the default values for these parameters defined?

Default values are defined in [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) where `max_page_num_each_node` defaults to `10` and `max_token_num_each_node` defaults to `20000`. The `ConfigLoader` class in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 81-112) loads these defaults and merges them with any user-provided configuration overrides.