# PageIndex Internal Tree Structure Format: Schema, Generation, and Implementation

> Explore PageIndex's internal tree structure format, learn how it generates nested JSON-like documents using PDF TOC parsers or Markdown analyzers, and implement its node metadata for efficient document representation.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: internals
- Published: 2026-02-16

---

**PageIndex represents documents as nested JSON-like trees where each node contains metadata fields including title, node_id, page indices, and a recursive nodes array, generated through either a PDF TOC parser or a Markdown heading analyzer.**

PageIndex is an open-source Python library developed by VectifyAI that transforms flat documents into navigable hierarchies. Understanding the internal tree structure format is crucial for developers building retrieval-augmented generation (RAG) systems or document analysis agents. This guide examines the exact node schema, field definitions, and the dual-generation pipelines that construct these hierarchical trees from both PDF and Markdown sources.

## Understanding the PageIndex Tree Structure Schema

The internal tree structure format in PageIndex follows a recursive dictionary schema. Each document becomes a list of root-level nodes, where each node represents a section or subsection with standardized metadata fields.

### Core Node Fields

Every node in the tree is a Python dictionary containing the following fields:

- **title**: The exact section heading text extracted from the document (e.g., `"Introduction"`)
- **node_id**: A unique, zero-padded four-digit identifier (e.g., `"0001"`)
- **start_index** / **end_index**: Integer page numbers indicating the section's physical span (added during post-processing)
- **physical_index**: Raw page tag extracted from the PDF (e.g., `12`), later converted to an integer
- **text**: Full text content of the section (optional, added on demand)
- **line_num**: Line number in the original Markdown file (Markdown mode only)
- **nodes**: Recursive list of child nodes following the same schema

### Top-Level Container Structure

The final output is either a raw list of root nodes or a dictionary wrapper containing a `structure` key (as returned by the high-level API). This flexibility allows PageIndex to interface with both direct API consumers and LLM-friendly JSON outputs.

## How PageIndex Generates the Internal Tree Structure

PageIndex implements two distinct pipelines for tree generation, depending on the source document type: PDF documents with Table of Contents (TOC) metadata, and Markdown files with heading hierarchies.

### PDF Pipeline: From Flat TOC to Nested Hierarchy

The PDF pipeline processes documents containing embedded TOC information. This workflow operates in three phases:

1. **Extraction**: The `toc_extractor` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) parses the PDF to identify TOC pages and extracts a flat list of entries. Each entry contains a `structure` field (e.g., `"1.2.3"`), `title`, and raw `physical_index` tags.

2. **Tree Construction**: The `list_to_tree` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) converts the flat list into the nested hierarchy. It uses the dot-notation in the `structure` field to infer parent-child relationships. The algorithm splits the structure code, drops the last segment to identify the parent, and appends the node to the parent's `nodes` list.

3. **Post-Processing**: The `post_processing` function adds `start_index` and `end_index` fields based on the physical page spans and optionally removes empty `nodes` arrays.

### Markdown Pipeline: Stack-Based Heading Parsing

For Markdown documents, PageIndex builds the tree by analyzing heading levels directly. This pipeline is implemented in [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py):

1. **Node Extraction**: The `extract_nodes_from_markdown` function scans the file for `#` headings, recording the heading level (number of `#` symbols) and line number.

2. **Text Attachment**: The `extract_node_text_content` function associates the subsequent text blocks with each heading node.

3. **Tree Building**: The `build_tree_from_nodes` function uses a **stack-based algorithm** to establish parent-child relationships. It maintains a stack of `(node, level)` tuples. When processing a new heading, it pops entries until finding a node with a lower heading level, then appends the new node to that parent's `nodes` list.

4. **Formatting**: The `format_structure` function reorders fields and prunes empty children arrays before final output.

## Key Implementation Details

Understanding the specific implementation mechanics helps developers customize PageIndex for specialized document formats.

### Parent Detection via Dot-Notation Codes

The core logic for determining parent-child relationships in PDF processing resides in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py). The `get_parent_structure` helper parses dot-notation structure codes:

```python

# pageindex/utils.py – core of list_to_tree

def get_parent_structure(structure):
    if not structure:
        return None
    parts = str(structure).split('.')
    return '.'.join(parts[:-1]) if len(parts) > 1 else None

```

This function enables the algorithm to build the hierarchy without requiring recursive parsing of the entire document text.

### Stack-Based Nesting Algorithm

The Markdown tree builder in [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) implements an efficient stack algorithm:

```python

# pageindex/page_index_md.py – building the tree

while stack and stack[-1][1] >= current_level:
    stack.pop()
if not stack:
    root_nodes.append(tree_node)
else:
    parent_node, _ = stack[-1]
    parent_node['nodes'].append(tree_node)
stack.append((tree_node, current_level))

```

This approach ensures O(n) complexity for tree construction from n headings, avoiding expensive recursive lookups.

### Node ID Assignment and Field Customization

The `write_node_id` function (called from `md_to_tree` and similar PDF functions) walks the completed tree and assigns sequential, zero-padded four-digit identifiers (e.g., "0001", "0002"), guaranteeing deterministic references for downstream applications.

Additionally, the `format_structure` function can be instructed to keep or drop fields (`summary`, `text`, etc.) by passing an ordered list, which is useful for compact output or for feeding only the hierarchy to another model.

## Practical Code Examples

### Generating Trees from PDF Documents

To process a PDF with embedded TOC data:

```python
from pageindex.page_index import check_toc, process_toc_with_page_numbers

# Assume `pdf_pages` is a list of (text, token_len) tuples from utils.get_page_tokens()

toc_info = check_toc(pdf_pages, opt)                # Finds TOC pages

tree = process_toc_with_page_numbers(
    toc_content=toc_info['toc_content'],
    toc_page_list=toc_info['toc_page_list'],
    page_list=pdf_pages,
    toc_check_page_num=opt.toc_check_page_num,
    model=opt.model,
    logger=opt.logger,
)

# `tree` is a list of dicts with the fields described above

print(tree[0])

```

*Source:* `check_toc` and `process_toc_with_page_numbers` in **pageindex/page_index.py** ([lines 88‑124](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L88-L124))

### Processing Markdown Files

For Markdown documents, use the async `md_to_tree` function:

```python
import asyncio
from pageindex.page_index_md import md_to_tree

tree = asyncio.run(md_to_tree(
    md_path='samples/example.md',
    if_thinning=False,
    if_add_node_summary='yes',
    summary_token_threshold=200,
    model='gpt-4o',
))

# The output is a dict: {'doc_name': 'example', 'structure': [...]}

print(tree['structure'][0])

```

*Source:* `md_to_tree` in **pageindex/page_index_md.py** ([lines 259‑285](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py#L259-L285))

### Converting Flat TOC Lists Manually

To manually convert a flat list of TOC entries (with `structure` fields like "1.2.3"):

```python
from pageindex.utils import list_to_tree

# `flat_toc` is the list returned from toc_extractor

tree = list_to_tree(flat_toc)
print(tree[0])          # First top‑level section

```

*Source:* `list_to_tree` implementation in **pageindex/utils.py** ([lines 50‑85](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py#L50-L85))

## Summary

- PageIndex uses a **recursive JSON-like structure** where each node contains metadata (title, node_id, page indices) and a nested `nodes` array for children.
- The **PDF pipeline** extracts a flat TOC list, then uses `list_to_tree` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) to build hierarchies based on dot-notation structure codes (e.g., "1.2.3").
- The **Markdown pipeline** parses heading levels directly, using a **stack-based algorithm** in `build_tree_from_nodes` to establish parent-child relationships in O(n) time.
- Both pipelines assign **zero-padded four-digit node IDs** for deterministic referencing and support optional fields like `text` content and page span indices.

## Frequently Asked Questions

### What is the exact schema of the PageIndex internal tree structure?

The PageIndex internal tree structure format consists of a list of dictionaries, where each dictionary represents a document section. Required fields include `title` (the heading text), `node_id` (a unique zero-padded four-digit identifier), and `nodes` (a list of child nodes). Optional fields include `start_index` and `end_index` for page spans, `physical_index` for raw PDF page tags, `text` for section content, and `line_num` for Markdown source references.

### How does PageIndex determine parent-child relationships in PDF documents?

For PDF documents, PageIndex uses the `structure` field from the Table of Contents, which contains dot-notation codes like "1.2.3". The `list_to_tree` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) splits these codes to identify parents—dropping the last segment yields the parent code (e.g., "1.2.3" becomes "1.2"). Nodes are then appended to their respective parent's `nodes` list, building the hierarchy recursively without parsing the full document text.

### Why does the Markdown pipeline use a stack-based algorithm?

The Markdown pipeline uses a stack-based approach in `build_tree_from_nodes` (located in [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py)) to achieve O(n) linear complexity when processing n headings. The algorithm maintains a stack of `(node, level)` tuples, popping entries until it finds a node with a lower heading level than the current one, then appending the new node to that parent's `nodes` list. This avoids expensive recursive lookups and efficiently handles arbitrarily deep nesting.

### Can I customize which fields are included in the final tree output?

Yes, PageIndex allows field customization through the `format_structure` function available in both pipelines. You can pass an ordered list of field names to include or exclude specific metadata such as `summary`, `text`, or `physical_index`. This is particularly useful when optimizing the output size for LLM context windows or when only the hierarchy (without content) is required for downstream processing.