PageIndex Internal Tree Structure Format: Schema, Generation, and Implementation

PageIndex represents documents as nested JSON-like trees where each node contains metadata fields including title, node_id, page indices, and a recursive nodes array, generated through either a PDF TOC parser or a Markdown heading analyzer.

PageIndex is an open-source Python library developed by VectifyAI that transforms flat documents into navigable hierarchies. Understanding the internal tree structure format is crucial for developers building retrieval-augmented generation (RAG) systems or document analysis agents. This guide examines the exact node schema, field definitions, and the dual-generation pipelines that construct these hierarchical trees from both PDF and Markdown sources.

Understanding the PageIndex Tree Structure Schema

The internal tree structure format in PageIndex follows a recursive dictionary schema. Each document becomes a list of root-level nodes, where each node represents a section or subsection with standardized metadata fields.

Core Node Fields

Every node in the tree is a Python dictionary containing the following fields:

  • title: The exact section heading text extracted from the document (e.g., "Introduction")
  • node_id: A unique, zero-padded four-digit identifier (e.g., "0001")
  • start_index / end_index: Integer page numbers indicating the section's physical span (added during post-processing)
  • physical_index: Raw page tag extracted from the PDF (e.g., 12), later converted to an integer
  • text: Full text content of the section (optional, added on demand)
  • line_num: Line number in the original Markdown file (Markdown mode only)
  • nodes: Recursive list of child nodes following the same schema

Top-Level Container Structure

The final output is either a raw list of root nodes or a dictionary wrapper containing a structure key (as returned by the high-level API). This flexibility allows PageIndex to interface with both direct API consumers and LLM-friendly JSON outputs.

How PageIndex Generates the Internal Tree Structure

PageIndex implements two distinct pipelines for tree generation, depending on the source document type: PDF documents with Table of Contents (TOC) metadata, and Markdown files with heading hierarchies.

PDF Pipeline: From Flat TOC to Nested Hierarchy

The PDF pipeline processes documents containing embedded TOC information. This workflow operates in three phases:

  1. Extraction: The toc_extractor function in pageindex/page_index.py parses the PDF to identify TOC pages and extracts a flat list of entries. Each entry contains a structure field (e.g., "1.2.3"), title, and raw physical_index tags.

  2. Tree Construction: The list_to_tree function in pageindex/utils.py converts the flat list into the nested hierarchy. It uses the dot-notation in the structure field to infer parent-child relationships. The algorithm splits the structure code, drops the last segment to identify the parent, and appends the node to the parent's nodes list.

  3. Post-Processing: The post_processing function adds start_index and end_index fields based on the physical page spans and optionally removes empty nodes arrays.

Markdown Pipeline: Stack-Based Heading Parsing

For Markdown documents, PageIndex builds the tree by analyzing heading levels directly. This pipeline is implemented in pageindex/page_index_md.py:

  1. Node Extraction: The extract_nodes_from_markdown function scans the file for # headings, recording the heading level (number of # symbols) and line number.

  2. Text Attachment: The extract_node_text_content function associates the subsequent text blocks with each heading node.

  3. Tree Building: The build_tree_from_nodes function uses a stack-based algorithm to establish parent-child relationships. It maintains a stack of (node, level) tuples. When processing a new heading, it pops entries until finding a node with a lower heading level, then appends the new node to that parent's nodes list.

  4. Formatting: The format_structure function reorders fields and prunes empty children arrays before final output.

Key Implementation Details

Understanding the specific implementation mechanics helps developers customize PageIndex for specialized document formats.

Parent Detection via Dot-Notation Codes

The core logic for determining parent-child relationships in PDF processing resides in pageindex/utils.py. The get_parent_structure helper parses dot-notation structure codes:


# pageindex/utils.py – core of list_to_tree

def get_parent_structure(structure):
    if not structure:
        return None
    parts = str(structure).split('.')
    return '.'.join(parts[:-1]) if len(parts) > 1 else None

This function enables the algorithm to build the hierarchy without requiring recursive parsing of the entire document text.

Stack-Based Nesting Algorithm

The Markdown tree builder in pageindex/page_index_md.py implements an efficient stack algorithm:


# pageindex/page_index_md.py – building the tree

while stack and stack[-1][1] >= current_level:
    stack.pop()
if not stack:
    root_nodes.append(tree_node)
else:
    parent_node, _ = stack[-1]
    parent_node['nodes'].append(tree_node)
stack.append((tree_node, current_level))

This approach ensures O(n) complexity for tree construction from n headings, avoiding expensive recursive lookups.

Node ID Assignment and Field Customization

The write_node_id function (called from md_to_tree and similar PDF functions) walks the completed tree and assigns sequential, zero-padded four-digit identifiers (e.g., "0001", "0002"), guaranteeing deterministic references for downstream applications.

Additionally, the format_structure function can be instructed to keep or drop fields (summary, text, etc.) by passing an ordered list, which is useful for compact output or for feeding only the hierarchy to another model.

Practical Code Examples

Generating Trees from PDF Documents

To process a PDF with embedded TOC data:

from pageindex.page_index import check_toc, process_toc_with_page_numbers

# Assume `pdf_pages` is a list of (text, token_len) tuples from utils.get_page_tokens()

toc_info = check_toc(pdf_pages, opt)                # Finds TOC pages

tree = process_toc_with_page_numbers(
    toc_content=toc_info['toc_content'],
    toc_page_list=toc_info['toc_page_list'],
    page_list=pdf_pages,
    toc_check_page_num=opt.toc_check_page_num,
    model=opt.model,
    logger=opt.logger,
)

# `tree` is a list of dicts with the fields described above

print(tree[0])

Source: check_toc and process_toc_with_page_numbers in pageindex/page_index.py (lines 88‑124)

Processing Markdown Files

For Markdown documents, use the async md_to_tree function:

import asyncio
from pageindex.page_index_md import md_to_tree

tree = asyncio.run(md_to_tree(
    md_path='samples/example.md',
    if_thinning=False,
    if_add_node_summary='yes',
    summary_token_threshold=200,
    model='gpt-4o',
))

# The output is a dict: {'doc_name': 'example', 'structure': [...]}

print(tree['structure'][0])

Source: md_to_tree in pageindex/page_index_md.py (lines 259‑285)

Converting Flat TOC Lists Manually

To manually convert a flat list of TOC entries (with structure fields like "1.2.3"):

from pageindex.utils import list_to_tree

# `flat_toc` is the list returned from toc_extractor

tree = list_to_tree(flat_toc)
print(tree[0])          # First top‑level section

Source: list_to_tree implementation in pageindex/utils.py (lines 50‑85)

Summary

  • PageIndex uses a recursive JSON-like structure where each node contains metadata (title, node_id, page indices) and a nested nodes array for children.
  • The PDF pipeline extracts a flat TOC list, then uses list_to_tree in pageindex/utils.py to build hierarchies based on dot-notation structure codes (e.g., "1.2.3").
  • The Markdown pipeline parses heading levels directly, using a stack-based algorithm in build_tree_from_nodes to establish parent-child relationships in O(n) time.
  • Both pipelines assign zero-padded four-digit node IDs for deterministic referencing and support optional fields like text content and page span indices.

Frequently Asked Questions

What is the exact schema of the PageIndex internal tree structure?

The PageIndex internal tree structure format consists of a list of dictionaries, where each dictionary represents a document section. Required fields include title (the heading text), node_id (a unique zero-padded four-digit identifier), and nodes (a list of child nodes). Optional fields include start_index and end_index for page spans, physical_index for raw PDF page tags, text for section content, and line_num for Markdown source references.

How does PageIndex determine parent-child relationships in PDF documents?

For PDF documents, PageIndex uses the structure field from the Table of Contents, which contains dot-notation codes like "1.2.3". The list_to_tree function in pageindex/utils.py splits these codes to identify parents—dropping the last segment yields the parent code (e.g., "1.2.3" becomes "1.2"). Nodes are then appended to their respective parent's nodes list, building the hierarchy recursively without parsing the full document text.

Why does the Markdown pipeline use a stack-based algorithm?

The Markdown pipeline uses a stack-based approach in build_tree_from_nodes (located in pageindex/page_index_md.py) to achieve O(n) linear complexity when processing n headings. The algorithm maintains a stack of (node, level) tuples, popping entries until it finds a node with a lower heading level than the current one, then appending the new node to that parent's nodes list. This avoids expensive recursive lookups and efficiently handles arbitrarily deep nesting.

Can I customize which fields are included in the final tree output?

Yes, PageIndex allows field customization through the format_structure function available in both pipelines. You can pass an ordered list of field names to include or exclude specific metadata such as summary, text, or physical_index. This is particularly useful when optimizing the output size for LLM context windows or when only the hierarchy (without content) is required for downstream processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →