PageIndex Internal Tree Structure Format: Schema, Generation, and Implementation
PageIndex represents documents as nested JSON-like trees where each node contains metadata fields including title, node_id, page indices, and a recursive nodes array, generated through either a PDF TOC parser or a Markdown heading analyzer.
PageIndex is an open-source Python library developed by VectifyAI that transforms flat documents into navigable hierarchies. Understanding the internal tree structure format is crucial for developers building retrieval-augmented generation (RAG) systems or document analysis agents. This guide examines the exact node schema, field definitions, and the dual-generation pipelines that construct these hierarchical trees from both PDF and Markdown sources.
Understanding the PageIndex Tree Structure Schema
The internal tree structure format in PageIndex follows a recursive dictionary schema. Each document becomes a list of root-level nodes, where each node represents a section or subsection with standardized metadata fields.
Core Node Fields
Every node in the tree is a Python dictionary containing the following fields:
- title: The exact section heading text extracted from the document (e.g.,
"Introduction") - node_id: A unique, zero-padded four-digit identifier (e.g.,
"0001") - start_index / end_index: Integer page numbers indicating the section's physical span (added during post-processing)
- physical_index: Raw page tag extracted from the PDF (e.g.,
12), later converted to an integer - text: Full text content of the section (optional, added on demand)
- line_num: Line number in the original Markdown file (Markdown mode only)
- nodes: Recursive list of child nodes following the same schema
Top-Level Container Structure
The final output is either a raw list of root nodes or a dictionary wrapper containing a structure key (as returned by the high-level API). This flexibility allows PageIndex to interface with both direct API consumers and LLM-friendly JSON outputs.
How PageIndex Generates the Internal Tree Structure
PageIndex implements two distinct pipelines for tree generation, depending on the source document type: PDF documents with Table of Contents (TOC) metadata, and Markdown files with heading hierarchies.
PDF Pipeline: From Flat TOC to Nested Hierarchy
The PDF pipeline processes documents containing embedded TOC information. This workflow operates in three phases:
-
Extraction: The
toc_extractorfunction inpageindex/page_index.pyparses the PDF to identify TOC pages and extracts a flat list of entries. Each entry contains astructurefield (e.g.,"1.2.3"),title, and rawphysical_indextags. -
Tree Construction: The
list_to_treefunction inpageindex/utils.pyconverts the flat list into the nested hierarchy. It uses the dot-notation in thestructurefield to infer parent-child relationships. The algorithm splits the structure code, drops the last segment to identify the parent, and appends the node to the parent'snodeslist. -
Post-Processing: The
post_processingfunction addsstart_indexandend_indexfields based on the physical page spans and optionally removes emptynodesarrays.
Markdown Pipeline: Stack-Based Heading Parsing
For Markdown documents, PageIndex builds the tree by analyzing heading levels directly. This pipeline is implemented in pageindex/page_index_md.py:
-
Node Extraction: The
extract_nodes_from_markdownfunction scans the file for#headings, recording the heading level (number of#symbols) and line number. -
Text Attachment: The
extract_node_text_contentfunction associates the subsequent text blocks with each heading node. -
Tree Building: The
build_tree_from_nodesfunction uses a stack-based algorithm to establish parent-child relationships. It maintains a stack of(node, level)tuples. When processing a new heading, it pops entries until finding a node with a lower heading level, then appends the new node to that parent'snodeslist. -
Formatting: The
format_structurefunction reorders fields and prunes empty children arrays before final output.
Key Implementation Details
Understanding the specific implementation mechanics helps developers customize PageIndex for specialized document formats.
Parent Detection via Dot-Notation Codes
The core logic for determining parent-child relationships in PDF processing resides in pageindex/utils.py. The get_parent_structure helper parses dot-notation structure codes:
# pageindex/utils.py – core of list_to_tree
def get_parent_structure(structure):
if not structure:
return None
parts = str(structure).split('.')
return '.'.join(parts[:-1]) if len(parts) > 1 else None
This function enables the algorithm to build the hierarchy without requiring recursive parsing of the entire document text.
Stack-Based Nesting Algorithm
The Markdown tree builder in pageindex/page_index_md.py implements an efficient stack algorithm:
# pageindex/page_index_md.py – building the tree
while stack and stack[-1][1] >= current_level:
stack.pop()
if not stack:
root_nodes.append(tree_node)
else:
parent_node, _ = stack[-1]
parent_node['nodes'].append(tree_node)
stack.append((tree_node, current_level))
This approach ensures O(n) complexity for tree construction from n headings, avoiding expensive recursive lookups.
Node ID Assignment and Field Customization
The write_node_id function (called from md_to_tree and similar PDF functions) walks the completed tree and assigns sequential, zero-padded four-digit identifiers (e.g., "0001", "0002"), guaranteeing deterministic references for downstream applications.
Additionally, the format_structure function can be instructed to keep or drop fields (summary, text, etc.) by passing an ordered list, which is useful for compact output or for feeding only the hierarchy to another model.
Practical Code Examples
Generating Trees from PDF Documents
To process a PDF with embedded TOC data:
from pageindex.page_index import check_toc, process_toc_with_page_numbers
# Assume `pdf_pages` is a list of (text, token_len) tuples from utils.get_page_tokens()
toc_info = check_toc(pdf_pages, opt) # Finds TOC pages
tree = process_toc_with_page_numbers(
toc_content=toc_info['toc_content'],
toc_page_list=toc_info['toc_page_list'],
page_list=pdf_pages,
toc_check_page_num=opt.toc_check_page_num,
model=opt.model,
logger=opt.logger,
)
# `tree` is a list of dicts with the fields described above
print(tree[0])
Source: check_toc and process_toc_with_page_numbers in pageindex/page_index.py (lines 88‑124)
Processing Markdown Files
For Markdown documents, use the async md_to_tree function:
import asyncio
from pageindex.page_index_md import md_to_tree
tree = asyncio.run(md_to_tree(
md_path='samples/example.md',
if_thinning=False,
if_add_node_summary='yes',
summary_token_threshold=200,
model='gpt-4o',
))
# The output is a dict: {'doc_name': 'example', 'structure': [...]}
print(tree['structure'][0])
Source: md_to_tree in pageindex/page_index_md.py (lines 259‑285)
Converting Flat TOC Lists Manually
To manually convert a flat list of TOC entries (with structure fields like "1.2.3"):
from pageindex.utils import list_to_tree
# `flat_toc` is the list returned from toc_extractor
tree = list_to_tree(flat_toc)
print(tree[0]) # First top‑level section
Source: list_to_tree implementation in pageindex/utils.py (lines 50‑85)
Summary
- PageIndex uses a recursive JSON-like structure where each node contains metadata (title, node_id, page indices) and a nested
nodesarray for children. - The PDF pipeline extracts a flat TOC list, then uses
list_to_treeinpageindex/utils.pyto build hierarchies based on dot-notation structure codes (e.g., "1.2.3"). - The Markdown pipeline parses heading levels directly, using a stack-based algorithm in
build_tree_from_nodesto establish parent-child relationships in O(n) time. - Both pipelines assign zero-padded four-digit node IDs for deterministic referencing and support optional fields like
textcontent and page span indices.
Frequently Asked Questions
What is the exact schema of the PageIndex internal tree structure?
The PageIndex internal tree structure format consists of a list of dictionaries, where each dictionary represents a document section. Required fields include title (the heading text), node_id (a unique zero-padded four-digit identifier), and nodes (a list of child nodes). Optional fields include start_index and end_index for page spans, physical_index for raw PDF page tags, text for section content, and line_num for Markdown source references.
How does PageIndex determine parent-child relationships in PDF documents?
For PDF documents, PageIndex uses the structure field from the Table of Contents, which contains dot-notation codes like "1.2.3". The list_to_tree function in pageindex/utils.py splits these codes to identify parents—dropping the last segment yields the parent code (e.g., "1.2.3" becomes "1.2"). Nodes are then appended to their respective parent's nodes list, building the hierarchy recursively without parsing the full document text.
Why does the Markdown pipeline use a stack-based algorithm?
The Markdown pipeline uses a stack-based approach in build_tree_from_nodes (located in pageindex/page_index_md.py) to achieve O(n) linear complexity when processing n headings. The algorithm maintains a stack of (node, level) tuples, popping entries until it finds a node with a lower heading level than the current one, then appending the new node to that parent's nodes list. This avoids expensive recursive lookups and efficiently handles arbitrarily deep nesting.
Can I customize which fields are included in the final tree output?
Yes, PageIndex allows field customization through the format_structure function available in both pipelines. You can pass an ordered list of field names to include or exclude specific metadata such as summary, text, or physical_index. This is particularly useful when optimizing the output size for LLM context windows or when only the hierarchy (without content) is required for downstream processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →