How PageIndex Handles Markdown Files Differently from PDF Documents
PageIndex processes Markdown files using simple regex-based header extraction on line structure, while PDF documents require LLM-driven Table-of-Contents detection, page-level tokenization, and physical index bookkeeping to reconstruct a navigable hierarchy.
PageIndex is an open-source document indexing library developed by VectifyAI that converts unstructured documents into hierarchical tree structures optimized for retrieval-augmented generation (RAG) pipelines. While both Markdown and PDF inputs ultimately produce JSON-compatible tree outputs, the underlying processing architectures differ significantly to accommodate the structural nature of each format.
Entry Points and Pipeline Architecture
The two formats enter entirely separate processing pipelines based on file extension detection.
For Markdown files, the entry point is pageindex/page_index_md.py through the md_to_tree() function. This module handles header parsing, node extraction, and tree building using only local text processing.
For PDF documents, processing begins in pageindex/page_index.py via the page_index() function, which delegates to page_index_main() and subsequently tree_parser(). This pipeline coordinates PDF-specific utilities from pageindex/utils.py for tokenization and page extraction.
Structure Discovery: Regex vs. LLM-Driven TOC
The fundamental difference lies in how each format discovers its hierarchical structure.
Markdown processing employs a simple header regex pattern (^(#{1,6})\s+(.+)$) to extract a flat list of headings and their corresponding line numbers. Since Markdown is inherently structured by header levels, no page concepts are involved, and no external AI inference is required for structure detection.
PDF processing requires significantly more sophistication. The system must first locate or generate a Table-of-Contents using LLM prompts. Physical page indices (marked as <physical_index_X> in the extracted text) are identified and validated. The add_page_number_to_toc() and convert_physical_index_to_int() functions in page_index.py manage this page-level metadata extraction.
Node Creation and Metadata Extraction
The resulting tree nodes contain different metadata fields reflecting their source formats.
For Markdown, the extract_nodes_from_markdown() function creates nodes containing title, level, line_num, and the raw markdown slice as text. The build_tree_from_nodes() function then assembles these into a hierarchical structure based on header levels.
For PDF, nodes are created via list_to_tree() and contain title, physical_index, start_index, end_index, and optionally text. The get_page_tokens() utility in utils.py handles the tokenization of PDF pages, and the tree structure is built by matching TOC items to physical page ranges.
Thinning and Summarization Strategies
Both formats support optional tree thinning and node summarization, but the operations apply to different granularities.
Markdown thinning operates on token counts of markdown sections using update_node_list_with_text_token_count() and tree_thinning_for_index(). Summaries are generated directly on the markdown text content via generate_node_summary().
PDF thinning operates on token counts of pages rather than sections, using the same update_node_list_with_text_token_count() function but applied to PDF page tokens. Summaries are generated after the PDF tree is fully constructed using generate_summaries_for_structure().
Physical Index Handling and Preface Detection
A critical distinction exists in how each format handles physical location mapping.
Markdown has no concept of physical indices—the index is purely hierarchical based on heading levels, starting from the first header encountered.
PDF requires rigorous physical index bookkeeping. The validate_and_truncate_physical_indices() function ensures page numbers map correctly to the document. Additionally, add_preface_if_needed() inserts a "Preface" node when the first physical index is greater than 1, accounting for introductory content that precedes the main TOC structure.
Practical Code Examples
The following examples demonstrate the different API entry points for each format.
Processing a Markdown file:
from pageindex import md_to_tree
# md_path points to a *.md or *.markdown file
tree = md_to_tree(
md_path="tests/markdowns/example.md",
if_thinning=False,
if_add_node_summary="yes",
summary_token_threshold=200,
model="gpt-4o-2024-11-20"
)
print(tree)
Processing a PDF file:
from pageindex import page_index
structure = page_index(
doc="tests/pdfs/example.pdf",
model="gpt-4o-2024-11-20",
if_add_node_summary="yes",
if_add_node_text="yes"
)
print(structure)
Both functions return nested JSON-compatible dictionaries, but the PDF version includes physical_index, start_index, and end_index fields that map each node back to the original PDF pages.
Summary
- Markdown processing in PageIndex uses regex-based header extraction (
pageindex/page_index_md.py) to build hierarchical trees from heading levels without page concepts. - PDF processing requires LLM-driven TOC detection, page tokenization via
get_page_tokens(), and physical index bookkeeping throughpageindex/page_index.pyto map content to specific pages. - Node metadata differs: Markdown nodes contain
line_numand raw text slices; PDF nodes containphysical_index,start_index, andend_indexfor page-level navigation. - Preface handling is unique to PDFs, inserting introductory nodes when the first content page exceeds index 1.
Frequently Asked Questions
Does PageIndex require LLM calls for Markdown processing?
No. Markdown processing in pageindex/page_index_md.py uses only local regex patterns to extract headers and build the tree structure. LLM calls are optional and only used if you enable node summarization via the if_add_node_summary parameter.
Why do PDF nodes have physical indices while Markdown nodes do not?
PDF documents have inherent pagination that users reference when navigating content. The physical_index field in PDF nodes (managed by add_page_number_to_toc() and validate_and_truncate_physical_indices()) maps tree nodes back to specific page numbers in the original document. Markdown files lack fixed pagination, so they use line_num for location reference instead.
Can PageIndex handle PDFs that lack a Table of Contents?
Yes. If a PDF does not contain an extractable TOC, PageIndex uses LLM prompts within page_index.py to generate a hierarchical structure based on the document's content and detected headings. The system still extracts physical page indices to maintain navigability even when working with generated TOCs.
What happens when a PDF starts with content before the first chapter?
PageIndex detects this scenario in add_preface_if_needed() within pageindex/page_index.py. When the first detected physical index is greater than 1, the system automatically inserts a "Preface" node covering pages 1 through (first_index - 1), ensuring no introductory content is lost in the tree structure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →