# How PageIndex Handles Markdown Files Differently from PDF Documents

> Discover how PageIndex processes Markdown via regex and PDFs using LLMs. Learn about header extraction, ToC detection, and index bookkeeping for navigable hierarchies.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: deep-dive
- Published: 2026-02-16

---

**PageIndex processes Markdown files using simple regex-based header extraction on line structure, while PDF documents require LLM-driven Table-of-Contents detection, page-level tokenization, and physical index bookkeeping to reconstruct a navigable hierarchy.**

PageIndex is an open-source document indexing library developed by VectifyAI that converts unstructured documents into hierarchical tree structures optimized for retrieval-augmented generation (RAG) pipelines. While both Markdown and PDF inputs ultimately produce JSON-compatible tree outputs, the underlying processing architectures differ significantly to accommodate the structural nature of each format.

## Entry Points and Pipeline Architecture

The two formats enter entirely separate processing pipelines based on file extension detection.

For **Markdown files**, the entry point is [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) through the `md_to_tree()` function. This module handles header parsing, node extraction, and tree building using only local text processing.

For **PDF documents**, processing begins in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) via the `page_index()` function, which delegates to `page_index_main()` and subsequently `tree_parser()`. This pipeline coordinates PDF-specific utilities from [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) for tokenization and page extraction.

## Structure Discovery: Regex vs. LLM-Driven TOC

The fundamental difference lies in how each format discovers its hierarchical structure.

**Markdown processing** employs a simple header regex pattern (`^(#{1,6})\s+(.+)$`) to extract a flat list of headings and their corresponding line numbers. Since Markdown is inherently structured by header levels, no page concepts are involved, and no external AI inference is required for structure detection.

**PDF processing** requires significantly more sophistication. The system must first locate or generate a Table-of-Contents using LLM prompts. Physical page indices (marked as `<physical_index_X>` in the extracted text) are identified and validated. The `add_page_number_to_toc()` and `convert_physical_index_to_int()` functions in [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py) manage this page-level metadata extraction.

## Node Creation and Metadata Extraction

The resulting tree nodes contain different metadata fields reflecting their source formats.

For **Markdown**, the `extract_nodes_from_markdown()` function creates nodes containing `title`, `level`, `line_num`, and the raw markdown slice as `text`. The `build_tree_from_nodes()` function then assembles these into a hierarchical structure based on header levels.

For **PDF**, nodes are created via `list_to_tree()` and contain `title`, `physical_index`, `start_index`, `end_index`, and optionally `text`. The `get_page_tokens()` utility in [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py) handles the tokenization of PDF pages, and the tree structure is built by matching TOC items to physical page ranges.

## Thinning and Summarization Strategies

Both formats support optional tree thinning and node summarization, but the operations apply to different granularities.

**Markdown thinning** operates on token counts of markdown sections using `update_node_list_with_text_token_count()` and `tree_thinning_for_index()`. Summaries are generated directly on the markdown text content via `generate_node_summary()`.

**PDF thinning** operates on token counts of pages rather than sections, using the same `update_node_list_with_text_token_count()` function but applied to PDF page tokens. Summaries are generated after the PDF tree is fully constructed using `generate_summaries_for_structure()`.

## Physical Index Handling and Preface Detection

A critical distinction exists in how each format handles physical location mapping.

**Markdown** has no concept of physical indices—the index is purely hierarchical based on heading levels, starting from the first header encountered.

**PDF** requires rigorous physical index bookkeeping. The `validate_and_truncate_physical_indices()` function ensures page numbers map correctly to the document. Additionally, `add_preface_if_needed()` inserts a "Preface" node when the first physical index is greater than 1, accounting for introductory content that precedes the main TOC structure.

## Practical Code Examples

The following examples demonstrate the different API entry points for each format.

Processing a Markdown file:

```python
from pageindex import md_to_tree

# md_path points to a *.md or *.markdown file

tree = md_to_tree(
    md_path="tests/markdowns/example.md",
    if_thinning=False,
    if_add_node_summary="yes",
    summary_token_threshold=200,
    model="gpt-4o-2024-11-20"
)
print(tree)

```

Processing a PDF file:

```python
from pageindex import page_index

structure = page_index(
    doc="tests/pdfs/example.pdf",
    model="gpt-4o-2024-11-20",
    if_add_node_summary="yes",
    if_add_node_text="yes"
)
print(structure)

```

Both functions return nested JSON-compatible dictionaries, but the PDF version includes `physical_index`, `start_index`, and `end_index` fields that map each node back to the original PDF pages.

## Summary

- **Markdown processing** in PageIndex uses regex-based header extraction ([`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py)) to build hierarchical trees from heading levels without page concepts.
- **PDF processing** requires LLM-driven TOC detection, page tokenization via `get_page_tokens()`, and physical index bookkeeping through [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) to map content to specific pages.
- **Node metadata differs**: Markdown nodes contain `line_num` and raw text slices; PDF nodes contain `physical_index`, `start_index`, and `end_index` for page-level navigation.
- **Preface handling** is unique to PDFs, inserting introductory nodes when the first content page exceeds index 1.

## Frequently Asked Questions

### Does PageIndex require LLM calls for Markdown processing?

No. Markdown processing in [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) uses only local regex patterns to extract headers and build the tree structure. LLM calls are optional and only used if you enable node summarization via the `if_add_node_summary` parameter.

### Why do PDF nodes have physical indices while Markdown nodes do not?

PDF documents have inherent pagination that users reference when navigating content. The `physical_index` field in PDF nodes (managed by `add_page_number_to_toc()` and `validate_and_truncate_physical_indices()`) maps tree nodes back to specific page numbers in the original document. Markdown files lack fixed pagination, so they use `line_num` for location reference instead.

### Can PageIndex handle PDFs that lack a Table of Contents?

Yes. If a PDF does not contain an extractable TOC, PageIndex uses LLM prompts within [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py) to generate a hierarchical structure based on the document's content and detected headings. The system still extracts physical page indices to maintain navigability even when working with generated TOCs.

### What happens when a PDF starts with content before the first chapter?

PageIndex detects this scenario in `add_preface_if_needed()` within [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py). When the first detected physical index is greater than 1, the system automatically inserts a "Preface" node covering pages 1 through (first_index - 1), ensuring no introductory content is lost in the tree structure.