# How LiteParse Reconstructs Headings, Tables, and Lists in Its Markdown Output

> Discover how LiteParse reconstructs headings, tables, and lists in its Markdown output. Learn about its three-phase pipeline: heading calibration, block classification, and rendering for clean documents.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-24

---

**LiteParse converts PDF projected lines into clean Markdown through a three-phase pipeline: document-wide heading calibration, per-page block classification, and final block rendering with fallback handling for unstructured content.**

LiteParse, an open-source PDF parser from the run-llama/liteparse repository, transforms raw `ProjectedLine` data into structured Markdown documents. The conversion relies on a sophisticated architecture that first calibrates heading thresholds across the entire document, then classifies individual pages into semantic blocks, and finally renders those blocks into standard Markdown syntax. This reconstruction process preserves the original document hierarchy while handling edge cases like merged table cells and multi-line headings.

## Phase 1: Document-Wide Heading Calibration

Before processing individual pages, LiteParse establishes a **heading-size map** by analyzing font statistics across the entire document. In [`crates/liteparse/src/markdown_layout/headings.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/markdown_layout/headings.rs), the `compute_body_size` function (≈ line 6780) builds a character-weighted histogram of font sizes, deliberately excluding rotated text, figure captions, and lines from detectable chrome.

Using this baseline, `build_heading_map` (≈ line 6440) identifies sizes that are **significantly larger** than the body size using thresholds like `HEADING_SIZE_EPSILON` or `ESTIMATED_HEADING_SIZE_MARGIN`. The function applies quality guards—total character counts, average characters per line, and alphabetic-ratio filters—to eliminate false positives before mapping the remaining sizes to heading levels 1‥6. This map drives *size-based* heading promotion during the per-page classification phase.

## Phase 2: Per-Page Block Classification

The heavy lifting occurs in `classify_page_with_filters` (≈ line 70 of [`src/markdown_layout/classify.rs`](https://github.com/run-llama/liteparse/blob/main/src/markdown_layout/classify.rs)). For each page, the system:

1. Filters repeating header and footer lines using `compute_header_footer_set` and `detect_single_page_chrome`.
2. Detects global ruled tables via `detect_ruled_tables_global` and removes their lines from the normal flow.
3. Splits remaining lines into region leaves based on `region_path`.

Inside each region, `classify_region` (≈ line 675) runs a state machine that categorizes lines into blocks: headings, paragraphs, list items, code blocks, or tables.

### Heading Detection

LiteParse employs four complementary strategies to identify headings, implemented in [`src/markdown_layout/classify.rs`](https://github.com/run-llama/liteparse/blob/main/src/markdown_layout/classify.rs):

- **Explicit PDF structure**: `struct_heading_level` (≈ line 7877) reads the PDF’s structure tree (`H1`…`H6` tags) and overrides other heuristics when present.
- **Document outline**: `outline_heading_level` (≈ line 8098) matches outline entries (bookmarks) to a line’s y‑coordinate and text prefix.
- **Size-based promotion**: `heading_level_for` (≈ line 690) looks up the line’s font size in the heading map built during Phase 1.
- **Bold-body heuristic**: `looks_like_bold_heading` (≈ line 5670) recognizes section headings that rely solely on bold styling at body size.

When a heading is identified, the system creates a `Block::Heading { level, text }`. Multi-line headings are merged via `append_inline_continuation` (≈ line 1010) to produce a single coherent block.

### Table Recognition

Table detection in [`src/markdown_layout/tables.rs`](https://github.com/run-llama/liteparse/blob/main/src/markdown_layout/tables.rs) proceeds through three stages:

1. **Cell splitting**: `split_cells` (≈ line 80) groups spans into cells using a gap threshold based on the dominant font size.
2. **Column track inference**: `infer_tracks_from_raw_items` (≈ line 4450) clusters raw span x‑positions across several rows to obtain robust column anchors, preventing tightly-kerned numeric columns from merging incorrectly.
3. **Row building**: `try_detect_table` (≈ line 6710) walks subsequent lines, assigning cells to inferred tracks using `cells_from_raw_items_with_tracks`. It tolerates merged cells via `recover_merged_cell`, sparse rows, and occasional extra columns caused by multi-column layouts.

If the run meets minimum thresholds (`TABLE_MIN_ROWS`, `TABLE_MIN_COLUMNS`) and passes the row‑spacing coefficient-of-variation test, `finalize_table_run` creates a `Block::Table { header, rows }`. When a table is too irregular, LiteParse falls back to a `Block::GridFallback` that preserves the raw text in a fenced block.

### List Item Parsing

Lists are identified in [`src/markdown_layout/lists.rs`](https://github.com/run-llama/liteparse/blob/main/src/markdown_layout/lists.rs) via `parse_list_marker` (≈ line 17), which recognizes:

- Unicode bullet glyphs (`•`, `·`, `◦`).
- Decimal markers (`1.`, `12)`) limited to three digits to avoid matching page numbers.

Upon detection, `classify_region` flushes any active paragraph, computes the nesting level from the indent using `LIST_INDENT_STEP_PT` (≈ 12 pt), and renders the item text via `render_list_item_text`. Continuation lines that satisfy `continues_paragraph` are appended to the previous list item through `append_inline_continuation`.

## Phase 3: Markdown Rendering

The final conversion happens in `format_markdown` (≈ line 20 of [`src/output/markdown.rs`](https://github.com/run-llama/liteparse/blob/main/src/output/markdown.rs)). This function orchestrates the assembly of the Markdown document:

```rust
let mut out = String::new();
for (i, page) in pages.iter().enumerate() {
    if i > 0 { out.push_str("\n\n-----\n\n"); }          // page separator
    if page.projected_lines.is_empty() {
        out.push_str("```text\n");
        out.push_str(&page.text);
        out.push_str("\n```\n");
        continue;
    }
    let blocks = classify_page_with_filters(...);
    dedupe_rules(&mut blocks);                         // collapse repeated HRs
    out.push_str(&render_blocks(&blocks));             // convert blocks → Markdown
}

```

The `render_blocks` function in [`src/markdown_layout/blocks.rs`](https://github.com/run-llama/liteparse/blob/main/src/markdown_layout/blocks.rs) converts each `Block` into its Markdown representation:

- **Headings** become ATX-style headers (`#`…`######`) according to the detected level.
- **Tables** become pipe-delimited rows with aligned headers.
- **List items** become `-` or `1.` prefixes with appropriate indentation.
- **Code blocks** receive fenced fences with an optional language hint from `detect_code_language`.

If a page contains **no projected lines** (e.g., a scanned page without OCR), the raw text is emitted inside a fenced `text` block to prevent silent data loss. The `dedupe_rules` function (≈ line 86) strips redundant horizontal rules that would appear adjacent to page separators.

## Practical Implementation Example

To extract Markdown from a PDF using LiteParse’s reconstruction pipeline:

```rust
use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;

let cfg = LiteParseConfig::default();
let parser = LiteParse::new(cfg);
let parsed = parser.parse_path("paper.pdf")?;          // → Vec<ParsedPage>
let outline = parser.outline();                       // optional PDF outline
let markdown = liteparse::output::markdown::format_markdown(
    &parsed,
    &outline,
    liteparse::config::ImageMode::Placeholder,
);
println!("{}", markdown);

```

Running this on a research article yields Markdown with reconstructed headings (`#` / `##`), pipe tables recovered from tightly-spaced numeric data, and nested lists (`-`, `1.`) generated from detected markers.

## Summary

- **Three-phase architecture**: LiteParse calibrates headings document-wide, classifies pages into semantic blocks, and renders clean Markdown.
- **Multi-strategy heading detection**: Combines PDF structure trees, document outlines, font-size mapping, and bold-body heuristics.
- **Robust table reconstruction**: Uses cell splitting, column track inference, and merged-cell recovery to handle complex tabular layouts.
- **List hierarchy preservation**: Parses Unicode and decimal markers, calculating nesting levels via fixed 12 pt indentation steps.
- **Graceful degradation**: Falls back to fenced text blocks when pages lack projectable lines or when table structures are too irregular.

## Frequently Asked Questions

### How does LiteParse handle scanned PDFs without OCR?

When a page contains no `projected_lines` (indicating a scanned image without text extraction), `format_markdown` emits the raw page text inside a fenced ` ```text ` block. This ensures no content is silently dropped while clearly indicating that the text was not structurally parsed.

### What triggers LiteParse to use a table fallback instead of a Markdown table?

If a potential table run fails to meet `TABLE_MIN_ROWS` and `TABLE_MIN_COLUMNS` thresholds, or if the row-spacing coefficient-of-variation exceeds the tolerance limit, the system creates a `Block::GridFallback` instead of a `Block::Table`. This fallback preserves the raw cell text in a preformatted block rather than risking data loss through incorrect column alignment.

### How does LiteParse distinguish between different heading levels?

LiteParse maps font sizes to heading levels 1‥6 using the `heading_map` built during document initialization. The `heading_level_for` function looks up a line’s size in this map. Additionally, explicit PDF structure tags (`H1`…`H6`) and document outline entries (bookmarks) can override size-based detection to ensure semantic accuracy.

### Can LiteParse handle nested lists with multiple indentation levels?

Yes. The `classify_region` state machine calculates list nesting depth by measuring indentation against `LIST_INDENT_STEP_PT` (approximately 12 points). Each detected level increments the indentation in the final Markdown output, preserving the hierarchical structure of bullet and ordered lists.