# How pdf-inspector Computes Font Statistics to Detect Heading Tiers (H1–H4)

> Learn how pdf-inspector computes font statistics to detect H1-H4 headings. It analyzes bold text, filters font sizes, and uses ratio heuristics to identify heading tiers. Understand the PDF heading structure.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**pdf-inspector derives heading tiers from the distribution of font sizes in a PDF by collecting candidate sizes from bold text lines, filtering for uniqueness within 0.5 pt tolerance, keeping the four largest distinct tiers, then mapping each line's font size against these tiers using ratio-based heuristics and boldness checks.**

The **pdf-inspector** library—a Rust-based PDF-to-Markdown converter from Firecrawl—determines document structure without relying on embedded PDF tags or visual layout analysis. Instead, it computes **font statistics** from the actual text content to distinguish headings from body text and assign appropriate H1–H4 levels. This article examines the two-phase algorithm implemented in [`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs) that powers this detection.

## Phase 1: Building the Heading Tier List with `compute_heading_tiers`

The first step in detecting heading tiers is constructing a document-specific list of prominent font sizes. The `compute_heading_tiers` function scans all extracted text lines and identifies sizes that likely represent headings.

### Candidate Selection Criteria

The function applies two filters when evaluating text lines:

- **Bold text only** – Non-bold text is skipped entirely, as boldness is a strong heading signal
- **Excludes pure digit strings** – Lines containing only digits (e.g., "42" or "2024") are rejected to prevent page numbers from polluting the tier list

### Deduplication with Tolerance

Collected sizes are deduplicated using a **0.5 pt tolerance** to prevent near-identical sizes from creating separate tiers:

```rust
// src/markdown/analysis.rs (lines 421-429)
if !tiers.iter().any(|&t| (t - size).abs() < 0.5) {
    tiers.push(size);
}

```

This threshold accounts for subtle font size variations that may appear across different PDF generators or font subsets.

### Tier Limits and Ordering

After collection, the algorithm:
- Sorts tiers in **descending order** (largest first)
- Truncates to **maximum four tiers** (H1 through H4)
- Maps tier index 0 → H1, tier index 1 → H2, and so on

```rust
// src/markdown/analysis.rs
tiers.sort_by(|a, b| b.partial_cmp(a).unwrap());
tiers.truncate(4);

```

The result is a document-specific, data-driven set of font-size thresholds tailored to that PDF's actual typography.

## Phase 2: Classifying Lines with `detect_header_level`

Once the tier list is established, `detect_header_level` assigns heading levels to individual text lines. The function signature reflects all inputs needed for this decision:

```rust
// src/markdown/analysis.rs (lines 442-493)
pub(crate) fn detect_header_level(
    font_size: f32,
    base_size: f32,        // Document's body text size
    heading_tiers: &[f32], // From compute_heading_tiers
    is_bold: bool,
) -> Option<usize>

```

### Ratio-Based Gate Checks

The algorithm begins by computing `ratio = font_size / base_size` and applying tiered logic:

**Small ratio with boldness (1.05–1.20)**

When the size is only marginally larger than body text but the line is bold, the function attempts to match against discovered tiers. A successful match returns the corresponding tier-based level.

**Large ratio fallback (≥1.20)**

If the ratio reaches 1.20 or higher and a tier match exists, the tier mapping is applied. If no tier matches but the ratio exceeds 1.5, the function extrapolates a level one step beyond the last discovered tier (capped at H4).

### Pure Ratio Fallback for Unstructured Documents

When `compute_heading_tiers` finds no valid tiers—common in documents with uniform typography—the algorithm falls back to classic ratio thresholds:

```rust
// src/markdown/analysis.rs
if ratio >= 2.0      { Some(1) }  // H1
else if ratio >= 1.5 { Some(2) }  // H2
else if ratio >= 1.25 { Some(3) } // H3
else                 { Some(4) }  // H4

```

This ensures robust heading detection even in poorly formatted PDFs lacking visual hierarchy cues.

## Complete Working Example

The following Rust code demonstrates the full pipeline from extracted text lines to classified headings:

```rust
use pdf_inspector::markdown::analysis::{
    compute_heading_tiers,
    detect_header_level,
};
use pdf_inspector::types::TextLine;

// 1. Extract lines from a PDF (via process_pdf_with_options)
let lines: Vec<TextLine> = /* extraction result */;

// 2. Determine body text size (most frequent size in document)
let base_font: f32 = 12.0;

// 3. Build document-specific heading tiers
let heading_tiers = compute_heading_tiers(&lines, base_font);

// 4. Classify each line
for line in &lines {
    let size = line.items[0].font_size;
    let bold = line.items[0].is_bold;
    
    if let Some(level) = detect_header_level(size, base_font, &heading_tiers, bold) {
        println!("Page {}: H{} → {}", line.page, level, line.text());
    }
}

```

## Architecture and Integration

The heading detection system spans three core files:

- **[`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs)** – Contains `compute_heading_tiers` and `detect_header_level`, implementing the font statistics logic described above
- **[`src/markdown/heading.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/heading.rs)** – Consumes detection results to merge heading fragments and generate final Markdown syntax
- **[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)** – Orchestrates the full conversion pipeline, invoking heading analysis during PDF-to-Markdown transformation

This separation allows [`heading.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/heading.rs) to handle structural decisions (resolving multi-line headings) while [`analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/analysis.rs) focuses purely on statistical font-size classification.

## Summary

- **pdf-inspector** computes font statistics through a two-phase approach: tier discovery followed by line-level classification
- **`compute_heading_tiers`** builds a document-specific list of up to four font sizes from bold, non-numeric text, using 0.5 pt deduplication tolerance
- **`detect_header_level`** applies ratio gates (1.05–1.20 and ≥1.20), tier matching, and fallback thresholds to assign H1–H4 levels
- The algorithm degrades gracefully to pure ratio-based detection when no distinct tiers exist in the source document
- All implementation resides in [`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs) with integration points in [`heading.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/heading.rs) and [`convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/convert.rs)

## Frequently Asked Questions

### Why does pdf-inspector exclude pure digit strings from heading detection?

Pure digit strings like "42" or "2024" frequently appear as page numbers, footnote references, or list markers—none of which should be treated as structural headings. The filter in `compute_heading_tiers` ensures these numeric artifacts don't contaminate the font-size tier list.

### What happens if a PDF uses only one font size throughout?

When `compute_heading_tiers` finds no valid tiers, `detect_header_level` falls back to fixed ratio thresholds: ≥2.0× base size becomes H1, ≥1.5× becomes H2, ≥1.25× becomes H3, and anything smaller becomes H4. This ensures some structural output even from uniformly-formatted documents.

### How does the 0.5 pt tolerance affect detection accuracy?

The tolerance prevents font size variations—caused by different PDF generators, font subsetting, or floating-point rounding—from fragmenting what should be a single heading tier. Sizes within 0.5 points are treated as identical, producing cleaner, more accurate tier boundaries without manual threshold tuning.