How pdf-inspector Computes Font Statistics to Detect Heading Tiers (H1–H4)
pdf-inspector derives heading tiers from the distribution of font sizes in a PDF by collecting candidate sizes from bold text lines, filtering for uniqueness within 0.5 pt tolerance, keeping the four largest distinct tiers, then mapping each line's font size against these tiers using ratio-based heuristics and boldness checks.
The pdf-inspector library—a Rust-based PDF-to-Markdown converter from Firecrawl—determines document structure without relying on embedded PDF tags or visual layout analysis. Instead, it computes font statistics from the actual text content to distinguish headings from body text and assign appropriate H1–H4 levels. This article examines the two-phase algorithm implemented in src/markdown/analysis.rs that powers this detection.
Phase 1: Building the Heading Tier List with compute_heading_tiers
The first step in detecting heading tiers is constructing a document-specific list of prominent font sizes. The compute_heading_tiers function scans all extracted text lines and identifies sizes that likely represent headings.
Candidate Selection Criteria
The function applies two filters when evaluating text lines:
- Bold text only – Non-bold text is skipped entirely, as boldness is a strong heading signal
- Excludes pure digit strings – Lines containing only digits (e.g., "42" or "2024") are rejected to prevent page numbers from polluting the tier list
Deduplication with Tolerance
Collected sizes are deduplicated using a 0.5 pt tolerance to prevent near-identical sizes from creating separate tiers:
// src/markdown/analysis.rs (lines 421-429)
if !tiers.iter().any(|&t| (t - size).abs() < 0.5) {
tiers.push(size);
}
This threshold accounts for subtle font size variations that may appear across different PDF generators or font subsets.
Tier Limits and Ordering
After collection, the algorithm:
- Sorts tiers in descending order (largest first)
- Truncates to maximum four tiers (H1 through H4)
- Maps tier index 0 → H1, tier index 1 → H2, and so on
// src/markdown/analysis.rs
tiers.sort_by(|a, b| b.partial_cmp(a).unwrap());
tiers.truncate(4);
The result is a document-specific, data-driven set of font-size thresholds tailored to that PDF's actual typography.
Phase 2: Classifying Lines with detect_header_level
Once the tier list is established, detect_header_level assigns heading levels to individual text lines. The function signature reflects all inputs needed for this decision:
// src/markdown/analysis.rs (lines 442-493)
pub(crate) fn detect_header_level(
font_size: f32,
base_size: f32, // Document's body text size
heading_tiers: &[f32], // From compute_heading_tiers
is_bold: bool,
) -> Option<usize>
Ratio-Based Gate Checks
The algorithm begins by computing ratio = font_size / base_size and applying tiered logic:
Small ratio with boldness (1.05–1.20)
When the size is only marginally larger than body text but the line is bold, the function attempts to match against discovered tiers. A successful match returns the corresponding tier-based level.
Large ratio fallback (≥1.20)
If the ratio reaches 1.20 or higher and a tier match exists, the tier mapping is applied. If no tier matches but the ratio exceeds 1.5, the function extrapolates a level one step beyond the last discovered tier (capped at H4).
Pure Ratio Fallback for Unstructured Documents
When compute_heading_tiers finds no valid tiers—common in documents with uniform typography—the algorithm falls back to classic ratio thresholds:
// src/markdown/analysis.rs
if ratio >= 2.0 { Some(1) } // H1
else if ratio >= 1.5 { Some(2) } // H2
else if ratio >= 1.25 { Some(3) } // H3
else { Some(4) } // H4
This ensures robust heading detection even in poorly formatted PDFs lacking visual hierarchy cues.
Complete Working Example
The following Rust code demonstrates the full pipeline from extracted text lines to classified headings:
use pdf_inspector::markdown::analysis::{
compute_heading_tiers,
detect_header_level,
};
use pdf_inspector::types::TextLine;
// 1. Extract lines from a PDF (via process_pdf_with_options)
let lines: Vec<TextLine> = /* extraction result */;
// 2. Determine body text size (most frequent size in document)
let base_font: f32 = 12.0;
// 3. Build document-specific heading tiers
let heading_tiers = compute_heading_tiers(&lines, base_font);
// 4. Classify each line
for line in &lines {
let size = line.items[0].font_size;
let bold = line.items[0].is_bold;
if let Some(level) = detect_header_level(size, base_font, &heading_tiers, bold) {
println!("Page {}: H{} → {}", line.page, level, line.text());
}
}
Architecture and Integration
The heading detection system spans three core files:
src/markdown/analysis.rs– Containscompute_heading_tiersanddetect_header_level, implementing the font statistics logic described abovesrc/markdown/heading.rs– Consumes detection results to merge heading fragments and generate final Markdown syntaxsrc/markdown/convert.rs– Orchestrates the full conversion pipeline, invoking heading analysis during PDF-to-Markdown transformation
This separation allows heading.rs to handle structural decisions (resolving multi-line headings) while analysis.rs focuses purely on statistical font-size classification.
Summary
- pdf-inspector computes font statistics through a two-phase approach: tier discovery followed by line-level classification
compute_heading_tiersbuilds a document-specific list of up to four font sizes from bold, non-numeric text, using 0.5 pt deduplication tolerancedetect_header_levelapplies ratio gates (1.05–1.20 and ≥1.20), tier matching, and fallback thresholds to assign H1–H4 levels- The algorithm degrades gracefully to pure ratio-based detection when no distinct tiers exist in the source document
- All implementation resides in
src/markdown/analysis.rswith integration points inheading.rsandconvert.rs
Frequently Asked Questions
Why does pdf-inspector exclude pure digit strings from heading detection?
Pure digit strings like "42" or "2024" frequently appear as page numbers, footnote references, or list markers—none of which should be treated as structural headings. The filter in compute_heading_tiers ensures these numeric artifacts don't contaminate the font-size tier list.
What happens if a PDF uses only one font size throughout?
When compute_heading_tiers finds no valid tiers, detect_header_level falls back to fixed ratio thresholds: ≥2.0× base size becomes H1, ≥1.5× becomes H2, ≥1.25× becomes H3, and anything smaller becomes H4. This ensures some structural output even from uniformly-formatted documents.
How does the 0.5 pt tolerance affect detection accuracy?
The tolerance prevents font size variations—caused by different PDF generators, font subsetting, or floating-point rounding—from fragmenting what should be a single heading tier. Sizes within 0.5 points are treated as identical, producing cleaner, more accurate tier boundaries without manual threshold tuning.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →