How Tagged PDF Support in pdf‑inspector Uses Structure Tree Roles vs Font‑Size Heuristics
pdf‑inspector prioritizes PDF Structure Tree roles (H1–H6, P, L, Code) for semantic extraction and only falls back to font‑size heuristics when tags are missing or the PDF is untagged.
Tagged PDFs contain machine‑readable semantic markup that preserves the author's intended document structure. The firecrawl/pdf-inspector library leverages this markup when available, ensuring headings remain headings and lists remain lists in the extracted Markdown. When a PDF lacks this metadata, the library gracefully degrades to analyzing visual properties like font size and weight to reconstruct the structure.
How Structure Tree Role Extraction Works
The primary extraction path in pdf‑inspector relies on the PDF Structure Tree, a standardized hierarchy of elements with assigned semantic roles.
Parsing the Structure Tree
In src/structure_tree.rs, the library parses the PDF's structure tree and builds a list of elements with their associated PDF roles. Each element carries a role name that directly maps to document semantics:
- H1–H6 → Markdown headings (
#through######) - P → Plain paragraphs
- L → List items
- Code → Fenced code blocks
This role‑driven approach preserves the original author's intent verbatim, without guessing based on visual appearance.
// Example: Extract a PDF while preserving tags when available
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::ProcessOptions;
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Options enable both tag‑aware and heuristic extraction
let opts = ProcessOptions {
preserve_structure: true, // try to use the structure tree first
..Default::default()
};
let md = process_pdf_with_options("report.pdf", opts)?;
println!("{}", md);
Ok(())
}
The mapping logic resides in src/markdown/classify.rs, which determines whether to use the extracted role or defer to heuristic classification for untagged elements.
When and How Font‑Size Heuristics Apply
When a PDF lacks a Structure Tree—or when specific elements within a tagged PDF carry no role assignment—pdf‑inspector activates its heuristic fallback system.
Heuristic Detection Logic
Two files handle the visual analysis:
src/markdown/heading.rs– Identifies headings by evaluating font size, weight, and spacing patternssrc/markdown/classify.rs– Handles broader classification for lists, code blocks, and paragraphs
The heuristics follow predictable patterns:
| Visual Cue | Inferred Structure |
|---|---|
| Larger font size + bold weight | Heading level (relative size determines H1 vs H2, etc.) |
| Indentation or bullet characters | List item |
| Monospaced font family | Code block |
| Default styling | Paragraph |
This fallback ensures robust extraction even from legacy PDFs created before tagging conventions existed or from scanned documents converted without semantic enrichment.
The Complete Decision Flow
The orchestration happens in src/lib.rs through the process_pdf_with_options function, which implements a clear priority order:
- Detect Structure Tree – Check for
/StructTreeRootin the PDF catalog - Extract role‑tagged elements – Map each role to its Markdown equivalent
- Classify untagged elements – Apply font‑size heuristics to remaining content
The command‑line interface in src/bin/pdf2md.rs exposes this same logic with optional overrides:
# Extract with structure‑tree support (default)
pdf2md report.pdf > report.md
# Force heuristic‑only extraction (skip tags)
pdf2md --no-structure report.pdf > report.md
Performance and Accuracy Trade‑offs
| Aspect | Structure Tree Approach | Font‑Size Heuristics |
|---|---|---|
| Speed | Faster (direct role lookup) | Slower (analysis required) |
| Accuracy | Perfect semantic fidelity | Approximate, prone to misclassification |
| Availability | Requires properly tagged PDF | Works with any PDF |
| Heading levels | Explicitly defined | Inferred from relative sizing |
The dual‑approach design means pdf‑inspector delivers optimal results for modern, accessibility‑compliant PDFs without sacrificing utility for older documents.
Summary
- Structure tree roles (H1–H6, P, L, Code) in
src/structure_tree.rsprovide the primary extraction path for tagged PDFs - Font‑size heuristics in
src/markdown/heading.rsandsrc/markdown/classify.rsserve as a fallback for untagged content - The
ProcessOptions.preserve_structureflag and--no-structureCLI option control which mode takes precedence - Role‑based extraction preserves author intent exactly; heuristics approximate structure from visual cues
Frequently Asked Questions
What PDF versions support Structure Trees for tagged extraction?
PDF 1.3 and later support the Structure Tree specification, though practical adoption varies. Academic publishers and government documents typically include full tagging, while older corporate PDFs often rely solely on visual formatting. pdf‑inspector handles both transparently.
Can I force heuristic extraction even when tags exist?
Yes. Pass --no-structure to the pdf2md CLI binary, or set preserve_structure: false in ProcessOptions when using the Rust API. This disables role‑based extraction entirely and routes all content through the font‑size analysis pipeline.
How does pdf‑inspector handle mixed tagging where some elements have roles and others don't?
The classifier in src/markdown/classify.rs processes each content block individually. Tagged elements use their explicit roles; untagged siblings trigger the heuristic engine. This hybrid approach ensures maximum coverage without requiring perfect tagging throughout a document.
Does font‑size heuristic detection require knowing the document's base font size?
The heuristics operate on relative comparisons rather than absolute thresholds. The algorithm identifies outlier sizes within the document itself—larger fonts become higher heading levels, smaller or monospaced fonts trigger code detection. This self‑referential approach avoids configuration per document.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →