How PDF-Inspector Handles Tagged PDFs Using the Structure Tree
PDF-Inspector detects tagged PDFs by locating the /StructTreeRoot entry in the document catalog and uses the structure tree to map MCIDs to semantic roles, enabling accurate Markdown generation without relying solely on visual heuristics.
PDF-Inspector is an open-source Rust library for converting PDFs to structured Markdown. When processing tagged PDFs, it leverages the embedded structure tree to extract semantic meaning directly from the document's markup rather than inferring layout from coordinates. This article explores how the library parses the /StructTreeRoot, maps structure elements to content streams, and integrates semantic roles into the extraction pipeline according to the firecrawl/pdf-inspector source code.
Detecting Tagged PDFs and Loading the Structure Tree
PDF-Inspector determines whether a PDF contains semantic markup by checking for the presence of a /StructTreeRoot entry in the document catalog. In src/structure_tree.rs, the StructTree::from_doc method performs this detection:
// src/structure_tree.rs – lines 44-49
pub fn from_doc(doc: &Document) -> Option<Self> {
let catalog = doc.catalog().ok()?;
let struct_root_obj = catalog.get(b"StructTreeRoot").ok()?;
…
}
If the /StructTreeRoot is missing, the method returns None, signaling that the PDF is untagged. In this case, the extractor falls back to visual heuristics such as column detection and line-based table detection. When the root is present, the parser proceeds to build the complete structure tree.
Parsing Standard Structure Roles
Once a tagged PDF is detected, PDF-Inspector interprets the semantic intent of each element using the StructRole enum defined in src/structure_tree.rs (lines 16-75). This enum maps every ISO-defined standard role—such as H1, P, Table, TD, and TH—and includes a catch-all Other(String) variant for custom tags.
The library resolves custom tag names to standard roles through a role map, ensuring that even documents with specialized tagging schemes produce consistent output. You can query roles via StructRole::from_name, which normalizes the incoming tag strings against the ISO vocabulary.
Building the Structure Tree
The tree construction process recursively traverses the /K (kids) entries starting from the StructTreeRoot. Each node becomes a StructElement that stores:
- Its semantic role
- Optional alt-text and actual-text properties
- Language attributes
- MCID (Marked Content ID) references
- Child element relationships
The parser in src/structure_tree.rs (lines 82-197) includes defensive measures against malformed PDFs, enforcing depth limits and properly handling MCR (Marked Content Reference) and OBJR (Object Reference) entries. This recursive descent builds an in-memory representation of the document's logical reading order, independent of the physical page layout.
Mapping MCIDs to Semantic Roles
After constructing the tree, PDF-Inspector creates a lookup mechanism to connect structure elements with their corresponding content in the page streams. The StructTree::mcid_to_roles method (lines 66-81) walks the tree and generates a per-page HashMap mapping each MCID to its StructRole:
use std::collections::BTreeMap;
use pdf_inspector::structure_tree::StructTree;
// Assume `doc` is a lopdf::Document already loaded
let page_ids: BTreeMap<u32, ObjectId> = doc.get_pages(); // page number → ObjectId
let struct_tree = StructTree::from_doc(&doc).unwrap();
let mcid_roles = struct_tree.mcid_to_roles(&page_ids);
// `mcid_roles[page][mcid]` now yields the `StructRole` (e.g. H1, P, TD)
This map is consulted by the content-stream extractor so that each extracted TextItem can be annotated with its semantic role, preventing visual heuristics from misclassifying short captions or table headers as headings.
Flattening for Linear Processing
To facilitate linear output formats like Markdown, the library provides StructTree::flatten (lines 115-124), which produces a vector of FlatStructElement objects preserving the document's reading order. This flattened view eliminates the need for a separate layout analysis pass when generating Markdown.
let flat = struct_tree.flatten();
for elem in flat {
println!("Depth {} – role: {:?}", elem.depth, elem.role);
}
The depth information maintains the hierarchical relationship between elements, allowing the Markdown generator to correctly nest headings and list items.
Table Extraction from Semantic Tags
When a tagged PDF contains explicit table markup (/Table, /TR, /TH, /TD), PDF-Inspector bypasses geometry-based detection in favor of semantic extraction. The StructTree::extract_tables method (lines 123-139) walks the structure tree, collects MCIDs for each cell, and returns a StructTable structure:
let tables = struct_tree.extract_tables(&page_ids);
for (i, table) in tables.iter().enumerate() {
println!("Table {} has {} rows", i + 1, table.rows.len());
}
This approach yields more reliable table reconstructions on documents that already describe them semantically, avoiding the errors common in purely visual table detectors.
Integration with Markdown Generation
The Markdown conversion layer in src/markdown/convert.rs (lines 6-12 and 1529-1534) consumes both the MCID-role map and the flattened structure to drive output decisions. By consulting the StructRole information, the converter distinguishes between body text, headings, list items, and code blocks without guessing based on font size or positioning.
The src/markdown/preprocess.rs module additionally uses StructRole to influence pre-processing steps such as drop-cap merging and heading line merging, ensuring that the final Markdown accurately reflects the document's intended semantic structure.
Summary
- Detection: PDF-Inspector checks for
/StructTreeRootin the document catalog viaStructTree::from_docto identify tagged PDFs. - Role Mapping: The
StructRoleenum standardizes ISO-defined tags and custom roles, mapped throughmcid_to_rolesto connect structure elements with content streams. - Tree Parsing: Recursive parsing of
/Kentries builds a completeStructElementhierarchy with MCID references, alt-text, and language attributes. - Linear Output: The
flattenmethod producesFlatStructElementvectors that preserve reading order for Markdown generation. - Semantic Tables:
extract_tablesuses explicit table tags (/Table,/TD,/TH) to bypass visual heuristics and extract accurate table structures. - Pipeline Integration:
src/markdown/convert.rsuses structure tree data to generate semantic Markdown, falling back to visual heuristics only for untagged documents.
Frequently Asked Questions
What is a tagged PDF?
A tagged PDF is a PDF document that includes a structure tree (rooted at /StructTreeRoot in the catalog) containing semantic markup that describes the document's logical structure—such as headings, paragraphs, tables, and lists—separate from the visual presentation. This allows software to understand that a piece of text is a heading rather than just a large font on a page.
How does PDF-Inspector handle untagged PDFs?
When StructTree::from_doc returns None (indicating no /StructTreeRoot entry), PDF-Inspector falls back to visual heuristics. It analyzes coordinates, font sizes, and spatial relationships to detect columns, tables, and headings using geometric algorithms rather than semantic markup.
What are MCIDs and why do they matter?
MCIDs (Marked Content IDs) are numeric identifiers in PDF content streams that link specific text runs to nodes in the structure tree. PDF-Inspector uses mcid_to_roles to build a lookup table mapping each MCID to its semantic role (e.g., H1, TD). This connection allows the extractor to tag individual text items with their correct structural meaning during the content extraction phase.
Can PDF-Inspector extract tables from tagged PDFs without visual heuristics?
Yes. When a tagged PDF contains explicit table markup (/Table, /TR, /TH, /TD), the extract_tables method walks the structure tree and collects MCIDs for each cell directly. This semantic table extraction bypasses geometry-based detectors entirely, producing more accurate results for documents that include proper table tagging.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →