How Tagged PDF Structure Tree Parsing Works in pdf-inspector for Semantic Roles (H1-H6, P, Code)

pdf-inspector extracts semantic meaning from tagged PDFs by parsing the /StructTreeRoot structure tree and mapping each Marked-Content ID (MCID) to its structural role, enabling accurate markdown conversion of headings, paragraphs, and code blocks.

The firecrawl/pdf-inspector Rust library treats PDFs as structured documents rather than plain text streams. For tagged PDFs—those containing a /StructTreeRoot in their document catalog—the library builds an in-memory representation that preserves the original author's intent for semantic elements. This article explains how pdf-inspector parses the structure tree to identify heading levels (H1-H6), paragraphs (P), and code blocks (Code).

Detecting and Loading a Tagged PDF

The parsing pipeline begins with StructTree::from_doc in src/structure_tree.rs. This function inspects the document catalog for a /StructTreeRoot entry. If present, the PDF is tagged and parsing proceeds; if absent, the library falls back to visual heuristics.

// src/structure_tree.rs, lines 42-48
let struct_tree_root = catalog
    .get(b"StructTreeRoot")
    .and_then(Object::as_reference)
    .and_then(|id| doc.get_object(id).ok())?;

Once confirmed, the parser loads custom role mappings via parse_role_map (lines 59-74). PDFs can define non-standard tags through a /RoleMap dictionary that maps custom names to standard PDF structure types. The function builds a HashMap<String, String> for resolving these aliases during parsing.

Walking the Structure Tree Recursively

The core parsing logic walks the /K (kids) entries recursively. Two functions handle this traversal:

  • parse_kids (lines 82-108) iterates over the /K array or dictionary, tracking page references and depth
  • parse_kid (lines 119-148) dispatches each child to appropriate handlers based on its type

The parse_kid function distinguishes three cases:

  • Integer → a bare MCID, wrapped in a Span element
  • Dictionary → either a structure element or a marked-content reference (MCR)
  • Stream → rarely used structural elements
// From parse_kid: handling different child types
match kid {
    Object::Integer(mcid) => {
        // Bare MCID - wrap in Span
        spans.push(Span::new(mcid as i64, current_page_id));
    }
    Object::Dictionary(dict) => {
        // Could be MCR, OBJR, or struct element
        elements.extend(parse_struct_element_dict(...));
    }
    // ...
}

Parsing Structure Elements and Resolving Roles

When parse_struct_element_dict (lines 150-207) encounters a genuine structure element, it performs several operations:

  1. Detects MCR/OBJR references by checking /Type
  2. Reads the /S (role) name from the dictionary
  3. Resolves the role through StructRole::from_name_with_role_map
  4. Extracts optional attributes: Alt, ActualText, Lang
  5. Recurses into nested /K children

The StructRole enum (lines 16-61) enumerates all standard PDF semantic roles:

pub enum StructRole {
    H1, H2, H3, H4, H5, H6,  // Heading levels
    P,                       // Paragraph
    Code,                    // Code block
    // ... additional roles (L, LI, Table, Figure, etc.)
}

Role resolution follows the role-map chain with a limit of 8 hops to prevent infinite loops from circular mappings.

Mapping Standard Roles to Output

PDF Role StructRole Variant Markdown Output
/S /H1 StructRole::H1 # Heading

| /S /H2 | StructRole::H2 | ## Heading |

| /S /H3 → /S /H6 | StructRole::H3 → H6 | ### → ###### | | /S /P | StructRole::P | Plain paragraph | | /S /Code | StructRole::Code | Fenced code block |

The Code role receives special handling through is_non_heading_content (lines 78-96). This prevents the visual heuristic from mistakenly promoting short, bold code fragments to headings—a common false positive in untagged PDFs.

Building the MCID-to-Role Lookup Table

After tree construction, mcid_to_roles (lines 66-82) flattens the hierarchy into a per-page lookup structure:

// HashMap structure: page_number -> HashMap<mcid, StructRole>
type McidRoleMap = HashMap<i64, HashMap<i64, StructRole>>;

let mcid_roles = struct_tree.mcid_to_roles(&page_ids);

This map enables O(1) role lookup during text extraction. When the extractor encounters a text item with MCID 42 on page 1, it queries mcid_roles[1][42] to retrieve the semantic role instantly.

Practical Code Example

use pdf_inspector::structure_tree::{StructTree, StructRole};

// Load PDF and parse structure tree
let doc = lopdf::Document::load("tagged_document.pdf").unwrap();
let struct_tree = StructTree::from_doc(&doc).expect("PDF must be tagged");

// Build MCID → role mapping
let page_ids = doc.get_pages();
let mcid_roles = struct_tree.mcid_to_roles(&page_ids);

// Iterate flattened structure for processing
for elem in struct_tree.flatten() {
    match elem.role {
        StructRole::Code => {
            println!("Code block: {} MCIDs", elem.content_refs.len());
        }
        StructRole::H1 | StructRole::H2 | StructRole::H3 => {
            println!("Heading {:?} at depth {}", elem.role, elem.depth);
        }
        StructRole::P => {
            println!("Paragraph with {} MCIDs", elem.content_refs.len());
        }
        _ => {}
    }
}

The flatten method (lines 15-22) produces a linear Vec<FlatStructElement> preserving order, depth, and MCID information—ideal for simple traversals like generating document outlines.

Integration with Markdown Conversion

The parsed roles flow into src/markdown/convert.rs, which makes rendering decisions:

  • Heading roles (H1-H6) → emit corresponding markdown heading markers
  • StructRole::P → wrap content as plain text blocks
  • StructRole::Code → wrap in triple backticks with optional language detection

This integration ensures that a PDF tagged with proper semantic structure converts to semantically equivalent markdown without visual guessing.

Summary

  • Detection: StructTree::from_doc identifies tagged PDFs by locating /StructTreeRoot
  • Role mapping: parse_role_map handles custom tag aliases via /RoleMap
  • Traversal: parse_kids and parse_kid recursively walk the /K hierarchy
  • Resolution: StructRole::from_name_with_role_map converts PDF names to typed enum variants
  • Lookup: mcid_to_roles builds per-page HashMaps for O(1) MCID-to-role access
  • Protection: is_non_heading_content prevents false heading promotion for Code and list elements

Frequently Asked Questions

What happens if a PDF is not tagged?

The parser returns None from StructTree::from_doc, and the library falls back to visual heuristics based on font size, weight, and positioning. These heuristics are less accurate than structure-tree parsing but handle legacy PDFs.

How does pdf-inspector handle custom role names in the /RoleMap?

The from_name_with_role_map function follows alias chains up to 8 hops deep. If a PDF defines /MyHeading /H1 in its role map, the parser resolves MyHeading to StructRole::H1 for proper heading treatment.

Can the structure tree parser handle deeply nested documents?

Yes. The recursive parse_kids/parse_kid functions maintain depth tracking throughout traversal. The flatten output includes depth fields for reconstructing hierarchy, and stack depth is naturally bounded by PDF implementation limits.

Why is Code specifically protected from heading promotion?

Short monospaced text fragments often appear visually similar to headings—bold, compact, and isolated. The is_non_heading_content check ensures these retain their semantic identity as code rather than being falsely promoted to H1-H6 during markdown conversion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →