# How Tagged PDF Structure Tree Parsing Works in pdf-inspector for Semantic Roles (H1-H6, P, Code)

> Explore how pdf-inspector parses tagged PDF structure trees for semantic roles like H1-H6, P, and Code. Learn how it maps MCIDs to roles for accurate markdown conversion.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-06

---

**pdf-inspector extracts semantic meaning from tagged PDFs by parsing the `/StructTreeRoot` structure tree and mapping each Marked-Content ID (MCID) to its structural role, enabling accurate markdown conversion of headings, paragraphs, and code blocks.**

The `firecrawl/pdf-inspector` Rust library treats PDFs as structured documents rather than plain text streams. For tagged PDFs—those containing a `/StructTreeRoot` in their document catalog—the library builds an in-memory representation that preserves the original author's intent for semantic elements. This article explains how pdf-inspector parses the structure tree to identify **heading levels (H1-H6)**, **paragraphs (P)**, and **code blocks (Code)**.

## Detecting and Loading a Tagged PDF

The parsing pipeline begins with `StructTree::from_doc` in [`src/structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/structure_tree.rs). This function inspects the document catalog for a `/StructTreeRoot` entry. If present, the PDF is tagged and parsing proceeds; if absent, the library falls back to visual heuristics.

```rust
// src/structure_tree.rs, lines 42-48
let struct_tree_root = catalog
    .get(b"StructTreeRoot")
    .and_then(Object::as_reference)
    .and_then(|id| doc.get_object(id).ok())?;

```

Once confirmed, the parser loads custom role mappings via `parse_role_map` (lines 59-74). PDFs can define non-standard tags through a `/RoleMap` dictionary that maps custom names to standard PDF structure types. The function builds a `HashMap<String, String>` for resolving these aliases during parsing.

## Walking the Structure Tree Recursively

The core parsing logic walks the `/K` (kids) entries recursively. Two functions handle this traversal:

- **`parse_kids`** (lines 82-108) iterates over the `/K` array or dictionary, tracking page references and depth
- **`parse_kid`** (lines 119-148) dispatches each child to appropriate handlers based on its type

The `parse_kid` function distinguishes three cases:

- **Integer** → a bare MCID, wrapped in a `Span` element
- **Dictionary** → either a structure element or a marked-content reference (MCR)
- **Stream** → rarely used structural elements

```rust
// From parse_kid: handling different child types
match kid {
    Object::Integer(mcid) => {
        // Bare MCID - wrap in Span
        spans.push(Span::new(mcid as i64, current_page_id));
    }
    Object::Dictionary(dict) => {
        // Could be MCR, OBJR, or struct element
        elements.extend(parse_struct_element_dict(...));
    }
    // ...
}

```

## Parsing Structure Elements and Resolving Roles

When `parse_struct_element_dict` (lines 150-207) encounters a genuine structure element, it performs several operations:

1. **Detects MCR/OBJR references** by checking `/Type`
2. **Reads the `/S` (role) name** from the dictionary
3. **Resolves the role** through `StructRole::from_name_with_role_map`
4. **Extracts optional attributes**: `Alt`, `ActualText`, `Lang`
5. **Recurses into nested `/K` children**

The `StructRole` enum (lines 16-61) enumerates all standard PDF semantic roles:

```rust
pub enum StructRole {
    H1, H2, H3, H4, H5, H6,  // Heading levels
    P,                       // Paragraph
    Code,                    // Code block
    // ... additional roles (L, LI, Table, Figure, etc.)
}

```

Role resolution follows the role-map chain with a limit of **8 hops** to prevent infinite loops from circular mappings.

### Mapping Standard Roles to Output

| PDF Role | StructRole Variant | Markdown Output |
|----------|-------------------|-----------------|
| `/S /H1` | `StructRole::H1` | `# Heading` |

| `/S /H2` | `StructRole::H2` | `## Heading` |

| `/S /H3` → `/S /H6` | `StructRole::H3` → `H6` | `###` → `######` |
| `/S /P` | `StructRole::P` | Plain paragraph |
| `/S /Code` | `StructRole::Code` | Fenced code block |

The `Code` role receives special handling through `is_non_heading_content` (lines 78-96). This prevents the visual heuristic from mistakenly promoting short, bold code fragments to headings—a common false positive in untagged PDFs.

## Building the MCID-to-Role Lookup Table

After tree construction, `mcid_to_roles` (lines 66-82) flattens the hierarchy into a per-page lookup structure:

```rust
// HashMap structure: page_number -> HashMap<mcid, StructRole>
type McidRoleMap = HashMap<i64, HashMap<i64, StructRole>>;

let mcid_roles = struct_tree.mcid_to_roles(&page_ids);

```

This map enables O(1) role lookup during text extraction. When the extractor encounters a text item with MCID 42 on page 1, it queries `mcid_roles[1][42]` to retrieve the semantic role instantly.

## Practical Code Example

```rust
use pdf_inspector::structure_tree::{StructTree, StructRole};

// Load PDF and parse structure tree
let doc = lopdf::Document::load("tagged_document.pdf").unwrap();
let struct_tree = StructTree::from_doc(&doc).expect("PDF must be tagged");

// Build MCID → role mapping
let page_ids = doc.get_pages();
let mcid_roles = struct_tree.mcid_to_roles(&page_ids);

// Iterate flattened structure for processing
for elem in struct_tree.flatten() {
    match elem.role {
        StructRole::Code => {
            println!("Code block: {} MCIDs", elem.content_refs.len());
        }
        StructRole::H1 | StructRole::H2 | StructRole::H3 => {
            println!("Heading {:?} at depth {}", elem.role, elem.depth);
        }
        StructRole::P => {
            println!("Paragraph with {} MCIDs", elem.content_refs.len());
        }
        _ => {}
    }
}

```

The `flatten` method (lines 15-22) produces a linear `Vec<FlatStructElement>` preserving order, depth, and MCID information—ideal for simple traversals like generating document outlines.

## Integration with Markdown Conversion

The parsed roles flow into [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs), which makes rendering decisions:

- **Heading roles (`H1`-`H6`)** → emit corresponding markdown heading markers
- **`StructRole::P`** → wrap content as plain text blocks
- **`StructRole::Code`** → wrap in triple backticks with optional language detection

This integration ensures that a PDF tagged with proper semantic structure converts to semantically equivalent markdown without visual guessing.

## Summary

- **Detection**: `StructTree::from_doc` identifies tagged PDFs by locating `/StructTreeRoot`
- **Role mapping**: `parse_role_map` handles custom tag aliases via `/RoleMap`
- **Traversal**: `parse_kids` and `parse_kid` recursively walk the `/K` hierarchy
- **Resolution**: `StructRole::from_name_with_role_map` converts PDF names to typed enum variants
- **Lookup**: `mcid_to_roles` builds per-page HashMaps for O(1) MCID-to-role access
- **Protection**: `is_non_heading_content` prevents false heading promotion for `Code` and list elements

## Frequently Asked Questions

### What happens if a PDF is not tagged?

The parser returns `None` from `StructTree::from_doc`, and the library falls back to visual heuristics based on font size, weight, and positioning. These heuristics are less accurate than structure-tree parsing but handle legacy PDFs.

### How does pdf-inspector handle custom role names in the `/RoleMap`?

The `from_name_with_role_map` function follows alias chains up to 8 hops deep. If a PDF defines `/MyHeading /H1` in its role map, the parser resolves `MyHeading` to `StructRole::H1` for proper heading treatment.

### Can the structure tree parser handle deeply nested documents?

Yes. The recursive `parse_kids`/`parse_kid` functions maintain depth tracking throughout traversal. The `flatten` output includes `depth` fields for reconstructing hierarchy, and stack depth is naturally bounded by PDF implementation limits.

### Why is `Code` specifically protected from heading promotion?

Short monospaced text fragments often appear visually similar to headings—bold, compact, and isolated. The `is_non_heading_content` check ensures these retain their semantic identity as code rather than being falsely promoted to `H1`-`H6` during markdown conversion.