# How PDF-Inspector Handles Tagged PDFs Using the Structure Tree

> Discover how PDF-Inspector uses the structure tree to process tagged PDFs. Learn how MCIDs and semantic roles enable precise Markdown generation beyond visual cues.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-08

---

**PDF-Inspector detects tagged PDFs by locating the `/StructTreeRoot` entry in the document catalog and uses the structure tree to map MCIDs to semantic roles, enabling accurate Markdown generation without relying solely on visual heuristics.**

PDF-Inspector is an open-source Rust library for converting PDFs to structured Markdown. When processing **tagged PDFs**, it leverages the embedded **structure tree** to extract semantic meaning directly from the document's markup rather than inferring layout from coordinates. This article explores how the library parses the `/StructTreeRoot`, maps structure elements to content streams, and integrates semantic roles into the extraction pipeline according to the `firecrawl/pdf-inspector` source code.

## Detecting Tagged PDFs and Loading the Structure Tree

PDF-Inspector determines whether a PDF contains semantic markup by checking for the presence of a `/StructTreeRoot` entry in the document catalog. In [`src/structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/structure_tree.rs), the `StructTree::from_doc` method performs this detection:

```rust
// src/structure_tree.rs – lines 44-49
pub fn from_doc(doc: &Document) -> Option<Self> {
    let catalog = doc.catalog().ok()?;
    let struct_root_obj = catalog.get(b"StructTreeRoot").ok()?;
    …
}

```

If the `/StructTreeRoot` is missing, the method returns `None`, signaling that the PDF is untagged. In this case, the extractor falls back to visual heuristics such as column detection and line-based table detection. When the root is present, the parser proceeds to build the complete structure tree.

## Parsing Standard Structure Roles

Once a tagged PDF is detected, PDF-Inspector interprets the semantic intent of each element using the `StructRole` enum defined in [`src/structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/structure_tree.rs) (lines 16-75). This enum maps every ISO-defined standard role—such as `H1`, `P`, `Table`, `TD`, and `TH`—and includes a catch-all `Other(String)` variant for custom tags.

The library resolves custom tag names to standard roles through a role map, ensuring that even documents with specialized tagging schemes produce consistent output. You can query roles via `StructRole::from_name`, which normalizes the incoming tag strings against the ISO vocabulary.

## Building the Structure Tree

The tree construction process recursively traverses the `/K` (kids) entries starting from the `StructTreeRoot`. Each node becomes a `StructElement` that stores:

- Its semantic role
- Optional alt-text and actual-text properties
- Language attributes
- MCID (Marked Content ID) references
- Child element relationships

The parser in [`src/structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/structure_tree.rs) (lines 82-197) includes defensive measures against malformed PDFs, enforcing depth limits and properly handling `MCR` (Marked Content Reference) and `OBJR` (Object Reference) entries. This recursive descent builds an in-memory representation of the document's logical reading order, independent of the physical page layout.

## Mapping MCIDs to Semantic Roles

After constructing the tree, PDF-Inspector creates a lookup mechanism to connect structure elements with their corresponding content in the page streams. The `StructTree::mcid_to_roles` method (lines 66-81) walks the tree and generates a per-page `HashMap` mapping each MCID to its `StructRole`:

```rust
use std::collections::BTreeMap;
use pdf_inspector::structure_tree::StructTree;

// Assume `doc` is a lopdf::Document already loaded
let page_ids: BTreeMap<u32, ObjectId> = doc.get_pages(); // page number → ObjectId
let struct_tree = StructTree::from_doc(&doc).unwrap();
let mcid_roles = struct_tree.mcid_to_roles(&page_ids);

// `mcid_roles[page][mcid]` now yields the `StructRole` (e.g. H1, P, TD)

```

This map is consulted by the content-stream extractor so that each extracted `TextItem` can be annotated with its semantic role, preventing visual heuristics from misclassifying short captions or table headers as headings.

## Flattening for Linear Processing

To facilitate linear output formats like Markdown, the library provides `StructTree::flatten` (lines 115-124), which produces a vector of `FlatStructElement` objects preserving the document's reading order. This flattened view eliminates the need for a separate layout analysis pass when generating Markdown.

```rust
let flat = struct_tree.flatten();
for elem in flat {
    println!("Depth {} – role: {:?}", elem.depth, elem.role);
}

```

The depth information maintains the hierarchical relationship between elements, allowing the Markdown generator to correctly nest headings and list items.

## Table Extraction from Semantic Tags

When a tagged PDF contains explicit table markup (`/Table`, `/TR`, `/TH`, `/TD`), PDF-Inspector bypasses geometry-based detection in favor of semantic extraction. The `StructTree::extract_tables` method (lines 123-139) walks the structure tree, collects MCIDs for each cell, and returns a `StructTable` structure:

```rust
let tables = struct_tree.extract_tables(&page_ids);
for (i, table) in tables.iter().enumerate() {
    println!("Table {} has {} rows", i + 1, table.rows.len());
}

```

This approach yields more reliable table reconstructions on documents that already describe them semantically, avoiding the errors common in purely visual table detectors.

## Integration with Markdown Generation

The Markdown conversion layer in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) (lines 6-12 and 1529-1534) consumes both the MCID-role map and the flattened structure to drive output decisions. By consulting the `StructRole` information, the converter distinguishes between body text, headings, list items, and code blocks without guessing based on font size or positioning.

The [`src/markdown/preprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/preprocess.rs) module additionally uses `StructRole` to influence pre-processing steps such as drop-cap merging and heading line merging, ensuring that the final Markdown accurately reflects the document's intended semantic structure.

## Summary

- **Detection**: PDF-Inspector checks for `/StructTreeRoot` in the document catalog via `StructTree::from_doc` to identify tagged PDFs.
- **Role Mapping**: The `StructRole` enum standardizes ISO-defined tags and custom roles, mapped through `mcid_to_roles` to connect structure elements with content streams.
- **Tree Parsing**: Recursive parsing of `/K` entries builds a complete `StructElement` hierarchy with MCID references, alt-text, and language attributes.
- **Linear Output**: The `flatten` method produces `FlatStructElement` vectors that preserve reading order for Markdown generation.
- **Semantic Tables**: `extract_tables` uses explicit table tags (`/Table`, `/TD`, `/TH`) to bypass visual heuristics and extract accurate table structures.
- **Pipeline Integration**: [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) uses structure tree data to generate semantic Markdown, falling back to visual heuristics only for untagged documents.

## Frequently Asked Questions

### What is a tagged PDF?

A tagged PDF is a PDF document that includes a **structure tree** (rooted at `/StructTreeRoot` in the catalog) containing semantic markup that describes the document's logical structure—such as headings, paragraphs, tables, and lists—separate from the visual presentation. This allows software to understand that a piece of text is a heading rather than just a large font on a page.

### How does PDF-Inspector handle untagged PDFs?

When `StructTree::from_doc` returns `None` (indicating no `/StructTreeRoot` entry), PDF-Inspector falls back to **visual heuristics**. It analyzes coordinates, font sizes, and spatial relationships to detect columns, tables, and headings using geometric algorithms rather than semantic markup.

### What are MCIDs and why do they matter?

**MCIDs** (Marked Content IDs) are numeric identifiers in PDF content streams that link specific text runs to nodes in the structure tree. PDF-Inspector uses `mcid_to_roles` to build a lookup table mapping each MCID to its semantic role (e.g., `H1`, `TD`). This connection allows the extractor to tag individual text items with their correct structural meaning during the content extraction phase.

### Can PDF-Inspector extract tables from tagged PDFs without visual heuristics?

Yes. When a tagged PDF contains explicit table markup (`/Table`, `/TR`, `/TH`, `/TD`), the `extract_tables` method walks the structure tree and collects MCIDs for each cell directly. This **semantic table extraction** bypasses geometry-based detectors entirely, producing more accurate results for documents that include proper table tagging.