# How the `use_struct_tree` Option Leverages Tagged PDF Structure Tags in OpenDataLoader

> Learn how OpenDataLoader's use_struct_tree option extracts semantic structure like headings and tables directly from PDF Structure Trees bypassing visual layout inference.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: internals
- Published: 2026-03-20

---

**The `use_struct_tree` option enables OpenDataLoader to extract semantic document structure—such as headings, paragraphs, and tables—directly from a PDF's internal Structure Tree rather than inferring layout from visual coordinates.**

OpenDataLoader is an open-source PDF extraction engine maintained in the `opendataloader-project/opendataloader-pdf` repository. When you enable the `use_struct_tree` flag, the system bypasses traditional XY‑Cut++ heuristics and instead parses the **Tagged PDF** markup that authors embed to define logical reading order and content roles.

## What Are Tagged PDF Structure Tags?

**Tagged PDF** is an ISO 32000 standard feature where PDF documents contain a **Structure Tree** (also called the *Structure Tree Root*). This tree maps content streams to semantic types such as `HEADING`, `PARAGRAPH`, `LIST`, `TABLE`, `CODE`, and `FOOTNOTE`.

Unlike visual analysis, which guesses structure from font sizes and coordinates, Tagged PDF structure tags carry the author’s explicit intent. When `use_struct_tree` is active, OpenDataLoader traverses this tree via `ITree` and `IStructElem` interfaces to build `SemanticHeading`, `SemanticParagraph`, and other typed objects.

## Configuration and API Usage

### Enabling the Option in Python

The Python wrapper exposes the flag as `use_struct_tree`, which internally calls `Config#setUseStructTree(true)`.

```python
import opendataloader_pdf

# Convert PDFs using Tagged PDF structure extraction

opendataloader_pdf.convert(
    input_path=["report1.pdf", "report2.pdf"],
    output_dir="extracted/",
    use_struct_tree=True,  # Leverage Tagged PDF structure tags

)

```

### Command Line Interface

The CLI flag `--use-struct-tree` is defined in [`node/opendataloader-pdf/src/cli-options.generated.ts`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/node/opendataloader-pdf/src/cli-options.generated.ts) at lines 18‑19, mapping directly to the configuration object.

```bash
opendataloader-pdf mydoc.pdf \
  --output-dir extracted/ \
  --use-struct-tree

```

### Java Programmatic Access

In the core Java backend, you interact with [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java) (lines 66‑71 declare the field; lines 350‑360 define accessors).

```java
import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.api.PdfLoader;

Config cfg = new Config();
cfg.setUseStructTree(true);  // Enable Tagged PDF processing
cfg.setGenerateMarkdown(true);
cfg.setOutputFolder("extracted/");

PdfLoader.convert("mydoc.pdf", cfg);

```

## Internal Implementation Flow

### Config.java – The Source of Truth

The boolean flag `useStructTree` resides in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java). It defaults to `false` to maintain backward compatibility with heuristic-based extraction. The `isUseStructTree()` getter and `setUseStructTree(boolean)` setter allow the CLI, Python bindings, and Java API to toggle the behavior uniformly.

### Preprocessing and Tree Detection

During document initialization, `DocumentProcessor.preprocessing()` (lines 54‑60) checks the configuration:

```java
if (config.isUseStructTree()) {
    document.parseStructureTreeRoot();  // Parses the PDF Structure Tree Root
    if (document.getTree() != null) {
        StaticLayoutContainers.setIsUseStructTree(true);
    } else {
        StaticLayoutContainers.setIsUseStructTree(false);
        LOGGER.log(Level.WARNING,
            "The document has no structure tree. The 'use-struct-tree' option will be ignored.");
    }
}

```

If the PDF lacks a structure tree, the system logs a warning and falls back to visual heuristics.

### Pipeline Selection in DocumentProcessor

The `processFile()` method (lines 78‑82) uses the thread‑local flag set during preprocessing to choose the extraction pipeline:

```java
if (StaticLayoutContainers.isUseStructTree()) {
    contents = TaggedDocumentProcessor.processDocument(inputPdfName, config, pagesToProcess);
} else if (config.isHybridEnabled()) {
    // Hybrid backend path
} else {
    contents = processDocument(inputPdfName, config, pagesToProcess);
}

```

When `isUseStructTree()` returns `true`, the `TaggedDocumentProcessor` receives control.

### TaggedDocumentProcessor and Semantic Extraction

`TaggedDocumentProcessor` (defined in [`TaggedDocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/TaggedDocumentProcessor.java)) operates on the `ITree` interface:

```java
ITree tree = StaticContainers.getDocument().getTree();  // Retrieve parsed structure tree
processStructElem(tree.getRoot());                       // Recursive traversal

```

The `processStructElem` method switches on `node.getInitialSemanticType()` to instantiate concrete semantic objects:

- `SemanticHeading` for `HEADING` tags
- `SemanticParagraph` for `PARAGRAPH` tags  
- `PDFList` for `LIST` tags
- `Table` for `TABLE` tags

This preserves the author’s intended reading order and hierarchy, eliminating the guesswork required by XY‑Cut++ layout analysis.

## Fallback Behavior and Edge Cases

If a user enables `use_struct_tree` but the input PDF is not tagged, OpenDataLoader logs a warning at `Level.WARNING` and continues with the standard visual pipeline. This ensures **graceful degradation**—you can batch‑process mixed collections of tagged and untagged PDFs without manual filtering.

Additionally, the engine adjusts **Marked Content ID (MCID)** handling based on the flag. When structure tags are present, `StaticStorages.setIsIgnoreMCIDs(!StaticLayoutContainers.isUseStructTree())` prevents duplicate text extraction from both the structure tree and the content stream (lines 66‑67 in [`DocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/DocumentProcessor.java)).

## Impact on Downstream Processing

Using `use_struct_tree` does not alter the final output formats—Markdown, HTML, JSON, and plain text remain available. However, the **quality** of the output improves in tagged documents:

- **Exact reading order**: The extraction follows the `StructTreeRoot` order rather than top‑to‑bottom coordinate scanning.
- **Semantic markup**: Headings retain their level hierarchy, lists preserve nesting, and tables maintain cell relationships.
- **Accessibility alignment**: Output mirrors the PDF/UA (Universal Accessibility) intent, producing documents that screen readers would navigate naturally.

Downstream modules—such as header/footer detection, table border processing, and list item clustering—receive the same `SemanticDocument` objects regardless of whether they originated from `TaggedDocumentProcessor` or the heuristic pipeline, ensuring consistent API behavior.

## Summary

- The **`use_struct_tree`** option (Java: `useStructTree`) activates **Tagged PDF** extraction in OpenDataLoader, reading semantic structure directly from the PDF Structure Tree.
- When enabled, `DocumentProcessor` parses the tree during preprocessing; if present, `TaggedDocumentProcessor` handles the extraction, creating `SemanticHeading`, `SemanticParagraph`, and other typed objects.
- If the PDF lacks a structure tree, the engine logs a warning and falls back to visual heuristics, ensuring robust batch processing.
- The flag is exposed uniformly across **Python** (`use_struct_tree=True`), **CLI** (`--use-struct-tree`), and **Java** (`setUseStructTree(true)`).

## Frequently Asked Questions

### What is the difference between `use_struct_tree` and the default layout analysis?

**The default pipeline uses visual heuristics (XY‑Cut++) to guess document structure from coordinates and font sizes, while `use_struct_tree` reads the author‑provided semantic tags embedded in Tagged PDFs.** The tagged approach preserves exact reading order and hierarchical relationships (e.g., heading levels, list nesting) that visual analysis might misinterpret, especially in complex multi‑column layouts.

### Can I use `use_struct_tree` with untagged PDFs?

**Yes, but it will have no effect.** During preprocessing, `DocumentProcessor` checks for the presence of a Structure Tree Root. If the PDF is not tagged, the engine logs a warning stating *"The document has no structure tree. The 'use-struct-tree' option will be ignored"* and proceeds with the standard visual extraction pipeline, ensuring you still receive usable output.

### How does the option affect table and list extraction?

**When `use_struct_tree` is enabled and the PDF is properly tagged, tables and lists are instantiated from `TABLE` and `LIST` structure elements rather than inferred from line drawings or indentation patterns.** `TaggedDocumentProcessor` creates `Table` and `PDFList` semantic objects directly from the tree nodes, preserving cell relationships and list hierarchies exactly as marked by the document author, which reduces errors in merged cells or nested lists.

### Is there a performance difference when using Tagged PDF extraction?

**Parsing the Structure Tree adds a small upfront cost during preprocessing (`parseStructureTreeRoot`), but subsequent processing is often faster because the engine skips complex visual heuristics.** The `TaggedDocumentProcessor` performs a single recursive traversal of the tree (`processStructElem`) rather than multiple passes of XY‑Cut clustering, which can improve throughput on documents with complex layouts while delivering higher accuracy.