How the `use_struct_tree` Option Leverages Tagged PDF Structure Tags in OpenDataLoader

The use_struct_tree option enables OpenDataLoader to extract semantic document structure—such as headings, paragraphs, and tables—directly from a PDF's internal Structure Tree rather than inferring layout from visual coordinates.

OpenDataLoader is an open-source PDF extraction engine maintained in the opendataloader-project/opendataloader-pdf repository. When you enable the use_struct_tree flag, the system bypasses traditional XY‑Cut++ heuristics and instead parses the Tagged PDF markup that authors embed to define logical reading order and content roles.

What Are Tagged PDF Structure Tags?

Tagged PDF is an ISO 32000 standard feature where PDF documents contain a Structure Tree (also called the Structure Tree Root). This tree maps content streams to semantic types such as HEADING, PARAGRAPH, LIST, TABLE, CODE, and FOOTNOTE.

Unlike visual analysis, which guesses structure from font sizes and coordinates, Tagged PDF structure tags carry the author’s explicit intent. When use_struct_tree is active, OpenDataLoader traverses this tree via ITree and IStructElem interfaces to build SemanticHeading, SemanticParagraph, and other typed objects.

Configuration and API Usage

Enabling the Option in Python

The Python wrapper exposes the flag as use_struct_tree, which internally calls Config#setUseStructTree(true).

import opendataloader_pdf

# Convert PDFs using Tagged PDF structure extraction

opendataloader_pdf.convert(
    input_path=["report1.pdf", "report2.pdf"],
    output_dir="extracted/",
    use_struct_tree=True,  # Leverage Tagged PDF structure tags

)

Command Line Interface

The CLI flag --use-struct-tree is defined in node/opendataloader-pdf/src/cli-options.generated.ts at lines 18‑19, mapping directly to the configuration object.

opendataloader-pdf mydoc.pdf \
  --output-dir extracted/ \
  --use-struct-tree

Java Programmatic Access

In the core Java backend, you interact with Config.java (lines 66‑71 declare the field; lines 350‑360 define accessors).

import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.api.PdfLoader;

Config cfg = new Config();
cfg.setUseStructTree(true);  // Enable Tagged PDF processing
cfg.setGenerateMarkdown(true);
cfg.setOutputFolder("extracted/");

PdfLoader.convert("mydoc.pdf", cfg);

Internal Implementation Flow

Config.java – The Source of Truth

The boolean flag useStructTree resides in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java. It defaults to false to maintain backward compatibility with heuristic-based extraction. The isUseStructTree() getter and setUseStructTree(boolean) setter allow the CLI, Python bindings, and Java API to toggle the behavior uniformly.

Preprocessing and Tree Detection

During document initialization, DocumentProcessor.preprocessing() (lines 54‑60) checks the configuration:

if (config.isUseStructTree()) {
    document.parseStructureTreeRoot();  // Parses the PDF Structure Tree Root
    if (document.getTree() != null) {
        StaticLayoutContainers.setIsUseStructTree(true);
    } else {
        StaticLayoutContainers.setIsUseStructTree(false);
        LOGGER.log(Level.WARNING,
            "The document has no structure tree. The 'use-struct-tree' option will be ignored.");
    }
}

If the PDF lacks a structure tree, the system logs a warning and falls back to visual heuristics.

Pipeline Selection in DocumentProcessor

The processFile() method (lines 78‑82) uses the thread‑local flag set during preprocessing to choose the extraction pipeline:

if (StaticLayoutContainers.isUseStructTree()) {
    contents = TaggedDocumentProcessor.processDocument(inputPdfName, config, pagesToProcess);
} else if (config.isHybridEnabled()) {
    // Hybrid backend path
} else {
    contents = processDocument(inputPdfName, config, pagesToProcess);
}

When isUseStructTree() returns true, the TaggedDocumentProcessor receives control.

TaggedDocumentProcessor and Semantic Extraction

TaggedDocumentProcessor (defined in TaggedDocumentProcessor.java) operates on the ITree interface:

ITree tree = StaticContainers.getDocument().getTree();  // Retrieve parsed structure tree
processStructElem(tree.getRoot());                       // Recursive traversal

The processStructElem method switches on node.getInitialSemanticType() to instantiate concrete semantic objects:

  • SemanticHeading for HEADING tags
  • SemanticParagraph for PARAGRAPH tags
  • PDFList for LIST tags
  • Table for TABLE tags

This preserves the author’s intended reading order and hierarchy, eliminating the guesswork required by XY‑Cut++ layout analysis.

Fallback Behavior and Edge Cases

If a user enables use_struct_tree but the input PDF is not tagged, OpenDataLoader logs a warning at Level.WARNING and continues with the standard visual pipeline. This ensures graceful degradation—you can batch‑process mixed collections of tagged and untagged PDFs without manual filtering.

Additionally, the engine adjusts Marked Content ID (MCID) handling based on the flag. When structure tags are present, StaticStorages.setIsIgnoreMCIDs(!StaticLayoutContainers.isUseStructTree()) prevents duplicate text extraction from both the structure tree and the content stream (lines 66‑67 in DocumentProcessor.java).

Impact on Downstream Processing

Using use_struct_tree does not alter the final output formats—Markdown, HTML, JSON, and plain text remain available. However, the quality of the output improves in tagged documents:

  • Exact reading order: The extraction follows the StructTreeRoot order rather than top‑to‑bottom coordinate scanning.
  • Semantic markup: Headings retain their level hierarchy, lists preserve nesting, and tables maintain cell relationships.
  • Accessibility alignment: Output mirrors the PDF/UA (Universal Accessibility) intent, producing documents that screen readers would navigate naturally.

Downstream modules—such as header/footer detection, table border processing, and list item clustering—receive the same SemanticDocument objects regardless of whether they originated from TaggedDocumentProcessor or the heuristic pipeline, ensuring consistent API behavior.

Summary

  • The use_struct_tree option (Java: useStructTree) activates Tagged PDF extraction in OpenDataLoader, reading semantic structure directly from the PDF Structure Tree.
  • When enabled, DocumentProcessor parses the tree during preprocessing; if present, TaggedDocumentProcessor handles the extraction, creating SemanticHeading, SemanticParagraph, and other typed objects.
  • If the PDF lacks a structure tree, the engine logs a warning and falls back to visual heuristics, ensuring robust batch processing.
  • The flag is exposed uniformly across Python (use_struct_tree=True), CLI (--use-struct-tree), and Java (setUseStructTree(true)).

Frequently Asked Questions

What is the difference between use_struct_tree and the default layout analysis?

The default pipeline uses visual heuristics (XY‑Cut++) to guess document structure from coordinates and font sizes, while use_struct_tree reads the author‑provided semantic tags embedded in Tagged PDFs. The tagged approach preserves exact reading order and hierarchical relationships (e.g., heading levels, list nesting) that visual analysis might misinterpret, especially in complex multi‑column layouts.

Can I use use_struct_tree with untagged PDFs?

Yes, but it will have no effect. During preprocessing, DocumentProcessor checks for the presence of a Structure Tree Root. If the PDF is not tagged, the engine logs a warning stating "The document has no structure tree. The 'use-struct-tree' option will be ignored" and proceeds with the standard visual extraction pipeline, ensuring you still receive usable output.

How does the option affect table and list extraction?

When use_struct_tree is enabled and the PDF is properly tagged, tables and lists are instantiated from TABLE and LIST structure elements rather than inferred from line drawings or indentation patterns. TaggedDocumentProcessor creates Table and PDFList semantic objects directly from the tree nodes, preserving cell relationships and list hierarchies exactly as marked by the document author, which reduces errors in merged cells or nested lists.

Is there a performance difference when using Tagged PDF extraction?

Parsing the Structure Tree adds a small upfront cost during preprocessing (parseStructureTreeRoot), but subsequent processing is often faster because the engine skips complex visual heuristics. The TaggedDocumentProcessor performs a single recursive traversal of the tree (processStructElem) rather than multiple passes of XY‑Cut clustering, which can improve throughput on documents with complex layouts while delivering higher accuracy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →