Performance Benchmarks for Extracting Complex PDFs in Hybrid Mode

OpenDataLoader PDF achieves 0.93 TEDS accuracy on complex tables and processes hybrid-mode pages at approximately 0.43 seconds per page while maintaining ≥95% triage recall, combining a fast Java-based local parser with selective AI backend invocation.

The opendataloader-project/opendataloader-pdf repository delivers a production-grade document extraction system that balances speed and accuracy through intelligent page routing. When working with performance benchmarks for extracting complex PDFs in hybrid mode, developers benefit from deterministic Java processing for simple content and AI-backed analysis for complex layouts containing tables, formulas, and scanned graphics.

Hybrid Architecture and Page Triage

The hybrid system operates through a sophisticated decision engine that minimizes unnecessary AI backend calls while maximizing extraction accuracy on difficult content.

Page-Level Triage Logic

At the core of the system, HybridDocumentProcessor.java (lines 544–560) implements the triage logic that examines each page individually. The decidePageRoute() method analyzes visual complexity indicators to determine whether a page can be handled locally or requires AI intervention. According to the source code in HybridDocumentProcessor.java#L544-L560, this decision happens in milliseconds before any heavy processing begins.

The triage decisions are persisted for benchmark analysis via TriageLogger.logToFile() (lines 551–564), writing JSON logs that enable the evaluation suite to calculate recall metrics and identify false negatives.

Dual-Path Processing Strategy

The architecture maintains two distinct execution paths:

  • Local path: Simple pages process entirely within the Java runtime at approximately 0.05 seconds per page, using deterministic parsing algorithms
  • Hybrid path: Complex pages batch to configurable AI backends (Docling, Hancom, or custom endpoints) returning structured JSON with bounding boxes, OCR results, and formula renderings

When the AI backend returns results, the Java side merges the structured output with local extraction data, producing unified Markdown or JSON documents.

Documented Performance Benchmarks

The repository maintains rigorous performance standards validated through continuous integration. The benchmark suite in tests/benchmark/ enforces specific thresholds for production deployments.

Overall Accuracy Metrics

  • Combined score (NID, TEDS, MHS): 0.90 (rank #1 on the benchmark leaderboard) according to README.md
  • Complex table extraction: 0.93 TEDS for borderless and structured tables, representing a significant improvement over the 0.49 TEDS score of the pure-Java implementation

Speed and Latency Measurements

  • Hybrid processing speed: Approximately 0.43 seconds per page (roughly 2 pages per second), accounting for network latency to the AI backend
  • Local baseline: 0.05 seconds per page for text-only or simple layout documents

Quality Assurance Thresholds

The tests/benchmark/thresholds.json file (lines 8–9) enforces strict CI requirements:

  • Triage recall: ≥ 0.95 for correctly identifying tables that require AI processing
  • False negatives: ≤ 5 per benchmark run (pages incorrectly routed to local processing)
  • Per-document elapsed time: ≤ 2.0 seconds maximum tolerated latency

Benchmark Suite Implementation

The performance validation relies on tests/benchmark/run.py, which orchestrates the evaluation pipeline across approximately 200 curated real-world PDFs stored in tests/benchmark/pdfs/.

Evaluation Components

  1. Reading order analysis: Validates Natural Reading Order Detection (NID) metrics
  2. Table structure evaluation: Calculates Table Detection and Structure recognition (TEDS) scores
  3. Heading hierarchy: Measures Metadata and Heading Structure (MHS) accuracy
  4. Speed aggregation: Generates summary.json with per-page and per-document timing data

The evaluator_triage.py script specifically analyzes triage.json against ground-truth reference.json to verify that the HybridDocumentProcessor correctly identifies complex content requiring AI assistance.

Enabling Hybrid Mode: Implementation Examples

Java API Configuration

Configure hybrid extraction programmatically using the builder pattern in HybridConfig.java:

import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.api.OpenDataLoaderPDF;
import java.nio.file.Path;

Config cfg = Config.builder()
    .setHybrid("docling-fast")
    .setHybridUrl("http://localhost:5001")
    .build();

Path pdf = Path.of("docs/complex-paper.pdf");
String markdown = OpenDataLoaderPDF.process(pdf, cfg);
System.out.println(markdown);

Command-Line Interface

Process complex PDFs from the shell with backend specification:


# Local-only processing for simple documents

opendataloader-pdf input.pdf -o out/

# Hybrid mode for complex layouts with tables and formulas

opendataloader-pdf --hybrid docling-fast \
    --hybrid-url http://localhost:5001 \
    input.pdf -o out/

Running Performance Benchmarks

Validate your deployment against the official benchmarks:


# Execute the benchmark suite with hybrid configuration

./scripts/bench.sh --hybrid docling-fast

# Review generated metrics

cat prediction/opendataloader-hybrid-docling-fast/summary.json
cat evaluation.json

After execution, examine prediction/opendataloader-hybrid-docling-fast/triage.json to review routing decisions, and evaluation.json for the complete accuracy report.

Summary

  • Hybrid mode in opendataloader-project/opendataloader-pdf achieves 0.93 TEDS on complex tables while maintaining 0.43s/page throughput
  • Page-level triage in HybridDocumentProcessor.java routes only complex content to AI backends, preserving the 0.05s/page speed for simple documents
  • CI thresholds enforce ≥95% triage recall and ≤2.0s per-document latency through tests/benchmark/thresholds.json
  • The system provides deterministic fallback to local processing when AI backends fail, ensuring robust production deployments
  • Benchmark validation occurs via run.py and evaluator_triage.py against a corpus of 200 real-world PDFs

Frequently Asked Questions

What is the processing speed difference between local mode and hybrid mode?

Local mode processes simple pages at approximately 0.05 seconds per page using pure Java parsing, while hybrid mode requires approximately 0.43 seconds per page due to AI backend latency and batching overhead. However, hybrid mode only activates for pages containing complex elements like large tables or scanned graphics, meaning mixed documents achieve blended throughput significantly faster than processing every page through the AI backend.

How does the system decide which pages require AI processing?

The HybridDocumentProcessor class implements page-level triage logic (lines 544–560) that analyzes visual complexity indicators including table density, image regions, and formula presence. Pages exceeding complexity thresholds route to the configured AI backend (Docling, Hancom, etc.), while simple text pages remain in the fast local path. The TriageLogger records these decisions to JSON for benchmark validation and recall calculation.

What happens if the AI backend fails during processing?

OpenDataLoader PDF implements a deterministic fallback mechanism accessible via the --hybrid-fallback flag or Config.builder().setHybridFallback(true). When enabled, pages that fail AI processing automatically route to the pure-Java extraction path, ensuring document completion even during backend outages. This fallback preserves system robustness while maintaining the speed benefits of hybrid triage for available pages.

Can I use custom AI backends with the hybrid benchmark suite?

Yes, the HybridConfig.java API supports arbitrary backend URLs via setHybridUrl(), allowing integration with internal AI services or alternative document analysis engines. To benchmark custom backends, modify the tests/benchmark/run.py configuration to point to your endpoint and ensure the output format matches the expected JSON schema containing bounding boxes, OCR text, and structural classifications. The thresholds.json file can be adjusted to establish appropriate baselines for your specific backend's performance characteristics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →