# OpenDataLoader PDF Hybrid Mode Auto vs Full: Triage Differences Explained

> Understand OpenDataLoader PDF hybrid mode auto vs full triage differences. Learn how content analysis or full backend processing impacts your data loading efficiency.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: deep-dive
- Published: 2026-03-20

---

**In OpenDataLoader PDF, `auto` mode dynamically routes each page through the Java pipeline or AI backend based on content analysis, while `full` mode skips triage entirely and forces every page to the external backend.**

The OpenDataLoader PDF library provides a hybrid processing architecture that balances speed and accuracy by routing PDF pages between a fast Java-based engine and an external AI service. Understanding the difference between `hybrid_mode` settings is critical for optimizing document processing pipelines, latency, and API costs.

## What is Hybrid Mode?

Hybrid mode in OpenDataLoader PDF determines how the `HybridDocumentProcessor` routes pages between the local Java pipeline and an external AI backend. The mode is configured via `HybridConfig`, which stores the setting alongside backend connection parameters like URL, timeout, and fallback options.

## Auto Mode: Dynamic Per-Page Triage

`auto` is the default hybrid mode that enables intelligent page routing based on content complexity.

### How Auto Mode Works

When configured to `MODE_AUTO`, the processor invokes `TriageProcessor.triageAllPages()` to examine every page individually. As implemented in [`HybridDocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridDocumentProcessor.java) (lines 30-33), the system checks `!config.getHybridConfig().isFullMode()` to determine whether to run the triage step.

Each page receives a classification of either **JAVA** (process locally) or **BACKEND** (send to AI service). Only pages requiring advanced layout analysis or OCR capabilities are forwarded to the external service, while simple pages remain in the fast Java path.

### Java Configuration Example

```java
import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.hybrid.HybridConfig;
import org.opendataloader.pdf.processors.HybridDocumentProcessor;

// Configure auto mode
Config cfg = new Config();
HybridConfig hybrid = new HybridConfig();
hybrid.setMode(HybridConfig.MODE_AUTO); // Dynamic triage per page
cfg.setHybridConfig(hybrid);

// Process with intelligent routing
List<List<IObject>> result = HybridDocumentProcessor.processDocument(
    "document.pdf", cfg, null, null);

```

## Full Mode: Backend-Only Processing

`full` mode guarantees uniform processing by bypassing the triage logic entirely.

### How Full Mode Works

When `MODE_FULL` is active, `HybridDocumentProcessor` (lines 19-28) creates backend-only `TriageResult` entries without invoking the triage engine. The `isFullMode()` helper in [`HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridConfig.java) (lines 97-100) returns `true` when `mode == "full"`, triggering this bypass.

This forces every page into the **BACKEND** path regardless of content, ensuring the external AI handles layout analysis, OCR, and extraction for the entire document.

### Java Configuration Example

```java
Config cfg = new Config();
HybridConfig hybrid = new HybridConfig();

// Force all pages to backend
hybrid.setMode(HybridConfig.MODE_FULL);
hybrid.setFallbackToJava(true); // Optional safety net
cfg.setHybridConfig(hybrid);

List<List<IObject>> result = HybridDocumentProcessor.processDocument(
    "complex-layout.pdf", cfg, null, Paths.get("triage-logs"));

```

## Key Differences at a Glance

| Feature | Auto Mode (`auto`) | Full Mode (`full`) |
|---------|-------------------|-------------------|
| **Triage Execution** | Runs `TriageProcessor.triageAllPages()` on every page | Skips triage entirely |
| **Routing Logic** | Dynamic per-page decision (Java vs Backend) | All pages forced to backend |
| **Latency** | Lower (Java handles simple pages) | Higher (all pages via API) |
| **API Usage** | Optimized (only complex pages) | Maximum (every page) |
| **Use Case** | Mixed documents with simple and complex pages | Advanced layout/OCR requirements |
| **Default** | Yes | No |

## Source Code Implementation

The triage logic resides in two primary files:

- **[`HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridConfig.java)** (lines 45-48): Declares `MODE_AUTO` and `MODE_FULL` constants, plus the `isFullMode()` validation method
- **[`HybridDocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridDocumentProcessor.java)** (lines 19-33): Implements the branching logic that either runs triage or creates forced backend results based on the configuration

For CLI users, the mode is exposed via `--hybrid-mode`:

```bash

# Dynamic triage (default)

opendataloader-pdf --input report.pdf --hybrid-mode auto

# Full backend processing

opendataloader-pdf --input scan.pdf --hybrid-mode full

```

## Summary

- **`auto` mode** triggers per-page triage analysis to minimize backend API calls and latency while maintaining accuracy for complex pages
- **`full` mode** bypasses triage to ensure every page receives AI backend processing, ideal for documents requiring advanced OCR or layout analysis
- Configure via `HybridConfig.setMode()` using `HybridConfig.MODE_AUTO` or `HybridConfig.MODE_FULL` constants
- The decision logic branches in `HybridDocumentProcessor` based on the `isFullMode()` boolean check

## Frequently Asked Questions

### When should I use full mode instead of auto mode?

Use **full** mode when your documents contain complex layouts, handwritten text, or low-quality scans that require advanced AI capabilities beyond the Java engine. Use **auto** mode for standard documents with mixed content types where you want to minimize API costs and processing time.

### Does auto mode support fallback to Java if the backend fails?

Yes. Both modes support `setFallbackToJava(true)` in `HybridConfig`. In **auto** mode, pages already routed to Java continue processing locally, while backend-bound pages can fall back to Java if the AI service times out or returns errors.

### How does triage classify pages in auto mode?

The `TriageProcessor.triageAllPages()` method analyzes each page's content complexity, text density, and structural elements to assign a **JAVA** or **BACKEND** classification. Simple text-heavy pages typically stay local, while pages with tables, images, or complex layouts route to the backend.

### Can I change the hybrid mode per document or is it global?

The `HybridConfig` object is set per `Config` instance, allowing you to specify different modes for each `HybridDocumentProcessor.processDocument()` call. This enables pipeline flexibility where some documents use `auto` and others use `full` within the same application.