# How Hybrid Mode Dynamically Routes PDF Pages Between Local Java and AI Backends

> Discover how opendataloader-pdf's hybrid mode uses signals to dynamically route PDF pages between local Java and AI backends based on content complexity, optimizing your data processing.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: internals
- Published: 2026-03-20

---

**Hybrid mode in opendataloader-pdf uses a signal-based triage system to route each PDF page to either the fast local Java pipeline or an external AI backend based on content complexity.**

The opendataloader-pdf library implements a sophisticated hybrid processing architecture that dynamically routes individual PDF pages between a high-performance local Java engine and external AI backends like Docling or Hancom. This hybrid mode automatically analyzes page content to determine the optimal processing path, ensuring simple documents remain fast while complex layouts get the accuracy of AI parsing.

## The Three-Step Hybrid Routing Pipeline

Hybrid mode operates through a coordinated three-step pipeline that evaluates backend health, classifies page content, and executes parallel processing paths.

### Step 1: Backend Availability Verification

Before processing begins, the system verifies the selected backend is reachable. In `HybridDocumentProcessor.processDocument` (lines 9-13), the code calls `HybridClient.checkAvailability()`. If the backend is unreachable and the `--hybrid-fallback` flag is enabled, the system falls back to pure Java processing; otherwise, it fails fast with a connectivity error.

### Step 2: Page Triage and Signal Analysis

The core routing logic occurs in the triage phase. `TriageProcessor.triageAllPages` (lines 17-34 in `HybridDocumentProcessor`) iterates through each page, calling `classifyPage` to analyze content signals. The `SignalAccumulator` extracts structural clues including:

- **TableBorder** objects
- Vector graphics and grid lines
- Text pattern densities
- Large image footprints (≥11% of page area with aspect ratio ≥1.7)

### Step 3: Parallel Processing and Result Merging

After triage, pages split into two execution pools:

- **Java path**: Pages marked `JAVA` undergo local processing via `processJavaPath` (lines 48-96), running standard processors like `ClusterTableProcessor` and `TableBorderProcessor`.
- **Backend path**: Pages marked `BACKEND` are batched via `processBackendPath` (lines 127-176) and sent to `HybridClient.convert`, with JSON responses transformed by backend-specific `HybridSchemaTransformer` implementations like `DoclingSchemaTransformer`.

Finally, `mergeResults` (lines 191-206) reassembles pages in original order.

## The Triage Engine: Signal-Based Page Classification

The `TriageProcessor.classifyPage` method implements a priority-ordered decision tree (lines 31-98) that evaluates signals from most to least reliable:

1. **TableBorder detection** (`hasTableBorder`): Highest confidence trigger for BACKEND routing
2. **Vector graphics cues** (`hasVectorTableSignal`): Grid lines and line-art patterns
3. **Text pattern analysis** (`hasTextTablePattern`): Repeated table-like structures
4. **Large image detection** (`hasLargeImage`): Images exceeding 11% of page area with aspect ratio ≥1.7
5. **LineChunk ratio** (`getLineToTextRatio`): High ratio of line objects to text content

If no signals exceed their configured thresholds (defined in `HybridConfig`), the page defaults to `TriageResult.java(pageNumber, 0.9, signals)`.

## Configuration Modes: Auto vs Full

Hybrid mode supports two operational modes controlled via `HybridConfig`:

- **Auto mode** (`--hybrid-mode auto`): Default behavior activates the full signal-based triage system, routing only complex pages to AI backends while keeping simple pages in the fast Java pipeline.
- **Full mode** (`--hybrid-mode full`): Bypasses triage entirely. `HybridConfig.isFullMode()` returns true, forcing every page into the backend path via `HybridDocumentProcessor.processDocument` (lines 19-29).

## Fallback Handling for Backend Failures

When backend processing fails or returns `partial_success` status, the system implements transparent fallback. If `HybridConfig.isFallbackToJava()` returns true (enabled via `--hybrid-fallback`), failed pages from the backend batch are automatically re-routed through `processJavaPath` (lines 66-74), ensuring document completion even during AI service outages.

## Code Examples

### Command Line Interface

```bash

# Automatic triage routing (default)

opendataloader-pdf --hybrid docling sample.pdf -o out/

# Force all pages to AI backend

opendataloader-pdf --hybrid docling --hybrid-mode full sample.pdf -o out/

# Enable Java fallback on backend failure

opendataloader-pdf --hybrid docling --hybrid-fallback sample.pdf -o out/

```

### Programmatic Java API

```java
import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.processors.HybridDocumentProcessor;
import java.util.*;

String pdfPath = "complex-document.pdf";
Config cfg = new Config();
cfg.setHybrid("docling");                    // Select backend
cfg.setHybridMode("auto");                   // Or "full"
cfg.setHybridFallback(true);                 // Enable fallback

Set<Integer> pages = null;                   // Process all pages
List<List<IObject>> result = HybridDocumentProcessor.processDocument(
        pdfPath, cfg, pages);

```

### Debugging Triage Decisions

```java
import org.opendataloader.pdf.hybrid.TriageProcessor;
import org.opendataloader.pdf.hybrid.TriageResult;

Map<Integer, TriageResult> triage = TriageProcessor.triageAllPages(
        filteredContentsMap, cfg.getHybridConfig());

triage.forEach((page, result) -> {
    System.out.printf("Page %d routed to %s (confidence %.2f)%n",
            page + 1, result.getDecision(), result.getConfidence());
    // Access raw signals: result.getSignals()
});

```

## Key Implementation Files

The hybrid routing system is implemented across these core files:

- [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/HybridDocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/HybridDocumentProcessor.java) – Orchestrates the entire hybrid workflow including backend verification, triage, parallel execution, and result merging.
- [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java) – Configuration container for backend selection, operational mode (`auto`/`full`), and fallback flags.
- [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/TriageProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/TriageProcessor.java) – Signal extraction and priority-based routing decisions.
- [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridClientFactory.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridClientFactory.java) – Factory for creating backend-specific clients.
- [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/DoclingSchemaTransformer.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/DoclingSchemaTransformer.java) – Transforms AI backend JSON into internal `IObject` models.

## Summary

- **Hybrid mode** dynamically routes individual PDF pages between local Java processing and external AI backends using a signal-based triage system.
- The **triage engine** evaluates six categories of structural signals (TableBorder, vector graphics, text patterns, large images, LineChunk ratios, and alignment cues) in priority order to determine routing.
- **Three-step pipeline**: Backend availability verification → Signal-based page triage → Parallel Java/AI processing with automatic result merging.
- **Operational flexibility**: Choose `auto` mode for intelligent routing or `full` mode to force all pages to AI backends.
- **Resilience**: Optional fallback to Java processing when AI services fail, ensuring document completion via `processJavaPath`.

## Frequently Asked Questions

### How does the triage processor decide between Java and AI backend processing?

The `TriageProcessor.classifyPage` method evaluates signals in strict priority order: first checking for `TableBorder` objects (highest confidence), then vector graphics cues, text pattern densities, large images (≥11% of page area with aspect ratio ≥1.7), and finally LineChunk-to-text ratios. If any signal exceeds its configured threshold, the page routes to the AI backend; otherwise, it defaults to the Java path.

### What happens if the AI backend is unavailable during processing?

Before processing begins, `HybridDocumentProcessor.processDocument` calls `HybridClient.checkAvailability()` (lines 9-13). If the backend is unreachable and `--hybrid-fallback` is enabled, the system falls back to pure Java processing for all pages. Without fallback enabled, the operation fails immediately with a connectivity error.

### Can I force all pages to use the AI backend regardless of content complexity?

Yes. Set `--hybrid-mode full` via CLI or `cfg.setHybridMode("full")` programmatically. This sets `HybridConfig.isFullMode()` to true, which bypasses the `TriageProcessor` entirely and forces every page into the backend path via `HybridDocumentProcessor.processDocument` (lines 19-29).

### What types of PDF content trigger AI backend routing?

Complex structural elements trigger backend routing: tables with borders (`TableBorder` objects), vector graphics resembling grids or charts, dense text patterns resembling tables, large images occupying ≥11% of page area with aspect ratios ≥1.7, and pages with high LineChunk-to-text ratios indicating complex line-based layouts.

### How are results from Java and AI backends merged?

The `mergeResults` method (lines 191-206 in `HybridDocumentProcessor`) recombines processed pages from both the Java and backend paths, preserving the original page order. Each page retains its position index regardless of processing path, ensuring the final output maintains document sequence integrity.