How Hybrid Mode Dynamically Routes PDF Pages Between Local Java and AI Backends
Hybrid mode in opendataloader-pdf uses a signal-based triage system to route each PDF page to either the fast local Java pipeline or an external AI backend based on content complexity.
The opendataloader-pdf library implements a sophisticated hybrid processing architecture that dynamically routes individual PDF pages between a high-performance local Java engine and external AI backends like Docling or Hancom. This hybrid mode automatically analyzes page content to determine the optimal processing path, ensuring simple documents remain fast while complex layouts get the accuracy of AI parsing.
The Three-Step Hybrid Routing Pipeline
Hybrid mode operates through a coordinated three-step pipeline that evaluates backend health, classifies page content, and executes parallel processing paths.
Step 1: Backend Availability Verification
Before processing begins, the system verifies the selected backend is reachable. In HybridDocumentProcessor.processDocument (lines 9-13), the code calls HybridClient.checkAvailability(). If the backend is unreachable and the --hybrid-fallback flag is enabled, the system falls back to pure Java processing; otherwise, it fails fast with a connectivity error.
Step 2: Page Triage and Signal Analysis
The core routing logic occurs in the triage phase. TriageProcessor.triageAllPages (lines 17-34 in HybridDocumentProcessor) iterates through each page, calling classifyPage to analyze content signals. The SignalAccumulator extracts structural clues including:
- TableBorder objects
- Vector graphics and grid lines
- Text pattern densities
- Large image footprints (≥11% of page area with aspect ratio ≥1.7)
Step 3: Parallel Processing and Result Merging
After triage, pages split into two execution pools:
- Java path: Pages marked
JAVAundergo local processing viaprocessJavaPath(lines 48-96), running standard processors likeClusterTableProcessorandTableBorderProcessor. - Backend path: Pages marked
BACKENDare batched viaprocessBackendPath(lines 127-176) and sent toHybridClient.convert, with JSON responses transformed by backend-specificHybridSchemaTransformerimplementations likeDoclingSchemaTransformer.
Finally, mergeResults (lines 191-206) reassembles pages in original order.
The Triage Engine: Signal-Based Page Classification
The TriageProcessor.classifyPage method implements a priority-ordered decision tree (lines 31-98) that evaluates signals from most to least reliable:
- TableBorder detection (
hasTableBorder): Highest confidence trigger for BACKEND routing - Vector graphics cues (
hasVectorTableSignal): Grid lines and line-art patterns - Text pattern analysis (
hasTextTablePattern): Repeated table-like structures - Large image detection (
hasLargeImage): Images exceeding 11% of page area with aspect ratio ≥1.7 - LineChunk ratio (
getLineToTextRatio): High ratio of line objects to text content
If no signals exceed their configured thresholds (defined in HybridConfig), the page defaults to TriageResult.java(pageNumber, 0.9, signals).
Configuration Modes: Auto vs Full
Hybrid mode supports two operational modes controlled via HybridConfig:
- Auto mode (
--hybrid-mode auto): Default behavior activates the full signal-based triage system, routing only complex pages to AI backends while keeping simple pages in the fast Java pipeline. - Full mode (
--hybrid-mode full): Bypasses triage entirely.HybridConfig.isFullMode()returns true, forcing every page into the backend path viaHybridDocumentProcessor.processDocument(lines 19-29).
Fallback Handling for Backend Failures
When backend processing fails or returns partial_success status, the system implements transparent fallback. If HybridConfig.isFallbackToJava() returns true (enabled via --hybrid-fallback), failed pages from the backend batch are automatically re-routed through processJavaPath (lines 66-74), ensuring document completion even during AI service outages.
Code Examples
Command Line Interface
# Automatic triage routing (default)
opendataloader-pdf --hybrid docling sample.pdf -o out/
# Force all pages to AI backend
opendataloader-pdf --hybrid docling --hybrid-mode full sample.pdf -o out/
# Enable Java fallback on backend failure
opendataloader-pdf --hybrid docling --hybrid-fallback sample.pdf -o out/
Programmatic Java API
import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.processors.HybridDocumentProcessor;
import java.util.*;
String pdfPath = "complex-document.pdf";
Config cfg = new Config();
cfg.setHybrid("docling"); // Select backend
cfg.setHybridMode("auto"); // Or "full"
cfg.setHybridFallback(true); // Enable fallback
Set<Integer> pages = null; // Process all pages
List<List<IObject>> result = HybridDocumentProcessor.processDocument(
pdfPath, cfg, pages);
Debugging Triage Decisions
import org.opendataloader.pdf.hybrid.TriageProcessor;
import org.opendataloader.pdf.hybrid.TriageResult;
Map<Integer, TriageResult> triage = TriageProcessor.triageAllPages(
filteredContentsMap, cfg.getHybridConfig());
triage.forEach((page, result) -> {
System.out.printf("Page %d routed to %s (confidence %.2f)%n",
page + 1, result.getDecision(), result.getConfidence());
// Access raw signals: result.getSignals()
});
Key Implementation Files
The hybrid routing system is implemented across these core files:
java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/HybridDocumentProcessor.java– Orchestrates the entire hybrid workflow including backend verification, triage, parallel execution, and result merging.java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java– Configuration container for backend selection, operational mode (auto/full), and fallback flags.java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/TriageProcessor.java– Signal extraction and priority-based routing decisions.java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridClientFactory.java– Factory for creating backend-specific clients.java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/DoclingSchemaTransformer.java– Transforms AI backend JSON into internalIObjectmodels.
Summary
- Hybrid mode dynamically routes individual PDF pages between local Java processing and external AI backends using a signal-based triage system.
- The triage engine evaluates six categories of structural signals (TableBorder, vector graphics, text patterns, large images, LineChunk ratios, and alignment cues) in priority order to determine routing.
- Three-step pipeline: Backend availability verification → Signal-based page triage → Parallel Java/AI processing with automatic result merging.
- Operational flexibility: Choose
automode for intelligent routing orfullmode to force all pages to AI backends. - Resilience: Optional fallback to Java processing when AI services fail, ensuring document completion via
processJavaPath.
Frequently Asked Questions
How does the triage processor decide between Java and AI backend processing?
The TriageProcessor.classifyPage method evaluates signals in strict priority order: first checking for TableBorder objects (highest confidence), then vector graphics cues, text pattern densities, large images (≥11% of page area with aspect ratio ≥1.7), and finally LineChunk-to-text ratios. If any signal exceeds its configured threshold, the page routes to the AI backend; otherwise, it defaults to the Java path.
What happens if the AI backend is unavailable during processing?
Before processing begins, HybridDocumentProcessor.processDocument calls HybridClient.checkAvailability() (lines 9-13). If the backend is unreachable and --hybrid-fallback is enabled, the system falls back to pure Java processing for all pages. Without fallback enabled, the operation fails immediately with a connectivity error.
Can I force all pages to use the AI backend regardless of content complexity?
Yes. Set --hybrid-mode full via CLI or cfg.setHybridMode("full") programmatically. This sets HybridConfig.isFullMode() to true, which bypasses the TriageProcessor entirely and forces every page into the backend path via HybridDocumentProcessor.processDocument (lines 19-29).
What types of PDF content trigger AI backend routing?
Complex structural elements trigger backend routing: tables with borders (TableBorder objects), vector graphics resembling grids or charts, dense text patterns resembling tables, large images occupying ≥11% of page area with aspect ratios ≥1.7, and pages with high LineChunk-to-text ratios indicating complex line-based layouts.
How are results from Java and AI backends merged?
The mergeResults method (lines 191-206 in HybridDocumentProcessor) recombines processed pages from both the Java and backend paths, preserving the original page order. Each page retains its position index regardless of processing path, ensuring the final output maintains document sequence integrity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →