When Does `hybrid_fallback` Engage in OpenDataLoader PDF and How Does It Handle Errors?
The hybrid_fallback mechanism activates when external document processing backends throw exceptions or return partial success, automatically rerouting failed pages to the built-in Java PDF pipeline while logging detailed warnings to prevent complete conversion failures.
The hybrid_fallback feature in the opendataloader-project/opendataloader-pdf repository provides a critical safety net for hybrid document processing workflows. When enabled, it intercepts failures from external backends—such as docling, hancom, Azure, or Google Cloud—and seamlessly transitions processing to the internal Java pipeline, ensuring robust error recovery without manual intervention.
Activation Conditions for Hybrid Fallback
The HybridDocumentProcessor class implements distinct trigger conditions that determine when the Java fallback pipeline engages. These conditions are evaluated sequentially during the document conversion lifecycle.
Backend Exceptions and Complete Failures
When an external backend request throws an exception—such as a network timeout, authentication error, or invalid code point—the processor catches the failure in HybridDocumentProcessor.java (lines 166–174). The system then evaluates HybridConfig.isFallbackToJava():
- If enabled: The processor immediately reroutes all pages assigned to the backend through the Java processing path (
processJavaPath), discarding the failed backend results. - If disabled: The processor re-throws an
IOException, causing the entire conversion operation to fail fast.
This behavior ensures that transient network issues or backend outages do not automatically terminate batch processing jobs when resilience is configured.
Partial Success Responses
External backends may return a partial_success status when individual pages fail processing independently. In this scenario, the processor records the specific failed page numbers in the backendFailedPages collection (lines 177–185 of HybridDocumentProcessor.java):
- With fallback enabled: After the backend call returns, only the failed pages are re-processed through the Java pipeline, preserving successful backend results for other pages.
- With fallback disabled: The failed pages are skipped entirely, and the conversion continues with only the successfully processed pages.
Configuration Requirements and Defaults
By default, hybrid_fallback operates in a strict fail-fast mode. In HybridConfig.java (lines 43–44), the fallbackToJava boolean initializes to false, ensuring that any backend error aborts the conversion unless explicitly overridden.
To enable the safety net:
- Command Line: The
--hybrid-fallbackflag inCLIOptions.java(lines 34–35) setsHybridConfig.fallbackToJavatotrue. - Python Wrapper: The
convert_generated.pymodule (lines 35–36) propagates thehybrid_fallback=Trueparameter to the underlying Java CLI.
Error Handling and Logging Capabilities
The error handling architecture distinguishes between complete backend failures and partial degradation, providing granular observability through structured logging.
Exception Management Strategy
The HybridDocumentProcessor implements a two-tier exception handling strategy:
- Critical failures (total backend unavailability) trigger immediate fallback evaluation for the entire document.
- Page-level failures (partial success) trigger selective fallback only for affected pages, minimizing reprocessing overhead.
When fallback is disabled, the system maintains atomicity by throwing IOException upstream, allowing calling applications to implement their own retry logic or failure handling.
Logging Levels and Diagnostic Output
The processor emits distinct log entries via the LOGGER instance to facilitate debugging:
- Backend exceptions: Logged at
Level.WARNINGwith the full exception message and stack trace. - Partial success (fallback disabled): Logged at
Level.WARNINGindicating which pages were skipped and the reason for omission. - Partial success (fallback enabled): Logged at
Level.WARNINGfor the initial failure, followed byLevel.INFOwhen the Java fallback pipeline successfully completes reprocessing.
These logs are captured in the standard output and error streams used by the CLI, making them accessible in containerized environments and CI/CD pipelines.
Implementation Examples
Enabling Fallback via Command Line
opendataloader-pdf \
--input mydoc.pdf \
--hybrid docling \
--hybrid-fallback # Activates Java fallback on backend errors
The --hybrid-fallback flag sets HybridConfig.fallbackToJava to true as defined in CLIOptions.java.
Configuring Fallback in Python
from opendataloader_pdf import convert
convert(
input_path="mydoc.pdf",
hybrid="docling",
hybrid_fallback=True, # Enables Java fallback routing
)
The Python wrapper in convert_generated.py appends the --hybrid-fallback argument to the Java process invocation.
Programmatic Configuration in Java
HybridConfig cfg = new HybridConfig();
cfg.setFallbackToJava(true); // Enable fallback safety net
Config config = new Config();
config.setHybridConfig(cfg);
config.setHybrid("docling"); // Select external backend
List<List<IObject>> result = HybridDocumentProcessor.processDocument(
"mydoc.pdf", config, null);
Setting fallbackToJava directly on the configuration object triggers the same conditional logic that the CLI flag controls.
Test-Driven Validation
The fallback behavior is rigorously validated across both Java and Python test suites:
- Python integration: The
test_partial_success_no_page_range_fallbacktest intest_hybrid_server_partial_success.py(lines 71–81) verifies that fallback correctly reports failed pages even when no explicit page range constraints are supplied. - Java unit tests:
HybridDocumentProcessorTest.java(lines 74–75) asserts that fallback remains disabled by default and validates the state transition when explicitly enabled via configuration.
Summary
- The
hybrid_fallbackoption defaults to disabled (false) inHybridConfig.java, enforcing a fail-fast policy for backend errors. - It engages on two specific conditions: when a backend throws an unhandled exception (triggering full document reprocessing) or returns partial success (triggering selective page reprocessing).
- Error handling is configurable: Enabled fallback routes failures to the Java pipeline; disabled fallback propagates
IOExceptionto callers. - Comprehensive logging at
WARNINGandINFOlevels provides visibility into fallback triggers and reprocessing outcomes. - Configuration is available via CLI flag, Python wrapper parameter, or direct Java API manipulation.
Frequently Asked Questions
Is hybrid_fallback enabled by default in OpenDataLoader PDF?
No. According to HybridConfig.java (lines 43–44), the fallbackToJava field initializes to false. The system follows a fail-fast policy by default, meaning any backend exception or partial success will either throw an IOException or skip failed pages rather than attempting automatic recovery.
What happens to document pages when a backend fails and fallback is disabled?
If hybrid_fallback is disabled and the backend throws an exception, the HybridDocumentProcessor re-throws an IOException and the entire conversion fails. In cases of partial success where only specific pages fail, those pages are skipped entirely—logged at WARNING level—while successfully processed pages are retained in the output.
Does hybrid_fallback handle partial success differently than complete backend failure?
Yes. Complete backend failures trigger reprocessing of all pages through the Java pipeline, effectively discarding any partial results. Partial success scenarios trigger selective reprocessing: only pages listed in backendFailedPages are routed to the Java path, while successfully processed pages from the backend are preserved, optimizing performance and maintaining output quality.
Which external backends support the hybrid fallback mechanism?
The fallback mechanism is backend-agnostic and functions with any external processor configured through the hybrid system, including docling, hancom, Azure Document Intelligence, and Google Cloud Document AI. The HybridDocumentProcessor intercepts failures at the communication layer before backend-specific processing occurs, making the safety net universally applicable across all supported integrations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →