When Does `hybrid_fallback` Engage in OpenDataLoader PDF and How Does It Handle Errors?

The hybrid_fallback mechanism activates when external document processing backends throw exceptions or return partial success, automatically rerouting failed pages to the built-in Java PDF pipeline while logging detailed warnings to prevent complete conversion failures.

The hybrid_fallback feature in the opendataloader-project/opendataloader-pdf repository provides a critical safety net for hybrid document processing workflows. When enabled, it intercepts failures from external backends—such as docling, hancom, Azure, or Google Cloud—and seamlessly transitions processing to the internal Java pipeline, ensuring robust error recovery without manual intervention.

Activation Conditions for Hybrid Fallback

The HybridDocumentProcessor class implements distinct trigger conditions that determine when the Java fallback pipeline engages. These conditions are evaluated sequentially during the document conversion lifecycle.

Backend Exceptions and Complete Failures

When an external backend request throws an exception—such as a network timeout, authentication error, or invalid code point—the processor catches the failure in HybridDocumentProcessor.java (lines 166–174). The system then evaluates HybridConfig.isFallbackToJava():

  • If enabled: The processor immediately reroutes all pages assigned to the backend through the Java processing path (processJavaPath), discarding the failed backend results.
  • If disabled: The processor re-throws an IOException, causing the entire conversion operation to fail fast.

This behavior ensures that transient network issues or backend outages do not automatically terminate batch processing jobs when resilience is configured.

Partial Success Responses

External backends may return a partial_success status when individual pages fail processing independently. In this scenario, the processor records the specific failed page numbers in the backendFailedPages collection (lines 177–185 of HybridDocumentProcessor.java):

  • With fallback enabled: After the backend call returns, only the failed pages are re-processed through the Java pipeline, preserving successful backend results for other pages.
  • With fallback disabled: The failed pages are skipped entirely, and the conversion continues with only the successfully processed pages.

Configuration Requirements and Defaults

By default, hybrid_fallback operates in a strict fail-fast mode. In HybridConfig.java (lines 43–44), the fallbackToJava boolean initializes to false, ensuring that any backend error aborts the conversion unless explicitly overridden.

To enable the safety net:

  • Command Line: The --hybrid-fallback flag in CLIOptions.java (lines 34–35) sets HybridConfig.fallbackToJava to true.
  • Python Wrapper: The convert_generated.py module (lines 35–36) propagates the hybrid_fallback=True parameter to the underlying Java CLI.

Error Handling and Logging Capabilities

The error handling architecture distinguishes between complete backend failures and partial degradation, providing granular observability through structured logging.

Exception Management Strategy

The HybridDocumentProcessor implements a two-tier exception handling strategy:

  1. Critical failures (total backend unavailability) trigger immediate fallback evaluation for the entire document.
  2. Page-level failures (partial success) trigger selective fallback only for affected pages, minimizing reprocessing overhead.

When fallback is disabled, the system maintains atomicity by throwing IOException upstream, allowing calling applications to implement their own retry logic or failure handling.

Logging Levels and Diagnostic Output

The processor emits distinct log entries via the LOGGER instance to facilitate debugging:

  • Backend exceptions: Logged at Level.WARNING with the full exception message and stack trace.
  • Partial success (fallback disabled): Logged at Level.WARNING indicating which pages were skipped and the reason for omission.
  • Partial success (fallback enabled): Logged at Level.WARNING for the initial failure, followed by Level.INFO when the Java fallback pipeline successfully completes reprocessing.

These logs are captured in the standard output and error streams used by the CLI, making them accessible in containerized environments and CI/CD pipelines.

Implementation Examples

Enabling Fallback via Command Line

opendataloader-pdf \
  --input mydoc.pdf \
  --hybrid docling \
  --hybrid-fallback          # Activates Java fallback on backend errors

The --hybrid-fallback flag sets HybridConfig.fallbackToJava to true as defined in CLIOptions.java.

Configuring Fallback in Python

from opendataloader_pdf import convert

convert(
    input_path="mydoc.pdf",
    hybrid="docling",
    hybrid_fallback=True,          # Enables Java fallback routing

)

The Python wrapper in convert_generated.py appends the --hybrid-fallback argument to the Java process invocation.

Programmatic Configuration in Java

HybridConfig cfg = new HybridConfig();
cfg.setFallbackToJava(true);          // Enable fallback safety net

Config config = new Config();
config.setHybridConfig(cfg);
config.setHybrid("docling");          // Select external backend

List<List<IObject>> result = HybridDocumentProcessor.processDocument(
        "mydoc.pdf", config, null);

Setting fallbackToJava directly on the configuration object triggers the same conditional logic that the CLI flag controls.

Test-Driven Validation

The fallback behavior is rigorously validated across both Java and Python test suites:

  • Python integration: The test_partial_success_no_page_range_fallback test in test_hybrid_server_partial_success.py (lines 71–81) verifies that fallback correctly reports failed pages even when no explicit page range constraints are supplied.
  • Java unit tests: HybridDocumentProcessorTest.java (lines 74–75) asserts that fallback remains disabled by default and validates the state transition when explicitly enabled via configuration.

Summary

  • The hybrid_fallback option defaults to disabled (false) in HybridConfig.java, enforcing a fail-fast policy for backend errors.
  • It engages on two specific conditions: when a backend throws an unhandled exception (triggering full document reprocessing) or returns partial success (triggering selective page reprocessing).
  • Error handling is configurable: Enabled fallback routes failures to the Java pipeline; disabled fallback propagates IOException to callers.
  • Comprehensive logging at WARNING and INFO levels provides visibility into fallback triggers and reprocessing outcomes.
  • Configuration is available via CLI flag, Python wrapper parameter, or direct Java API manipulation.

Frequently Asked Questions

Is hybrid_fallback enabled by default in OpenDataLoader PDF?

No. According to HybridConfig.java (lines 43–44), the fallbackToJava field initializes to false. The system follows a fail-fast policy by default, meaning any backend exception or partial success will either throw an IOException or skip failed pages rather than attempting automatic recovery.

What happens to document pages when a backend fails and fallback is disabled?

If hybrid_fallback is disabled and the backend throws an exception, the HybridDocumentProcessor re-throws an IOException and the entire conversion fails. In cases of partial success where only specific pages fail, those pages are skipped entirely—logged at WARNING level—while successfully processed pages are retained in the output.

Does hybrid_fallback handle partial success differently than complete backend failure?

Yes. Complete backend failures trigger reprocessing of all pages through the Java pipeline, effectively discarding any partial results. Partial success scenarios trigger selective reprocessing: only pages listed in backendFailedPages are routed to the Java path, while successfully processed pages from the backend are preserved, optimizing performance and maintaining output quality.

Which external backends support the hybrid fallback mechanism?

The fallback mechanism is backend-agnostic and functions with any external processor configured through the hybrid system, including docling, hancom, Azure Document Intelligence, and Google Cloud Document AI. The HybridDocumentProcessor intercepts failures at the communication layer before backend-specific processing occurs, making the safety net universally applicable across all supported integrations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →