# Limitations of the Local Java-Only Processing Mode in OpenDataLoader PDF

> Explore the limitations of OpenDataLoader PDF's local Java-only processing mode. Discover what it lacks in AI features like OCR and GPU acceleration while achieving fast native PDF parsing.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: performance
- Published: 2026-03-20

---

**The local Java-only processing mode in OpenDataLoader PDF delivers high-speed CPU-based parsing at approximately 0.05 seconds per page, but lacks AI-driven features including OCR, formula extraction, and complex table detection, while supporting only native PDF documents without GPU acceleration.**

The `opendataloader-project/opendataloader-pdf` repository provides a hybrid parsing architecture that defaults to pure Java execution. Understanding the **limitations of the local Java-only processing mode** is essential for architects deciding between offline speed and AI-enhanced accuracy. This mode processes documents entirely within the JVM without external network calls or GPU acceleration.

## What Is the Local Java-Only Processing Mode?

By default, OpenDataLoader PDF operates in Java-only mode when the `--hybrid` flag is omitted or explicitly set to `off`. This configuration is defined in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java):

```java
public static final String HYBRID_OFF = "off";               // default: Java‑only
private String hybrid = HYBRID_OFF;                         // current mode

```

In this mode, the pipeline consists solely of the **Java Path** as documented in [`docs/hybrid/hybrid-mode-design.md`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/docs/hybrid/hybrid-mode-design.md). After the `ContentFilterProcessor` and `TriageProcessor`, every page proceeds directly to local Java processors such as `TableBorderProcessor`, `TextLineProcessor`, and `ParagraphProcessor`. No batch API calls, schema transformers, or backend merges are involved.

## Key Limitations of the Local Java-Only Processing Mode

### No AI-Driven Enrichments

The Java-only mode cannot perform AI-driven enrichments such as **OCR**, **formula extraction**, **picture description**, or **complex borderless-table detection**. These capabilities require enabling a hybrid backend such as `docling`. When running with `--hybrid off`, these features are unavailable because the pipeline lacks the neural network components necessary for computer vision tasks.

### Restricted Document Type Support

Only **native PDF content** is processed. The Java-only mode cannot handle Microsoft Word, Excel, or PowerPoint files. According to the [`README.md`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/README.md) limitations section, attempting to process these file types in Java-only mode will fail or produce empty results because the local engine lacks the format-specific parsers required for Office documents.

### CPU-Only Execution

The engine runs **purely on CPU** without GPU acceleration. While this eliminates GPU hardware requirements, it also means that any processor-intensive tasks that could benefit from parallel GPU computation—such as image analysis or deep learning inference—are bound to single-threaded or multi-threaded CPU performance only.

### Limited Complex Layout Handling

Simple heuristics handle standard tables and headings, but **complex table structures**, **scanned images**, and **embedded formulas** often fail to parse correctly. The local Java processors rely on geometric analysis and rule-based extraction rather than machine learning models that can interpret visual context. This results in lower accuracy for documents with multi-column layouts, merged cells, or mixed content types.

### No External Network Capabilities

By design, the Java-only mode makes **no external network calls**. While this ensures complete offline operation and data privacy, it also prevents the system from leveraging external AI services for tasks such as cloud-based OCR, image recognition, or content enrichment via third-party APIs.

### Silently Ignored Extension Flags

Adding new AI-based processors via flags such as `--enrich-formula` or `--enrich-picture-description` requires hybrid mode. When `--hybrid off` is set, these flags are **silently ignored** and the system continues with standard Java processing without warning the user that enrichment features are inactive.

## Performance Characteristics

The primary advantage of accepting these limitations is **speed**. The local Java-only mode achieves approximately **0.05 seconds per page**, making it significantly faster than hybrid modes that require HTTP requests to external services or containerized backends. This performance profile suits high-volume batch processing of clean, text-based PDFs where AI enrichment is unnecessary.

## Architecture Overview

When operating in Java-only mode, the pipeline follows a strictly local path. Documents flow through `ContentFilterProcessor` and `TriageProcessor` before entering the core Java extraction layer. The `TableBorderProcessor` handles geometric table detection, `TextLineProcessor` manages text extraction, and `ParagraphProcessor` structures content blocks. As documented in [`docs/hybrid/hybrid-mode-design.md`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/docs/hybrid/hybrid-mode-design.md), this path bypasses entirely the schema transformers and batch API components required for AI processing.

## How to Overcome Limitations by Enabling Hybrid Mode

To access OCR, formula extraction, and Office file support, switch from Java-only mode to a hybrid backend:

```bash

# Use docling backend for OCR, complex tables, formulas, etc.

opendataloader-pdf --hybrid docling --hybrid-url http://localhost:5001 input.pdf

```

This command activates the hybrid pipeline, routing pages through external AI services that provide the enrichments unavailable in pure Java mode.

## Summary

- **The local Java-only processing mode** is the default configuration in OpenDataLoader PDF, activated when `--hybrid` is omitted or set to `off` in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java).
- **Functional constraints** include the absence of OCR, formula extraction, picture description, and complex table detection, which require AI backends.
- **Document restrictions** limit processing to native PDFs only, excluding Microsoft Office formats and scanned image documents.
- **Hardware and network isolation** means pure CPU execution with no GPU acceleration and zero external network calls.
- **Performance trade-off** yields approximately 0.05 seconds per page processing speed at the cost of advanced extraction capabilities.

## Frequently Asked Questions

### Can the Java-only mode perform OCR on scanned PDFs?

No, the Java-only mode cannot perform OCR. Optical character recognition requires computer vision models that are only available when using a hybrid backend such as `docling`. In Java-only mode, scanned images are processed using simple heuristics that cannot extract text from images.

### Why does the Java-only mode ignore the `--enrich-formula` flag?

The flags `--enrich-formula` and `--enrich-picture-description` are designed to activate AI-based processors that require neural network inference. When `--hybrid off` is set, these flags are silently ignored because the local Java pipeline lacks the infrastructure to execute deep learning models. To use these features, enable hybrid mode with `--hybrid docling`.

### Does the Java-only mode support Microsoft Word or Excel files?

No, the Java-only mode only supports native PDF content. According to the repository's [`README.md`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/README.md), Word, Excel, and PowerPoint files are not supported in this mode. Processing these formats requires the hybrid backend which contains the necessary format-specific parsers and converters.

### Is GPU acceleration available in Java-only mode?

No, GPU acceleration is not available. The Java-only engine runs purely on CPU by design. While this eliminates GPU hardware requirements and ensures consistent behavior across environments, it also means that computationally intensive tasks cannot benefit from GPU parallelization.