# How to Configure EasyOCR for Less Common Languages Like Arabic in PDF Extraction

> Extract Arabic text from PDFs using EasyOCR. Learn how to configure the opendataloader-pdf-hybrid server with the --ocr-lang flag for seamless PDF to text conversion.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: how-to-guide
- Published: 2026-03-20

---

**Start the `opendataloader-pdf-hybrid` server with the `--ocr-lang` flag set to EasyOCR language codes like `"ar"` for Arabic, and all PDF conversions will automatically use that language model.**

The `opendataloader-project/opendataloader-pdf` repository extracts text from PDFs using either standard embedded text reading or hybrid OCR-based processing. For scanned documents or image-based PDFs containing less common languages such as Arabic, you must configure the OCR engine at the server level rather than through traditional CLI arguments.

## How OCR Language Configuration Works in opendataloader-pdf

The codebase operates in two distinct extraction modes. **Standard mode** handles digital PDFs by reading embedded text directly through the Java layer without invoking OCR. **Hybrid mode** processes scanned or image-based PDFs through a lightweight FastAPI server that runs the Docling pipeline, which internally calls EasyOCR for text recognition.

In hybrid mode, the OCR language is no longer a Java runtime option. Instead, it is a **server-side configuration** passed to the Docling pipeline via `EasyOcrOptions`. The relevant implementation lives in [`python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py), where three key operations occur:

- **Argument parsing** (lines 54-58): The server defines the `--ocr-lang` flag and splits the input string into a Python list of EasyOCR language codes.
- **Converter construction** (lines 21-23): The `create_converter` function receives the language list as the `ocr_lang` parameter and assigns it to `EasyOcrOptions.lang`.
- **Pipeline wiring** (lines 34-38): The code builds `PdfPipelineOptions` with these OCR settings and instantiates a `DocumentConverter` that persists across HTTP requests.

EasyOCR ships with extensive language model support. To extract Arabic text, you simply pass the EasyOCR language code **`ar`** to the server. For other less common languages, consult the EasyOCR documentation for the appropriate two-letter or three-letter codes.

## Starting the Hybrid Server with Custom Languages

All OCR configuration has migrated to the hybrid server side to avoid JVM restarts and ensure model reuse across requests. Older CLI flags such as `--hybrid-ocr` remain in [`java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java) but are **deprecated**.

### CLI Configuration for Arabic OCR

Launch the hybrid server with your desired language codes before running any conversions:

```bash

# Start the backend (listens on http://0.0.0.0:5002)

opendataloader-pdf-hybrid \
    --port 5002 \
    --force-ocr \
    --ocr-lang "ar,en"

```

With the server running, invoke the standard CLI and specify hybrid mode:

```bash
opendataloader-pdf --hybrid docling-fast arabic-document.pdf

```

The `--ocr-lang` parameter accepts comma-separated EasyOCR codes. For multilingual documents, combine codes such as `"ja,ko,ch_sim"` for Japanese, Korean, and Simplified Chinese.

### Python Subprocess Automation

You can programmatically manage the server lifecycle from Python scripts:

```python
import subprocess
import time
import opendataloader_pdf

# Start the hybrid server in the background

server = subprocess.Popen([
    "opendataloader-pdf-hybrid",
    "--port", "5002",
    "--force-ocr",
    "--ocr-lang", "ar,en"
])

# Allow initialization time

time.sleep(2)

# Convert PDF using the Python wrapper

opendataloader_pdf.convert(
    input_path="arabic-document.pdf",
    output_dir="out/",
    hybrid="docling-fast"
)

# Clean up

server.terminate()

```

## Programmatic Configuration with create_converter

For advanced use cases where you embed the conversion logic directly within a Python service, instantiate the `DocumentConverter` manually using the `create_converter` factory function:

```python
from opendataloader_pdf.hybrid_server import create_converter

converter = create_converter(
    force_full_page_ocr=True,
    ocr_lang=["ar", "en"],
    enrich_formula=False,
    enrich_picture_description=False,
)

# Process PDFs directly

doc = converter.convert_file("arabic-document.pdf")
print(doc.json_content)

```

This approach bypasses the HTTP server layer and passes the `ocr_lang` list directly to `EasyOcrOptions`, then into `PdfPipelineOptions` as implemented in lines 34-38 of [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py).

## Limitations for Right-to-Left Scripts

When configuring EasyOCR for Arabic, Hebrew, or other right-to-left (RTL) scripts, be aware of the current reading-order limitation. The extraction algorithm processes visual coordinates only, meaning characters are recognized correctly but may appear in **visual order** rather than logical order in the JSON output. According to the user guide in `content/docs/hybrid-mode.mdx` (lines 96-100), the text flow follows the physical layout on the page rather than the logical reading sequence for RTL languages.

## Summary

- **Server-side configuration**: OCR languages are set via `--ocr-lang` on the `opendataloader-pdf-hybrid` server, not through deprecated Java CLI flags.
- **EasyOCR codes**: Use standard codes like `"ar"` for Arabic, passing multiple languages as comma-separated values.
- **Implementation location**: Language lists flow through `create_converter` in [`python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py) into `EasyOcrOptions.lang`.
- **RTL caveat**: Right-to-left scripts extract accurately but may require post-processing to restore logical reading order.
- **Integration**: Both CLI and Python wrappers automatically route OCR requests to the configured hybrid server once started.

## Frequently Asked Questions

### What EasyOCR language code should I use for Arabic PDF extraction?

Use the code **`ar`**. When starting the hybrid server, pass `--ocr-lang "ar"` for Arabic-only documents or `--ocr-lang "ar,en"` for mixed-language content. You can find the complete list of supported codes in the EasyOCR documentation or the `content/docs/hybrid-mode.mdx` file in the repository.

### Why does the extracted Arabic text appear in the wrong order?

The current implementation processes text based on **visual coordinates** on the page. For right-to-left scripts like Arabic, EasyOCR recognizes individual characters accurately, but the JSON output may list them in visual order (left-to-right as they appear spatially) rather than logical reading order. You may need to implement additional post-processing to reorder RTL text correctly.

### Can I no longer set OCR languages from the Java CLI?

Correct. The `--hybrid-ocr` and related flags in [`java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java) are **deprecated**. All OCR configuration—including language selection—now happens server-side via `opendataloader-pdf-hybrid` startup flags. This architecture ensures the OCR model loads once and serves multiple requests efficiently without JVM restarts.

### How do I extract text from a PDF containing multiple less common languages?

Pass a comma-separated list of EasyOCR codes to the `--ocr-lang` flag when starting the hybrid server. For example, use `--ocr-lang "ar,fa,ur"` to enable Arabic, Persian, and Urdu simultaneously. The Docling pipeline will attempt recognition using all specified language models for each page.