How to Configure EasyOCR for Less Common Languages Like Arabic in PDF Extraction
Start the opendataloader-pdf-hybrid server with the --ocr-lang flag set to EasyOCR language codes like "ar" for Arabic, and all PDF conversions will automatically use that language model.
The opendataloader-project/opendataloader-pdf repository extracts text from PDFs using either standard embedded text reading or hybrid OCR-based processing. For scanned documents or image-based PDFs containing less common languages such as Arabic, you must configure the OCR engine at the server level rather than through traditional CLI arguments.
How OCR Language Configuration Works in opendataloader-pdf
The codebase operates in two distinct extraction modes. Standard mode handles digital PDFs by reading embedded text directly through the Java layer without invoking OCR. Hybrid mode processes scanned or image-based PDFs through a lightweight FastAPI server that runs the Docling pipeline, which internally calls EasyOCR for text recognition.
In hybrid mode, the OCR language is no longer a Java runtime option. Instead, it is a server-side configuration passed to the Docling pipeline via EasyOcrOptions. The relevant implementation lives in python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py, where three key operations occur:
- Argument parsing (lines 54-58): The server defines the
--ocr-langflag and splits the input string into a Python list of EasyOCR language codes. - Converter construction (lines 21-23): The
create_converterfunction receives the language list as theocr_langparameter and assigns it toEasyOcrOptions.lang. - Pipeline wiring (lines 34-38): The code builds
PdfPipelineOptionswith these OCR settings and instantiates aDocumentConverterthat persists across HTTP requests.
EasyOCR ships with extensive language model support. To extract Arabic text, you simply pass the EasyOCR language code ar to the server. For other less common languages, consult the EasyOCR documentation for the appropriate two-letter or three-letter codes.
Starting the Hybrid Server with Custom Languages
All OCR configuration has migrated to the hybrid server side to avoid JVM restarts and ensure model reuse across requests. Older CLI flags such as --hybrid-ocr remain in java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java but are deprecated.
CLI Configuration for Arabic OCR
Launch the hybrid server with your desired language codes before running any conversions:
# Start the backend (listens on http://0.0.0.0:5002)
opendataloader-pdf-hybrid \
--port 5002 \
--force-ocr \
--ocr-lang "ar,en"
With the server running, invoke the standard CLI and specify hybrid mode:
opendataloader-pdf --hybrid docling-fast arabic-document.pdf
The --ocr-lang parameter accepts comma-separated EasyOCR codes. For multilingual documents, combine codes such as "ja,ko,ch_sim" for Japanese, Korean, and Simplified Chinese.
Python Subprocess Automation
You can programmatically manage the server lifecycle from Python scripts:
import subprocess
import time
import opendataloader_pdf
# Start the hybrid server in the background
server = subprocess.Popen([
"opendataloader-pdf-hybrid",
"--port", "5002",
"--force-ocr",
"--ocr-lang", "ar,en"
])
# Allow initialization time
time.sleep(2)
# Convert PDF using the Python wrapper
opendataloader_pdf.convert(
input_path="arabic-document.pdf",
output_dir="out/",
hybrid="docling-fast"
)
# Clean up
server.terminate()
Programmatic Configuration with create_converter
For advanced use cases where you embed the conversion logic directly within a Python service, instantiate the DocumentConverter manually using the create_converter factory function:
from opendataloader_pdf.hybrid_server import create_converter
converter = create_converter(
force_full_page_ocr=True,
ocr_lang=["ar", "en"],
enrich_formula=False,
enrich_picture_description=False,
)
# Process PDFs directly
doc = converter.convert_file("arabic-document.pdf")
print(doc.json_content)
This approach bypasses the HTTP server layer and passes the ocr_lang list directly to EasyOcrOptions, then into PdfPipelineOptions as implemented in lines 34-38 of hybrid_server.py.
Limitations for Right-to-Left Scripts
When configuring EasyOCR for Arabic, Hebrew, or other right-to-left (RTL) scripts, be aware of the current reading-order limitation. The extraction algorithm processes visual coordinates only, meaning characters are recognized correctly but may appear in visual order rather than logical order in the JSON output. According to the user guide in content/docs/hybrid-mode.mdx (lines 96-100), the text flow follows the physical layout on the page rather than the logical reading sequence for RTL languages.
Summary
- Server-side configuration: OCR languages are set via
--ocr-langon theopendataloader-pdf-hybridserver, not through deprecated Java CLI flags. - EasyOCR codes: Use standard codes like
"ar"for Arabic, passing multiple languages as comma-separated values. - Implementation location: Language lists flow through
create_converterinpython/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.pyintoEasyOcrOptions.lang. - RTL caveat: Right-to-left scripts extract accurately but may require post-processing to restore logical reading order.
- Integration: Both CLI and Python wrappers automatically route OCR requests to the configured hybrid server once started.
Frequently Asked Questions
What EasyOCR language code should I use for Arabic PDF extraction?
Use the code ar. When starting the hybrid server, pass --ocr-lang "ar" for Arabic-only documents or --ocr-lang "ar,en" for mixed-language content. You can find the complete list of supported codes in the EasyOCR documentation or the content/docs/hybrid-mode.mdx file in the repository.
Why does the extracted Arabic text appear in the wrong order?
The current implementation processes text based on visual coordinates on the page. For right-to-left scripts like Arabic, EasyOCR recognizes individual characters accurately, but the JSON output may list them in visual order (left-to-right as they appear spatially) rather than logical reading order. You may need to implement additional post-processing to reorder RTL text correctly.
Can I no longer set OCR languages from the Java CLI?
Correct. The --hybrid-ocr and related flags in java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java are deprecated. All OCR configuration—including language selection—now happens server-side via opendataloader-pdf-hybrid startup flags. This architecture ensures the OCR model loads once and serves multiple requests efficiently without JVM restarts.
How do I extract text from a PDF containing multiple less common languages?
Pass a comma-separated list of EasyOCR codes to the --ocr-lang flag when starting the hybrid server. For example, use --ocr-lang "ar,fa,ur" to enable Arabic, Persian, and Urdu simultaneously. The Docling pipeline will attempt recognition using all specified language models for each page.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →