What AI Models Are Utilized for Complex Page Processing in Hybrid Mode
Hybrid mode in OpenDataLoader PDF routes complex pages containing tables, OCR-heavy scans, formulas, and images to specialized AI models including EasyOCR for text recognition, TableFormer for table structure extraction, SmolVLM-256M-Instruct for picture description, and an internal Docling pipeline for LaTeX formula extraction.
The opendataloader-project/opendataloader-pdf repository implements a hybrid architecture that combines fast local Java parsing with AI-driven services for document pages requiring advanced processing. When the local parser encounters tables, scanned text, mathematical formulas, or embedded pictures, it delegates these complex elements to the Docling-Fast backend, which orchestrates several specialized machine learning models to extract structured content accurately.
Hybrid Mode Architecture Overview
Hybrid mode operates as a two-tier system where the Java frontend handles simple text extraction locally while delegating complex content to an external AI backend.
Page Routing Logic
The Java client evaluates each PDF page to determine processing complexity. Pages containing only standard text stream through the local parser, while pages with tables, scanned images, formulas, or charts are flagged for backend processing. The HybridClient.java interface defines this contract, with DoclingFastServerClient.java implementing the HTTP transport layer to communicate with the Python-based Docling-Fast server.
Backend Configuration
The HybridConfig class manages supported backend options including docling-fast, hancom, azure, and google. As implemented in HybridConfig.java at lines 141-150, only the docling-fast backend is fully operational; other providers exist as placeholders for future extensions. The configuration maps logical backend names to concrete server URLs, enabling the Java client to route requests to the appropriate AI processing endpoint.
AI Models for Complex Page Processing
The Docling-Fast backend integrates four specialized AI components, each configured in python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py.
EasyOCR for Text Recognition
EasyOCR handles optical character recognition for scanned documents and image-based text. Configured via EasyOcrOptions at lines 20-23 of hybrid_server.py, the system loads language-specific models (such as en, ko, or ar) to extract text from rasterized content. This component activates when the --force-ocr flag is enabled or when the system detects image-based text requiring transcription.
TableFormer for Table Structure
TableFormer extracts semantic table structures from visual table representations. The backend initializes this model through TableStructureOptions(mode=TableFormerMode.ACCURATE) at lines 38-39 of hybrid_server.py. Running in accurate mode, TableFormer analyzes cell boundaries, row/column relationships, and header hierarchies to convert visual tables into structured data formats compatible with the final JSON, Markdown, or HTML output.
Formula Extraction Pipeline
Mathematical formulas undergo enrichment through an internal Docling pipeline component triggered by the do_formula_enrichment=True flag at lines 39-40 of hybrid_server.py. When enabled, this pipeline processes LaTeX expressions and mathematical notation embedded in PDFs, converting visual formula representations into machine-readable LaTeX markup that preserves the semantic structure of equations.
SmolVLM for Picture Description
SmolVLM-256M-Instruct, a lightweight vision-language model, generates textual descriptions of images, charts, and diagrams. Configured through PictureDescriptionVlmOptions at lines 28-31 of hybrid_server.py, this model loads from the HuggingFace repository HuggingFaceTB/SmolVLM-256M-Instruct to produce captions for visual elements. This capability requires explicit activation via the --enrich-picture-description CLI flag or equivalent API parameter.
Configuration and Implementation Details
The system exposes these AI capabilities through both command-line interfaces and programmatic APIs.
Python Backend Setup
The FastAPI server defined in hybrid_server.py wires all four AI components together. Server initialization accepts parameters controlling which models load into memory:
# Launch the Docling Fast server with all AI enrichments enabled
opendataloader-pdf-hybrid \
--port 5002 \
--force-ocr \
--enrich-formula \
--enrich-picture-description
Java Client Integration
The Java side packages requests through HybridRequest, which bundles PDF bytes and requested output formats (JSON, MARKDOWN, HTML) as defined in HybridClient.java at lines 70-78. The client sends these payloads to the configured backend URL, receiving enriched content that the Java processor merges with locally-extracted text.
Practical Usage Examples
CLI Processing
Process a document requiring full AI backend processing:
# Client requesting full backend mode for picture descriptions
opendataloader-pdf \
--hybrid docling-fast \
--hybrid-mode full \
file-with-charts.pdf
Python API Integration
Programmatically convert documents with all AI enrichments:
import opendataloader_pdf
# Convert with OCR, formula, and picture description enabled
result = opendataloader_pdf.convert(
input_path="file-with-charts.pdf",
output_dir="out/",
hybrid="docling-fast",
hybrid_mode="full",
hybrid_url="http://localhost:5002"
)
# Access generated picture descriptions from JSON output
for element in result["elements"]:
if element["type"] == "picture":
print("Caption:", element["description"])
Summary
- EasyOCR powers text extraction from scanned documents, configured in
hybrid_server.pylines 20-23 with multilingual language model support. - TableFormer operates in accurate mode for semantic table structure extraction, defined at lines 38-39 of the backend server.
- Internal Docling pipeline handles LaTeX formula enrichment when
do_formula_enrichment=Trueis set at lines 39-40. - SmolVLM-256M-Instruct generates picture and chart descriptions via
PictureDescriptionVlmOptionsat lines 28-31. - HybridConfig.java (lines 141-150) manages backend selection, currently supporting only the docling-fast implementation.
- Full backend mode requires explicit CLI flags (
--hybrid-mode full) or API parameters to activate picture description capabilities.
Frequently Asked Questions
What is hybrid mode in OpenDataLoader PDF?
Hybrid mode is a dual-phase processing architecture where simple text extraction occurs locally in Java, while complex content including tables, OCR-heavy pages, formulas, and pictures routes to the Docling-Fast Python backend. This approach balances processing speed with AI-powered accuracy for challenging document elements.
Which AI model handles table extraction?
TableFormer manages table structure extraction, operating in accurate mode via TableStructureOptions(mode=TableFormerMode.ACCURATE) as configured in hybrid_server.py. This model analyzes visual table layouts to determine cell relationships and hierarchical structures.
How do I enable picture description generation?
Picture description requires starting the backend with --enrich-picture-description and requesting full hybrid mode via --hybrid-mode full on the client side. The SmolVLM-256M-Instruct model (256M parameter vision-language model from HuggingFace) processes these requests according to configuration at lines 28-31 of hybrid_server.py.
What file configures the backend server URL?
HybridConfig.java defines backend URL mappings at lines 141-150, supporting logical names like docling-fast, hancom, azure, and google. Currently, only docling-fast is fully implemented, with the configuration resolving logical backend identifiers to concrete HTTP endpoints that DoclingFastServerClient.java uses for transport.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →