How to Set the Language for OCR in MinerU: CLI, Python API, and HTTP Methods
You can set the language for OCR in MinerU by passing the lang parameter via the CLI (--lang), Python API (lang_list), or HTTP API request body, which propagates through the pipeline to initialize the Pytorch-Paddle-OCR engine with language-specific weights.
MinerU is an open-source document parsing toolkit developed by OpenDataLab that extracts text, formulas, and tables from PDFs and images. When processing non-English documents, you must explicitly set the language for OCR in MinerU to ensure accurate text recognition, as the underlying Pytorch-Paddle-OCR engine loads different neural network weights based on the specified language code.
How the OCR Language Flows Through MinerU
Understanding how MinerU handles language settings helps debug recognition issues. The parameter propagates through three architectural layers before reaching the OCR engine.
CLI and API Entry Points
In mineru/cli/client.py (lines 73-78), the CLI defines the --lang argument that accepts a language code and forwards it as a list. When using the Python API, the doc_analyze function in mineru/backend/pipeline/pipeline_analyze.py (lines 70-78) accepts the lang_list parameter.
Pipeline Model Initialization
The MineruPipelineModel class in mineru/backend/pipeline/model_init.py (lines 48-52) stores the language in self.lang and passes it to the OCR atom model via atom_model_manager.get_atom_model(..., lang=self.lang).
OCR Engine Language Mapping
The actual language mapping occurs in mineru/model/ocr/pytorch_paddle.py (lines 45-64). This wrapper normalizes user-provided codes (e.g., en, ch) to internal Paddle-OCR language families and loads the appropriate detection and recognition weights.
How to Set the OCR Language in MinerU
You can configure the OCR language through three interfaces depending on your deployment scenario.
Command-Line Interface (CLI)
Use the --lang flag when running the mineru command. This is defined in mineru/cli/client.py:
mineru \
-p /path/to/document.pdf \
-o /tmp/output \
--method ocr \
--lang en
The --lang flag accepts any code from the supported language table (e.g., ch for Simplified Chinese, japan, korean, arabic).
Python API
When using the pipeline programmatically, pass the lang_list parameter to doc_analyze in mineru/backend/pipeline/pipeline_analyze.py:
from mineru.backend.pipeline.pipeline_analyze import doc_analyze
# Load PDF into memory
with open("invoice.pdf", "rb") as f:
pdf_bytes = f.read()
# Analyze with French OCR
results, _, _, _, _ = doc_analyze(
pdf_bytes_list=[pdf_bytes],
lang_list=["fr"], # Language code for French
parse_method="ocr",
formula_enable=False,
table_enable=False,
)
The lang_list parameter is forwarded to MineruPipelineModel, which initializes the OCR atom model with the specified language.
HTTP API (FastAPI)
If running MinerU as a service, send a POST request with the lang field. The FastAPI endpoint in mineru/cli/fast_api.py extracts this value and builds the pipeline:
import requests
files = {"file": open("handwritten.png", "rb")}
data = {
"method": "ocr",
"lang": "korean", # Set OCR language to Korean
"formula_enable": "false",
"table_enable": "false"
}
response = requests.post("http://localhost:8000/analyze", files=files, data=data)
print(response.json())
Supported OCR Languages and Auto-Fallback Behavior
MinerU supports multiple languages through the Pytorch-Paddle-OCR backend. The complete list of language codes is defined in projects/mcp/src/mineru/language.py (lines 5-20), which includes common options like:
en– Englishch– Simplified Chinesech_server– Server-optimized Chinesejapan– Japanesekorean– Koreanarabic– Arabicfr– Frenchgerman– German
CPU Optimization Fallback: In mineru/model/ocr/pytorch_paddle.py (lines 48-52), MinerU implements an automatic fallback for CPU devices. If you specify a heavyweight Chinese family language (ch, ch_server, japan, or chinese_cht) but the device is cpu, the system automatically switches to the lighter ch_lite model to improve processing speed without requiring manual configuration.
Summary
- Set the language for OCR in MinerU using the
langparameter across CLI (--lang), Python API (lang_list), or HTTP API (langfield). - The parameter propagates through
mineru/cli/client.py,mineru/backend/pipeline/model_init.py, and finally tomineru/model/ocr/pytorch_paddle.pywhere it maps to Paddle-OCR language families. - Supported languages include
en,ch,japan,korean,arabic, and others defined inprojects/mcp/src/mineru/language.py. - When running on CPU with heavyweight Chinese models, MinerU automatically falls back to
ch_litefor better performance.
Frequently Asked Questions
What language codes does MinerU support for OCR?
MinerU supports all language codes defined in projects/mcp/src/mineru/language.py, including en (English), ch (Simplified Chinese), japan (Japanese), korean (Korean), arabic (Arabic), fr (French), and german (German). These codes map to specific Paddle-OCR model weights in mineru/model/ocr/pytorch_paddle.py.
How do I set the OCR language when using the MinerU Python API?
Pass the lang_list parameter to the doc_analyze function in mineru/backend/pipeline/pipeline_analyze.py. For example: doc_analyze(pdf_bytes_list=[pdf_bytes], lang_list=["en"], parse_method="ocr"). This list is forwarded to MineruPipelineModel, which initializes the OCR atom model with the specified language code.
Does MinerU automatically optimize OCR models for CPU devices?
Yes. According to the implementation in mineru/model/ocr/pytorch_paddle.py (lines 48-52), when the device is cpu and you specify a heavyweight Chinese language family (ch, ch_server, japan, or chinese_cht), MinerU automatically falls back to the lighter ch_lite model to improve processing speed without requiring manual configuration.
Can I specify multiple languages for OCR in a single MinerU request?
The parameter structure in mineru/cli/client.py and mineru/backend/pipeline/pipeline_analyze.py accepts lists of languages (--lang or lang_list), suggesting the API supports multiple language codes. However, the underlying Paddle-OCR engine typically processes one primary language per model instance. You should verify specific multi-language behavior in your deployment, as the automatic fallback logic in mineru/model/ocr/pytorch_paddle.py primarily handles single-language selection based on the first specified code or device constraints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →