# How to Set the Language for OCR in MinerU: CLI, Python API, and HTTP Methods

> Easily set the OCR language in MinerU using CLI, Python API, or HTTP requests. Learn how to specify languages for accurate text recognition with Pytorch-Paddle-OCR.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: how-to-guide
- Published: 2026-02-23

---

**You can set the language for OCR in MinerU by passing the `lang` parameter via the CLI (`--lang`), Python API (`lang_list`), or HTTP API request body, which propagates through the pipeline to initialize the Pytorch-Paddle-OCR engine with language-specific weights.**

MinerU is an open-source document parsing toolkit developed by OpenDataLab that extracts text, formulas, and tables from PDFs and images. When processing non-English documents, you must explicitly set the language for OCR in MinerU to ensure accurate text recognition, as the underlying Pytorch-Paddle-OCR engine loads different neural network weights based on the specified language code.

## How the OCR Language Flows Through MinerU

Understanding how MinerU handles language settings helps debug recognition issues. The parameter propagates through three architectural layers before reaching the OCR engine.

### CLI and API Entry Points

In [`mineru/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/client.py) (lines 73-78), the CLI defines the `--lang` argument that accepts a language code and forwards it as a list. When using the Python API, the `doc_analyze` function in [`mineru/backend/pipeline/pipeline_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_analyze.py) (lines 70-78) accepts the `lang_list` parameter.

### Pipeline Model Initialization

The `MineruPipelineModel` class in [`mineru/backend/pipeline/model_init.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/model_init.py) (lines 48-52) stores the language in `self.lang` and passes it to the OCR atom model via `atom_model_manager.get_atom_model(..., lang=self.lang)`.

### OCR Engine Language Mapping

The actual language mapping occurs in [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py) (lines 45-64). This wrapper normalizes user-provided codes (e.g., `en`, `ch`) to internal Paddle-OCR language families and loads the appropriate detection and recognition weights.

## How to Set the OCR Language in MinerU

You can configure the OCR language through three interfaces depending on your deployment scenario.

### Command-Line Interface (CLI)

Use the `--lang` flag when running the `mineru` command. This is defined in [`mineru/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/client.py):

```bash
mineru \
  -p /path/to/document.pdf \
  -o /tmp/output \
  --method ocr \
  --lang en

```

The `--lang` flag accepts any code from the supported language table (e.g., `ch` for Simplified Chinese, `japan`, `korean`, `arabic`).

### Python API

When using the pipeline programmatically, pass the `lang_list` parameter to `doc_analyze` in [`mineru/backend/pipeline/pipeline_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_analyze.py):

```python
from mineru.backend.pipeline.pipeline_analyze import doc_analyze

# Load PDF into memory

with open("invoice.pdf", "rb") as f:
    pdf_bytes = f.read()

# Analyze with French OCR

results, _, _, _, _ = doc_analyze(
    pdf_bytes_list=[pdf_bytes],
    lang_list=["fr"],          # Language code for French

    parse_method="ocr",
    formula_enable=False,
    table_enable=False,
)

```

The `lang_list` parameter is forwarded to `MineruPipelineModel`, which initializes the OCR atom model with the specified language.

### HTTP API (FastAPI)

If running MinerU as a service, send a POST request with the `lang` field. The FastAPI endpoint in [`mineru/cli/fast_api.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/fast_api.py) extracts this value and builds the pipeline:

```python
import requests

files = {"file": open("handwritten.png", "rb")}
data = {
    "method": "ocr",
    "lang": "korean",           # Set OCR language to Korean

    "formula_enable": "false",
    "table_enable": "false"
}

response = requests.post("http://localhost:8000/analyze", files=files, data=data)
print(response.json())

```

## Supported OCR Languages and Auto-Fallback Behavior

MinerU supports multiple languages through the Pytorch-Paddle-OCR backend. The complete list of language codes is defined in [`projects/mcp/src/mineru/language.py`](https://github.com/opendatalab/MinerU/blob/main/projects/mcp/src/mineru/language.py) (lines 5-20), which includes common options like:

- `en` – English
- `ch` – Simplified Chinese
- `ch_server` – Server-optimized Chinese
- `japan` – Japanese
- `korean` – Korean
- `arabic` – Arabic
- `fr` – French
- `german` – German

**CPU Optimization Fallback:** In [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py) (lines 48-52), MinerU implements an automatic fallback for CPU devices. If you specify a heavyweight Chinese family language (`ch`, `ch_server`, `japan`, or `chinese_cht`) but the device is `cpu`, the system automatically switches to the lighter `ch_lite` model to improve processing speed without requiring manual configuration.

## Summary

- **Set the language for OCR in MinerU** using the `lang` parameter across CLI (`--lang`), Python API (`lang_list`), or HTTP API (`lang` field).
- The parameter propagates through [`mineru/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/client.py), [`mineru/backend/pipeline/model_init.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/model_init.py), and finally to [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py) where it maps to Paddle-OCR language families.
- Supported languages include `en`, `ch`, `japan`, `korean`, `arabic`, and others defined in [`projects/mcp/src/mineru/language.py`](https://github.com/opendatalab/MinerU/blob/main/projects/mcp/src/mineru/language.py).
- When running on CPU with heavyweight Chinese models, MinerU automatically falls back to `ch_lite` for better performance.

## Frequently Asked Questions

### What language codes does MinerU support for OCR?

MinerU supports all language codes defined in [`projects/mcp/src/mineru/language.py`](https://github.com/opendatalab/MinerU/blob/main/projects/mcp/src/mineru/language.py), including `en` (English), `ch` (Simplified Chinese), `japan` (Japanese), `korean` (Korean), `arabic` (Arabic), `fr` (French), and `german` (German). These codes map to specific Paddle-OCR model weights in [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py).

### How do I set the OCR language when using the MinerU Python API?

Pass the `lang_list` parameter to the `doc_analyze` function in [`mineru/backend/pipeline/pipeline_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_analyze.py). For example: `doc_analyze(pdf_bytes_list=[pdf_bytes], lang_list=["en"], parse_method="ocr")`. This list is forwarded to `MineruPipelineModel`, which initializes the OCR atom model with the specified language code.

### Does MinerU automatically optimize OCR models for CPU devices?

Yes. According to the implementation in [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py) (lines 48-52), when the device is `cpu` and you specify a heavyweight Chinese language family (`ch`, `ch_server`, `japan`, or `chinese_cht`), MinerU automatically falls back to the lighter `ch_lite` model to improve processing speed without requiring manual configuration.

### Can I specify multiple languages for OCR in a single MinerU request?

The parameter structure in [`mineru/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/client.py) and [`mineru/backend/pipeline/pipeline_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_analyze.py) accepts lists of languages (`--lang` or `lang_list`), suggesting the API supports multiple language codes. However, the underlying Paddle-OCR engine typically processes one primary language per model instance. You should verify specific multi-language behavior in your deployment, as the automatic fallback logic in [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py) primarily handles single-language selection based on the first specified code or device constraints.