How to Enable or Disable OCR in MinerU: Complete Configuration Guide
You can enable or disable OCR in MinerU using the --method flag in the CLI (ocr, txt, or auto), the enable_ocr parameter in the HTTP API, or the is_ocr checkbox in the Gradio UI.
MinerU is an open-source document parsing toolkit that supports multiple input methods for controlling optical character recognition (OCR). Whether processing scanned PDFs or digitally generated documents, you can explicitly force OCR on, force it off, or let the system automatically detect text layers. The configuration flows through three distinct architectural layers: the user-facing interface, the pipeline core, and the OCR execution engine.
Understanding OCR Control Architecture
MinerU implements OCR toggling across three layers that propagate the configuration from user input to execution:
- Interface Layer: Accepts
--method(CLI),enable_ocr(API), oris_ocr(UI) inputs - Pipeline Layer: Translates methods into Boolean
_ocr_enableflags inpipeline_analyze.py - Engine Layer: Conditionally executes OCR in
ocr_utils.pybased on the final Boolean value
The pipeline stores per-document OCR decisions in ocr_enabled_list and passes the final ocr_enable Boolean to mineru/utils/ocr_utils.py::get_ocr_result_list, which skips OCR recognition when the value is False.
Method 1: Command-Line Interface Control
The CLI provides the most direct method to enable or disable OCR using the --method argument defined in mineru/cli/client.py (lines 45-55).
Force OCR for All Documents
Use --method ocr to force OCR processing on every page, regardless of whether extractable text layers exist:
mineru --method ocr -p /path/to/file.pdf -o ./output
This setting bypasses automatic detection and invokes the OCR engine for all content.
Disable OCR Completely
Use --method txt to treat every PDF as text-only, disabling OCR entirely:
mineru --method txt -p ./documents -o ./output
This mode extracts only embedded text streams and ignores image-based content.
Automatic Detection (Default)
Use --method auto to let MinerU decide per-document based on content analysis:
mineru --method auto -p ./documents -o ./output
In auto mode, the pipeline calls pdf_classify.classify() from mineru/utils/pdf_classify.py to determine if the document requires OCR.
Method 2: HTTP API and Python Client
For programmatic access, the HTTP API exposes OCR control through the enable_ocr parameter in projects/mcp/src/mineru/api.py (lines 95-100).
Global OCR Toggle
Set the enable_ocr argument when submitting batch tasks:
from mineru.api import MinerUClient
client = MinerUClient(base_url="http://localhost:30000")
response = await client.submit_file_url_task(
urls=["https://example.com/report.pdf"],
enable_ocr=True, # Forces OCR for all URLs in this request
language="ch",
)
Per-URL Control
You can override the global setting for individual documents using the is_ocr field:
response = await client.submit_file_url_task(
urls=[
{"url": "https://example.com/scan.pdf", "is_ocr": True},
{"url": "https://example.com/digital.pdf", "is_ocr": False},
],
enable_ocr=False, # Default fallback
)
The is_ocr value is injected into each URL configuration and consumed downstream by the pipeline logic in pipeline_analyze.py.
Method 3: Gradio Web Interface
The Gradio frontend in mineru/cli/gradio_app.py exposes a "Force enable OCR" checkbox that maps to the is_ocr parameter:
is_ocr = gr.Checkbox(
label=i18n("force_ocr"),
value=False,
info=i18n("force_ocr_info")
)
When checked, the checkbox value passes to the backend via:
await client.call("convert_file_url", url=urls, enable_ocr=is_ocr)
Internal Pipeline Logic
When processing documents, mineru/backend/pipeline/pipeline_analyze.py evaluates the OCR configuration through the following flow:
- Method Translation: The
parse_methodvalue (auto,txt, orocr) is evaluated per PDF - Classification: For
automode,pdf_classify.classify(pdf_bytes)returns"ocr"or"txt" - Boolean Assignment: The pipeline sets
_ocr_enable = Truefor"ocr"andFalsefor"txt" - List Storage: Results are stored in
ocr_enabled_list(line 99 ofpipeline_analyze.py) - Execution Gate:
mineru/utils/ocr_utils.py::get_ocr_result_listreads the finalocr_enableBoolean and only populates OCR results whenTrue
Summary
- CLI: Use
--method ocrto enable,--method txtto disable, or--method autofor automatic detection viamineru/cli/client.py - API: Pass
enable_ocr=True/Falseglobally oris_ocr=True/Falseper URL throughprojects/mcp/src/mineru/api.py - UI: Toggle the Force enable OCR checkbox in the Gradio interface defined in
mineru/cli/gradio_app.py - Auto Logic: The pipeline uses
pdf_classify.classify()inmineru/utils/pdf_classify.pyto determine OCR necessity when method is set toauto - Execution: Final OCR skipping occurs in
mineru/utils/ocr_utils.py::get_ocr_result_listbased on the propagated Boolean flag
Frequently Asked Questions
What is the default OCR behavior in MinerU?
By default, MinerU uses --method auto, which analyzes each PDF with pdf_classify.classify() to determine if the document contains scanned images requiring OCR or selectable text that can be extracted directly. The classifier returns "ocr" for image-heavy documents and "txt" for digitally generated PDFs.
Can I enable OCR for some pages and disable for others in the same PDF?
Currently, MinerU determines OCR at the document level rather than the page level. The _ocr_enable Boolean applies to the entire PDF as stored in ocr_enabled_list. While the pipeline processes individual pages, the OCR decision is uniform across the document based on the initial classification or forced method.
How does MinerU decide whether to use OCR in auto mode?
In auto mode, mineru/backend/pipeline/pipeline_analyze.py invokes pdf_classify.classify(pdf_bytes) from mineru/utils/pdf_classify.py. This function analyzes the PDF structure and content to classify it as requiring OCR ("ocr") or not ("txt"). The result determines the _ocr_enable flag that controls whether ocr_utils.py executes recognition.
Is there a performance difference between forced OCR and text extraction?
Yes. Forced OCR (--method ocr) invokes the full recognition pipeline in mineru/utils/ocr_utils.py::get_ocr_result_list, which processes images through the OCR engine and is computationally expensive. Text extraction (--method txt) bypasses OCR entirely and extracts embedded text streams directly, significantly reducing processing time for digitally generated documents.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →