How to Enable or Disable OCR in MinerU: Complete Configuration Guide

You can enable or disable OCR in MinerU using the --method flag in the CLI (ocr, txt, or auto), the enable_ocr parameter in the HTTP API, or the is_ocr checkbox in the Gradio UI.

MinerU is an open-source document parsing toolkit that supports multiple input methods for controlling optical character recognition (OCR). Whether processing scanned PDFs or digitally generated documents, you can explicitly force OCR on, force it off, or let the system automatically detect text layers. The configuration flows through three distinct architectural layers: the user-facing interface, the pipeline core, and the OCR execution engine.

Understanding OCR Control Architecture

MinerU implements OCR toggling across three layers that propagate the configuration from user input to execution:

  • Interface Layer: Accepts --method (CLI), enable_ocr (API), or is_ocr (UI) inputs
  • Pipeline Layer: Translates methods into Boolean _ocr_enable flags in pipeline_analyze.py
  • Engine Layer: Conditionally executes OCR in ocr_utils.py based on the final Boolean value

The pipeline stores per-document OCR decisions in ocr_enabled_list and passes the final ocr_enable Boolean to mineru/utils/ocr_utils.py::get_ocr_result_list, which skips OCR recognition when the value is False.

Method 1: Command-Line Interface Control

The CLI provides the most direct method to enable or disable OCR using the --method argument defined in mineru/cli/client.py (lines 45-55).

Force OCR for All Documents

Use --method ocr to force OCR processing on every page, regardless of whether extractable text layers exist:

mineru --method ocr -p /path/to/file.pdf -o ./output

This setting bypasses automatic detection and invokes the OCR engine for all content.

Disable OCR Completely

Use --method txt to treat every PDF as text-only, disabling OCR entirely:

mineru --method txt -p ./documents -o ./output

This mode extracts only embedded text streams and ignores image-based content.

Automatic Detection (Default)

Use --method auto to let MinerU decide per-document based on content analysis:

mineru --method auto -p ./documents -o ./output

In auto mode, the pipeline calls pdf_classify.classify() from mineru/utils/pdf_classify.py to determine if the document requires OCR.

Method 2: HTTP API and Python Client

For programmatic access, the HTTP API exposes OCR control through the enable_ocr parameter in projects/mcp/src/mineru/api.py (lines 95-100).

Global OCR Toggle

Set the enable_ocr argument when submitting batch tasks:

from mineru.api import MinerUClient

client = MinerUClient(base_url="http://localhost:30000")
response = await client.submit_file_url_task(
    urls=["https://example.com/report.pdf"],
    enable_ocr=True,  # Forces OCR for all URLs in this request

    language="ch",
)

Per-URL Control

You can override the global setting for individual documents using the is_ocr field:

response = await client.submit_file_url_task(
    urls=[
        {"url": "https://example.com/scan.pdf", "is_ocr": True},
        {"url": "https://example.com/digital.pdf", "is_ocr": False},
    ],
    enable_ocr=False,  # Default fallback

)

The is_ocr value is injected into each URL configuration and consumed downstream by the pipeline logic in pipeline_analyze.py.

Method 3: Gradio Web Interface

The Gradio frontend in mineru/cli/gradio_app.py exposes a "Force enable OCR" checkbox that maps to the is_ocr parameter:

is_ocr = gr.Checkbox(
    label=i18n("force_ocr"), 
    value=False, 
    info=i18n("force_ocr_info")
)

When checked, the checkbox value passes to the backend via:

await client.call("convert_file_url", url=urls, enable_ocr=is_ocr)

Internal Pipeline Logic

When processing documents, mineru/backend/pipeline/pipeline_analyze.py evaluates the OCR configuration through the following flow:

  1. Method Translation: The parse_method value (auto, txt, or ocr) is evaluated per PDF
  2. Classification: For auto mode, pdf_classify.classify(pdf_bytes) returns "ocr" or "txt"
  3. Boolean Assignment: The pipeline sets _ocr_enable = True for "ocr" and False for "txt"
  4. List Storage: Results are stored in ocr_enabled_list (line 99 of pipeline_analyze.py)
  5. Execution Gate: mineru/utils/ocr_utils.py::get_ocr_result_list reads the final ocr_enable Boolean and only populates OCR results when True

Summary

  • CLI: Use --method ocr to enable, --method txt to disable, or --method auto for automatic detection via mineru/cli/client.py
  • API: Pass enable_ocr=True/False globally or is_ocr=True/False per URL through projects/mcp/src/mineru/api.py
  • UI: Toggle the Force enable OCR checkbox in the Gradio interface defined in mineru/cli/gradio_app.py
  • Auto Logic: The pipeline uses pdf_classify.classify() in mineru/utils/pdf_classify.py to determine OCR necessity when method is set to auto
  • Execution: Final OCR skipping occurs in mineru/utils/ocr_utils.py::get_ocr_result_list based on the propagated Boolean flag

Frequently Asked Questions

What is the default OCR behavior in MinerU?

By default, MinerU uses --method auto, which analyzes each PDF with pdf_classify.classify() to determine if the document contains scanned images requiring OCR or selectable text that can be extracted directly. The classifier returns "ocr" for image-heavy documents and "txt" for digitally generated PDFs.

Can I enable OCR for some pages and disable for others in the same PDF?

Currently, MinerU determines OCR at the document level rather than the page level. The _ocr_enable Boolean applies to the entire PDF as stored in ocr_enabled_list. While the pipeline processes individual pages, the OCR decision is uniform across the document based on the initial classification or forced method.

How does MinerU decide whether to use OCR in auto mode?

In auto mode, mineru/backend/pipeline/pipeline_analyze.py invokes pdf_classify.classify(pdf_bytes) from mineru/utils/pdf_classify.py. This function analyzes the PDF structure and content to classify it as requiring OCR ("ocr") or not ("txt"). The result determines the _ocr_enable flag that controls whether ocr_utils.py executes recognition.

Is there a performance difference between forced OCR and text extraction?

Yes. Forced OCR (--method ocr) invokes the full recognition pipeline in mineru/utils/ocr_utils.py::get_ocr_result_list, which processes images through the OCR engine and is computationally expensive. Text extraction (--method txt) bypasses OCR entirely and extracts embedded text streams directly, significantly reducing processing time for digitally generated documents.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →