How to Specify the Parsing Method (Auto, TXT, OCR) in MinerU
You can specify the parsing method in MinerU using the -m or --method flag in the CLI, the parse_method parameter in the Python API, or the parse_method form field in the FastAPI endpoint to choose between auto, txt, or ocr modes.
MinerU is an open-source document parsing tool developed by OpenDataLab that extracts structured data from PDFs and images. When processing documents, you can control whether the tool uses text extraction, OCR, or automatic detection by specifying the parsing method parameter.
Understanding the Three Parsing Methods
MinerU supports three distinct parsing strategies that determine how content is extracted from PDF documents.
Auto Mode (Default)
The auto method analyzes the PDF at runtime to determine the optimal extraction strategy. If the document contains extractable text layers, MinerU uses the text-extraction path; if the PDF is image-based or lacks text layers, it automatically falls back to OCR. This mode is implemented in mineru/utils/pdf_classify.py, where the classify function inspects PDF bytes to determine if OCR is required.
TXT Mode (Text Extraction)
The txt method forces MinerU to use the text-extraction pipeline exclusively, even for image-only PDFs. This mode extracts selectable text directly from the PDF without invoking OCR models, making it faster for documents that already contain text layers but potentially returning empty results for scanned documents.
OCR Mode (Optical Character Recognition)
The ocr method forces the OCR pipeline regardless of whether the PDF contains extractable text. This is useful when you need to process scanned documents or when the existing text layer in a PDF is corrupted or incomplete. In this mode, MinerU runs layout detection first, then crops text blocks and sends them to the OCR model for recognition.
How to Specify the Parsing Method in MinerU
You can specify the parsing method through three different interfaces depending on your integration needs.
Command Line Interface (CLI)
When using the mineru command, pass the -m or --method flag followed by your chosen method. The argument parsing is defined in mineru/cli/client.py at lines 44-48, where the value is passed to the do_parse function.
# Use auto-detection (default)
mineru -p document.pdf -o ./output
# Force text extraction
mineru -p document.pdf -o ./output -m txt
# Force OCR for scanned documents
mineru -p scanned.pdf -o ./output -m ocr
FastAPI HTTP Endpoint
When using the FastAPI server, include the parse_method form field in your POST request to the /file_parse endpoint. This parameter is declared in mineru/cli/fast_api.py at lines 63-70 and forwarded to the aio_do_parse function.
curl -X POST http://localhost:8000/file_parse \
-F "files=@/path/to/document.pdf" \
-F "output_dir=./output" \
-F "parse_method=ocr" \
-F "backend=hybrid-auto-engine"
Python API
For programmatic usage, pass the parse_method argument directly to the do_parse or aio_do_parse functions imported from mineru.cli.common. This value is then forwarded to the backend analyzers.
from mineru.cli.common import do_parse
from mineru.utils.cli_parser import arg_parse
# Force text extraction mode
do_parse(
output_dir="./output",
pdf_file_names=["document"],
pdf_bytes_list=[pdf_bytes],
p_lang_list=["ch"],
backend="pipeline",
parse_method="txt", # Options: "auto", "txt", "ocr"
formula_enable=True,
table_enable=True,
server_url=None,
start_page_id=0,
end_page_id=None,
**arg_parse(None)
)
How the Parsing Method Works Internally
When you specify a parsing method, the value flows through several layers of the MinerU architecture before determining the actual processing path.
The parse_method parameter first reaches the backend analyzer—either mineru/backend/pipeline/pipeline_analyze.py or mineru/backend/hybrid/hybrid_analyze.py. Both modules contain logic to determine whether OCR should be enabled based on your specification and the document content.
In mineru/backend/hybrid/hybrid_analyze.py (lines 33-41), the ocr_classify helper function evaluates the method:
- If
parse_method == "auto", it calls the classifier inmineru/utils/pdf_classify.pyto inspect the PDF bytes and determine if the document is image-only. - If
parse_method == "txt", it forces_ocr_enabletoFalse, bypassing OCR regardless of content. - If
parse_method == "ocr", it forces_ocr_enabletoTrue, enabling OCR for all pages.
This boolean flag then determines whether the pipeline extracts text directly from PDF text layers or crops text blocks for OCR processing. Both paths generate a middle-JSON structure that is later converted to your desired output format (Markdown, JSON, etc.).
Summary
- Three methods available:
auto(default),txt(force text extraction), andocr(force OCR). - CLI usage: Use
mineru -m txtormineru --method ocrwhen running the command. - API usage: Pass
parse_method="ocr"todo_parse()or include it as a form field in FastAPI requests. - Internal logic: The
ocr_classifyfunction inmineru/backend/hybrid/hybrid_analyze.pyevaluates your choice against the PDF content to set the_ocr_enableflag. - Performance implications:
txtis fastest for text-based PDFs,ocris necessary for scanned documents, andautoprovides the best balance by detecting document types at runtime.
Frequently Asked Questions
What is the default parsing method in MinerU?
The default parsing method is auto, which automatically detects whether a PDF contains extractable text layers or requires OCR. When set to auto, MinerU uses the classify function in mineru/utils/pdf_classify.py to inspect the document and choose the appropriate processing path at runtime.
When should I use the TXT parsing method instead of AUTO?
Use the txt method when you know your PDFs contain clean, selectable text layers and you want to maximize processing speed by skipping the automatic detection step. This forces MinerU to use the text-extraction pipeline exclusively, bypassing OCR even if the classifier would normally recommend it. However, if the PDF is actually image-based, this method will return empty results.
How does MinerU decide between text extraction and OCR in AUTO mode?
In auto mode, MinerU calls the ocr_classify helper function (located in mineru/backend/hybrid/hybrid_analyze.py), which invokes the classify function from mineru/utils/pdf_classify.py. This classifier analyzes the PDF bytes to determine if the document is image-only. If the classifier returns "ocr", the system sets _ocr_enable to True and processes the document with OCR; otherwise, it extracts text directly from the PDF layers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →