How to Use MinerU's Python API for Programmatic PDF Parsing
MinerU exposes synchronous do_parse and asynchronous aio_do_parse functions in mineru/cli/common.py that let you parse PDFs programmatically without touching the command line.
The opendatalab/MinerU repository provides a high-level Python API for extracting structured Markdown, JSON, and images from PDF documents. Whether you need local OCR processing, remote Vision-Language Model (VLM) inference, or a hybrid approach, you can orchestrate the entire workflow through Python function calls rather than shell commands.
Core API Components
The public API surface is intentionally narrow, consisting of two primary entry points and several helper utilities located in mineru/cli/common.py.
do_parse – The synchronous entry point for local pipeline processing. It handles backend selection, PDF normalization, and output generation in a blocking call.
aio_do_parse – The asynchronous counterpart designed for I/O-bound backends like hybrid-* or vlm-http-client that communicate with remote servers.
read_fn – A utility function that detects file extensions and automatically converts image inputs (PNG, JPG) into PDF byte streams before processing.
Parsing Workflow
When you invoke do_parse or aio_do_parse, MinerU executes a five-stage pipeline internally. Understanding these stages helps you configure parameters correctly.
Reading and Normalizing Input
First, read_fn (lines 32-43 in mineru/cli/common.py) ingests the source file. If you provide an image, it converts the bytes to a single-page PDF using utilities in mineru/utils/pdf_image_tools.py.
Next, _prepare_pdf_bytes (lines 54-81) handles page range slicing. If you specify start_page_id and end_page_id, this function uses pypdfium2 via convert_pdf_bytes_to_bytes_by_pypdfium2 to trim the PDF before analysis.
Selecting the Backend
The backend parameter determines which analysis engine runs. Around lines 39-45 in mineru/cli/common.py, the code branches based on your selection:
pipeline– Default local processing using OCR, formula, and table detection models.vlm-*– Routes tovlm_doc_analyzeinmineru/backend/vlm/vlm_analyze.pyfor remote VLM processing.hybrid-*– Combines local OCR with remote VLM viaaio_hybrid_doc_analyzeinmineru/backend/hybrid/hybrid_analyze.py.
Processing and Output Generation
For the pipeline backend, _process_pipeline calls pipeline_doc_analyze (imported from mineru/backend/pipeline/pipeline_analyze.py). This returns three key structures: infer_results (model predictions), all_image_lists (rendered page images), and all_pdf_docs (PDF objects).
Finally, _process_output (lines 94-168 in mineru/cli/common.py) writes the results to disk. It generates Markdown via pipeline_union_make or vlm_union_make, content-list JSON, middle-JSON, and optional visualization PDFs (layout, span, line-sort).
Code Examples
Synchronous Parsing with the Pipeline Backend
The following example demonstrates the most common use case: local PDF parsing with full feature extraction enabled.
from pathlib import Path
from mineru.cli.common import do_parse, read_fn
# Define input path
pdf_path = Path("documents/input.pdf")
# Convert file to PDF bytes (handles images automatically)
pdf_bytes = read_fn(pdf_path)
# Execute parsing
do_parse(
output_dir="./results",
pdf_file_names=[pdf_path.stem],
pdf_bytes_list=[pdf_bytes],
p_lang_list=["ch"], # Language hint: "ch" for Chinese
backend="pipeline", # Local processing backend
parse_method="auto", # "auto", "txt", or "ocr"
formula_enable=True, # Enable formula detection
table_enable=True, # Enable table structure recognition
start_page_id=0, # Start from first page
end_page_id=None, # Process until end
)
This creates a directory structure under ./results/input/auto/ containing input.md, input_middle.json, and extracted images.
Asynchronous Parsing with Hybrid Backend
Use aio_do_parse when integrating with remote VLM services for hybrid processing that combines local OCR with cloud-based layout analysis.
import asyncio
from pathlib import Path
from mineru.cli.common import aio_do_parse, read_fn
async def parse_remote():
pdf_path = Path("documents/complex_layout.pdf")
pdf_bytes = read_fn(pdf_path)
await aio_do_parse(
output_dir="./hybrid_results",
pdf_file_names=[pdf_path.stem],
pdf_bytes_list=[pdf_bytes],
p_lang_list=["en"],
backend="hybrid-auto-engine", # Hybrid backend
parse_method="auto",
server_url="http://127.0.0.1:30000", # VLM server endpoint
formula_enable=True,
table_enable=True,
)
# Run the async workflow
asyncio.run(parse_remote())
The hybrid backend in mineru/backend/hybrid/hybrid_analyze.py manages concurrent local OCR and remote VLM inference, making this approach suitable for high-accuracy extraction of complex academic papers.
Custom Page Ranges and Output Configuration
Control which pages get processed and which output formats get generated using optional flags.
from pathlib import Path
from mineru.cli.common import do_parse, read_fn
pdf_path = Path("documents/long_document.pdf")
pdf_bytes = read_fn(pdf_path)
do_parse(
output_dir="partial_results",
pdf_file_names=[pdf_path.stem],
pdf_bytes_list=[pdf_bytes],
p_lang_list=["en"],
start_page_id=10, # Skip first 10 pages
end_page_id=20, # Stop after page 21
f_draw_layout_bbox=False, # Disable layout visualization PDF
f_draw_span_bbox=False, # Disable span visualization PDF
f_dump_md=True, # Enable Markdown output
f_dump_middle_json=True, # Enable middle-JSON output
f_dump_content_list=True, # Enable content-list JSON
)
The _prepare_pdf_bytes function in mineru/cli/common.py uses pypdfium2 to efficiently slice the PDF before analysis, reducing memory usage for large documents.
Key Source Files and Architecture
Understanding the repository structure helps you extend or debug the API:
| Component | Path | Responsibility |
|---|---|---|
| Public API | mineru/cli/common.py |
Exports do_parse, aio_do_parse, read_fn, and orchestration logic |
| CLI Wrapper | mineru/cli/client.py |
Command-line interface that calls the common API |
| Pipeline Backend | mineru/backend/pipeline/pipeline_analyze.py |
Local OCR, formula, and table extraction via pipeline_doc_analyze |
| VLM Backend | mineru/backend/vlm/vlm_analyze.py |
Remote Vision-Language Model processing via vlm_doc_analyze |
| Hybrid Backend | mineru/backend/hybrid/hybrid_analyze.py |
Combines local OCR with remote VLM via aio_hybrid_doc_analyze |
| PDF Utilities | mineru/utils/pdf_page_id.py |
Page range calculations and indexing |
| Image Conversion | mineru/utils/pdf_image_tools.py |
Converts images to PDF bytes for uniform processing |
| Output Enums | mineru/utils/enum_class.py |
Defines MakeMode constants for output format selection |
Summary
- MinerU's Python API centers on two functions: synchronous
do_parseand asynchronousaio_do_parse, both defined inmineru/cli/common.py. - Input handling is managed by
read_fn, which normalizes images to PDF bytes and supports automatic format detection. - Three backends are available:
pipeline(fully local),vlm-*(remote Vision-Language Model), andhybrid-*(combined approach), selected via thebackendparameter. - Page ranges can be sliced efficiently using
start_page_idandend_page_id, which invokepypdfium2via_prepare_pdf_bytesto reduce memory footprint. - Output artifacts include Markdown, middle-JSON, content-list JSON, and optional visualization PDFs, controlled by boolean flags like
f_dump_mdandf_draw_layout_bbox.
Frequently Asked Questions
How do I parse an image file instead of a PDF using MinerU's Python API?
Use the read_fn utility from mineru/cli/common.py to handle image inputs. This function automatically detects image file extensions (PNG, JPG, etc.) and converts them to PDF byte streams using the conversion utilities in mineru/utils/pdf_image_tools.py before passing them to do_parse or aio_do_parse.
What is the difference between the pipeline, VLM, and hybrid backends in MinerU?
The pipeline backend runs entirely locally using OCR, formula detection, and table structure models defined in mineru/backend/pipeline/pipeline_analyze.py. The VLM backend sends document images to a remote Vision-Language Model server via vlm_doc_analyze in mineru/backend/vlm/vlm_analyze.py. The hybrid backend combines both approaches, using local OCR for text and remote VLM for layout analysis via aio_hybrid_doc_analyze in mineru/backend/hybrid/hybrid_analyze.py.
How can I process only specific pages of a large PDF to save memory?
Pass the start_page_id and end_page_id parameters to do_parse or aio_do_parse. These trigger _prepare_pdf_bytes in mineru/cli/common.py, which uses pypdfium2 (via convert_pdf_bytes_to_bytes_by_pypdfium2) to slice the PDF before analysis. This prevents loading the entire document into memory and significantly reduces processing time for large files.
Which output formats does MinerU's Python API generate, and how do I control them?
The API generates four primary output types: Markdown (*.md), middle-JSON (*_middle.json), content-list JSON (*_content_list.json), and model-output JSON (*_model.json). Additionally, you can enable visualization PDFs showing layout and span bounding boxes. Control these via boolean flags in do_parse or aio_do_parse: f_dump_md, f_dump_middle_json, f_dump_content_list, f_draw_layout_bbox, and f_draw_span_bbox. These flags are processed by _process_output in mineru/cli/common.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →