How to Use MinerU's Python API for Programmatic PDF Parsing

MinerU exposes synchronous do_parse and asynchronous aio_do_parse functions in mineru/cli/common.py that let you parse PDFs programmatically without touching the command line.

The opendatalab/MinerU repository provides a high-level Python API for extracting structured Markdown, JSON, and images from PDF documents. Whether you need local OCR processing, remote Vision-Language Model (VLM) inference, or a hybrid approach, you can orchestrate the entire workflow through Python function calls rather than shell commands.

Core API Components

The public API surface is intentionally narrow, consisting of two primary entry points and several helper utilities located in mineru/cli/common.py.

do_parse – The synchronous entry point for local pipeline processing. It handles backend selection, PDF normalization, and output generation in a blocking call.

aio_do_parse – The asynchronous counterpart designed for I/O-bound backends like hybrid-* or vlm-http-client that communicate with remote servers.

read_fn – A utility function that detects file extensions and automatically converts image inputs (PNG, JPG) into PDF byte streams before processing.

Parsing Workflow

When you invoke do_parse or aio_do_parse, MinerU executes a five-stage pipeline internally. Understanding these stages helps you configure parameters correctly.

Reading and Normalizing Input

First, read_fn (lines 32-43 in mineru/cli/common.py) ingests the source file. If you provide an image, it converts the bytes to a single-page PDF using utilities in mineru/utils/pdf_image_tools.py.

Next, _prepare_pdf_bytes (lines 54-81) handles page range slicing. If you specify start_page_id and end_page_id, this function uses pypdfium2 via convert_pdf_bytes_to_bytes_by_pypdfium2 to trim the PDF before analysis.

Selecting the Backend

The backend parameter determines which analysis engine runs. Around lines 39-45 in mineru/cli/common.py, the code branches based on your selection:

Processing and Output Generation

For the pipeline backend, _process_pipeline calls pipeline_doc_analyze (imported from mineru/backend/pipeline/pipeline_analyze.py). This returns three key structures: infer_results (model predictions), all_image_lists (rendered page images), and all_pdf_docs (PDF objects).

Finally, _process_output (lines 94-168 in mineru/cli/common.py) writes the results to disk. It generates Markdown via pipeline_union_make or vlm_union_make, content-list JSON, middle-JSON, and optional visualization PDFs (layout, span, line-sort).

Code Examples

Synchronous Parsing with the Pipeline Backend

The following example demonstrates the most common use case: local PDF parsing with full feature extraction enabled.

from pathlib import Path
from mineru.cli.common import do_parse, read_fn

# Define input path

pdf_path = Path("documents/input.pdf")

# Convert file to PDF bytes (handles images automatically)

pdf_bytes = read_fn(pdf_path)

# Execute parsing

do_parse(
    output_dir="./results",
    pdf_file_names=[pdf_path.stem],
    pdf_bytes_list=[pdf_bytes],
    p_lang_list=["ch"],              # Language hint: "ch" for Chinese

    backend="pipeline",              # Local processing backend

    parse_method="auto",             # "auto", "txt", or "ocr"

    formula_enable=True,             # Enable formula detection

    table_enable=True,               # Enable table structure recognition

    start_page_id=0,                 # Start from first page

    end_page_id=None,                # Process until end

)

This creates a directory structure under ./results/input/auto/ containing input.md, input_middle.json, and extracted images.

Asynchronous Parsing with Hybrid Backend

Use aio_do_parse when integrating with remote VLM services for hybrid processing that combines local OCR with cloud-based layout analysis.

import asyncio
from pathlib import Path
from mineru.cli.common import aio_do_parse, read_fn

async def parse_remote():
    pdf_path = Path("documents/complex_layout.pdf")
    pdf_bytes = read_fn(pdf_path)
    
    await aio_do_parse(
        output_dir="./hybrid_results",
        pdf_file_names=[pdf_path.stem],
        pdf_bytes_list=[pdf_bytes],
        p_lang_list=["en"],
        backend="hybrid-auto-engine",        # Hybrid backend

        parse_method="auto",
        server_url="http://127.0.0.1:30000", # VLM server endpoint

        formula_enable=True,
        table_enable=True,
    )

# Run the async workflow

asyncio.run(parse_remote())

The hybrid backend in mineru/backend/hybrid/hybrid_analyze.py manages concurrent local OCR and remote VLM inference, making this approach suitable for high-accuracy extraction of complex academic papers.

Custom Page Ranges and Output Configuration

Control which pages get processed and which output formats get generated using optional flags.

from pathlib import Path
from mineru.cli.common import do_parse, read_fn

pdf_path = Path("documents/long_document.pdf")
pdf_bytes = read_fn(pdf_path)

do_parse(
    output_dir="partial_results",
    pdf_file_names=[pdf_path.stem],
    pdf_bytes_list=[pdf_bytes],
    p_lang_list=["en"],
    start_page_id=10,              # Skip first 10 pages

    end_page_id=20,                # Stop after page 21

    f_draw_layout_bbox=False,      # Disable layout visualization PDF

    f_draw_span_bbox=False,        # Disable span visualization PDF

    f_dump_md=True,                # Enable Markdown output

    f_dump_middle_json=True,       # Enable middle-JSON output

    f_dump_content_list=True,      # Enable content-list JSON

)

The _prepare_pdf_bytes function in mineru/cli/common.py uses pypdfium2 to efficiently slice the PDF before analysis, reducing memory usage for large documents.

Key Source Files and Architecture

Understanding the repository structure helps you extend or debug the API:

Component Path Responsibility
Public API mineru/cli/common.py Exports do_parse, aio_do_parse, read_fn, and orchestration logic
CLI Wrapper mineru/cli/client.py Command-line interface that calls the common API
Pipeline Backend mineru/backend/pipeline/pipeline_analyze.py Local OCR, formula, and table extraction via pipeline_doc_analyze
VLM Backend mineru/backend/vlm/vlm_analyze.py Remote Vision-Language Model processing via vlm_doc_analyze
Hybrid Backend mineru/backend/hybrid/hybrid_analyze.py Combines local OCR with remote VLM via aio_hybrid_doc_analyze
PDF Utilities mineru/utils/pdf_page_id.py Page range calculations and indexing
Image Conversion mineru/utils/pdf_image_tools.py Converts images to PDF bytes for uniform processing
Output Enums mineru/utils/enum_class.py Defines MakeMode constants for output format selection

Summary

  • MinerU's Python API centers on two functions: synchronous do_parse and asynchronous aio_do_parse, both defined in mineru/cli/common.py.
  • Input handling is managed by read_fn, which normalizes images to PDF bytes and supports automatic format detection.
  • Three backends are available: pipeline (fully local), vlm-* (remote Vision-Language Model), and hybrid-* (combined approach), selected via the backend parameter.
  • Page ranges can be sliced efficiently using start_page_id and end_page_id, which invoke pypdfium2 via _prepare_pdf_bytes to reduce memory footprint.
  • Output artifacts include Markdown, middle-JSON, content-list JSON, and optional visualization PDFs, controlled by boolean flags like f_dump_md and f_draw_layout_bbox.

Frequently Asked Questions

How do I parse an image file instead of a PDF using MinerU's Python API?

Use the read_fn utility from mineru/cli/common.py to handle image inputs. This function automatically detects image file extensions (PNG, JPG, etc.) and converts them to PDF byte streams using the conversion utilities in mineru/utils/pdf_image_tools.py before passing them to do_parse or aio_do_parse.

What is the difference between the pipeline, VLM, and hybrid backends in MinerU?

The pipeline backend runs entirely locally using OCR, formula detection, and table structure models defined in mineru/backend/pipeline/pipeline_analyze.py. The VLM backend sends document images to a remote Vision-Language Model server via vlm_doc_analyze in mineru/backend/vlm/vlm_analyze.py. The hybrid backend combines both approaches, using local OCR for text and remote VLM for layout analysis via aio_hybrid_doc_analyze in mineru/backend/hybrid/hybrid_analyze.py.

How can I process only specific pages of a large PDF to save memory?

Pass the start_page_id and end_page_id parameters to do_parse or aio_do_parse. These trigger _prepare_pdf_bytes in mineru/cli/common.py, which uses pypdfium2 (via convert_pdf_bytes_to_bytes_by_pypdfium2) to slice the PDF before analysis. This prevents loading the entire document into memory and significantly reduces processing time for large files.

Which output formats does MinerU's Python API generate, and how do I control them?

The API generates four primary output types: Markdown (*.md), middle-JSON (*_middle.json), content-list JSON (*_content_list.json), and model-output JSON (*_model.json). Additionally, you can enable visualization PDFs showing layout and span bounding boxes. Control these via boolean flags in do_parse or aio_do_parse: f_dump_md, f_dump_middle_json, f_dump_content_list, f_draw_layout_bbox, and f_draw_span_bbox. These flags are processed by _process_output in mineru/cli/common.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →