# How to Use MinerU's Python API for Programmatic PDF Parsing

> Easily parse PDFs programmatically with MinerUs Python API. Utilize do_parse and aio_do_parse functions for seamless document extraction without the command line.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: how-to-guide
- Published: 2026-02-22

---

**MinerU exposes synchronous `do_parse` and asynchronous `aio_do_parse` functions in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py) that let you parse PDFs programmatically without touching the command line.**

The [opendatalab/MinerU](https://github.com/opendatalab/MinerU) repository provides a high-level Python API for extracting structured Markdown, JSON, and images from PDF documents. Whether you need local OCR processing, remote Vision-Language Model (VLM) inference, or a hybrid approach, you can orchestrate the entire workflow through Python function calls rather than shell commands.

## Core API Components

The public API surface is intentionally narrow, consisting of two primary entry points and several helper utilities located in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py).

**`do_parse`** – The synchronous entry point for local pipeline processing. It handles backend selection, PDF normalization, and output generation in a blocking call.

**`aio_do_parse`** – The asynchronous counterpart designed for I/O-bound backends like `hybrid-*` or `vlm-http-client` that communicate with remote servers.

**`read_fn`** – A utility function that detects file extensions and automatically converts image inputs (PNG, JPG) into PDF byte streams before processing.

## Parsing Workflow

When you invoke `do_parse` or `aio_do_parse`, MinerU executes a five-stage pipeline internally. Understanding these stages helps you configure parameters correctly.

### Reading and Normalizing Input

First, `read_fn` (lines 32-43 in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py)) ingests the source file. If you provide an image, it converts the bytes to a single-page PDF using utilities in [`mineru/utils/pdf_image_tools.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_image_tools.py).

Next, `_prepare_pdf_bytes` (lines 54-81) handles page range slicing. If you specify `start_page_id` and `end_page_id`, this function uses `pypdfium2` via `convert_pdf_bytes_to_bytes_by_pypdfium2` to trim the PDF before analysis.

### Selecting the Backend

The `backend` parameter determines which analysis engine runs. Around lines 39-45 in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py), the code branches based on your selection:

- **`pipeline`** – Default local processing using OCR, formula, and table detection models.
- **`vlm-*`** – Routes to `vlm_doc_analyze` in [`mineru/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/vlm/vlm_analyze.py) for remote VLM processing.
- **`hybrid-*`** – Combines local OCR with remote VLM via `aio_hybrid_doc_analyze` in [`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py).

### Processing and Output Generation

For the **pipeline** backend, `_process_pipeline` calls `pipeline_doc_analyze` (imported from [`mineru/backend/pipeline/pipeline_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_analyze.py)). This returns three key structures: `infer_results` (model predictions), `all_image_lists` (rendered page images), and `all_pdf_docs` (PDF objects).

Finally, `_process_output` (lines 94-168 in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py)) writes the results to disk. It generates Markdown via `pipeline_union_make` or `vlm_union_make`, content-list JSON, middle-JSON, and optional visualization PDFs (layout, span, line-sort).

## Code Examples

### Synchronous Parsing with the Pipeline Backend

The following example demonstrates the most common use case: local PDF parsing with full feature extraction enabled.

```python
from pathlib import Path
from mineru.cli.common import do_parse, read_fn

# Define input path

pdf_path = Path("documents/input.pdf")

# Convert file to PDF bytes (handles images automatically)

pdf_bytes = read_fn(pdf_path)

# Execute parsing

do_parse(
    output_dir="./results",
    pdf_file_names=[pdf_path.stem],
    pdf_bytes_list=[pdf_bytes],
    p_lang_list=["ch"],              # Language hint: "ch" for Chinese

    backend="pipeline",              # Local processing backend

    parse_method="auto",             # "auto", "txt", or "ocr"

    formula_enable=True,             # Enable formula detection

    table_enable=True,               # Enable table structure recognition

    start_page_id=0,                 # Start from first page

    end_page_id=None,                # Process until end

)

```

This creates a directory structure under `./results/input/auto/` containing [`input.md`](https://github.com/opendatalab/MinerU/blob/main/input.md), [`input_middle.json`](https://github.com/opendatalab/MinerU/blob/main/input_middle.json), and extracted images.

### Asynchronous Parsing with Hybrid Backend

Use `aio_do_parse` when integrating with remote VLM services for hybrid processing that combines local OCR with cloud-based layout analysis.

```python
import asyncio
from pathlib import Path
from mineru.cli.common import aio_do_parse, read_fn

async def parse_remote():
    pdf_path = Path("documents/complex_layout.pdf")
    pdf_bytes = read_fn(pdf_path)
    
    await aio_do_parse(
        output_dir="./hybrid_results",
        pdf_file_names=[pdf_path.stem],
        pdf_bytes_list=[pdf_bytes],
        p_lang_list=["en"],
        backend="hybrid-auto-engine",        # Hybrid backend

        parse_method="auto",
        server_url="http://127.0.0.1:30000", # VLM server endpoint

        formula_enable=True,
        table_enable=True,
    )

# Run the async workflow

asyncio.run(parse_remote())

```

The hybrid backend in [`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py) manages concurrent local OCR and remote VLM inference, making this approach suitable for high-accuracy extraction of complex academic papers.

### Custom Page Ranges and Output Configuration

Control which pages get processed and which output formats get generated using optional flags.

```python
from pathlib import Path
from mineru.cli.common import do_parse, read_fn

pdf_path = Path("documents/long_document.pdf")
pdf_bytes = read_fn(pdf_path)

do_parse(
    output_dir="partial_results",
    pdf_file_names=[pdf_path.stem],
    pdf_bytes_list=[pdf_bytes],
    p_lang_list=["en"],
    start_page_id=10,              # Skip first 10 pages

    end_page_id=20,                # Stop after page 21

    f_draw_layout_bbox=False,      # Disable layout visualization PDF

    f_draw_span_bbox=False,        # Disable span visualization PDF

    f_dump_md=True,                # Enable Markdown output

    f_dump_middle_json=True,       # Enable middle-JSON output

    f_dump_content_list=True,      # Enable content-list JSON

)

```

The `_prepare_pdf_bytes` function in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py) uses `pypdfium2` to efficiently slice the PDF before analysis, reducing memory usage for large documents.

## Key Source Files and Architecture

Understanding the repository structure helps you extend or debug the API:

| Component | Path | Responsibility |
|-----------|------|----------------|
| **Public API** | [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py) | Exports `do_parse`, `aio_do_parse`, `read_fn`, and orchestration logic |
| **CLI Wrapper** | [`mineru/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/client.py) | Command-line interface that calls the common API |
| **Pipeline Backend** | [`mineru/backend/pipeline/pipeline_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_analyze.py) | Local OCR, formula, and table extraction via `pipeline_doc_analyze` |
| **VLM Backend** | [`mineru/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/vlm/vlm_analyze.py) | Remote Vision-Language Model processing via `vlm_doc_analyze` |
| **Hybrid Backend** | [`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py) | Combines local OCR with remote VLM via `aio_hybrid_doc_analyze` |
| **PDF Utilities** | [`mineru/utils/pdf_page_id.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_page_id.py) | Page range calculations and indexing |
| **Image Conversion** | [`mineru/utils/pdf_image_tools.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_image_tools.py) | Converts images to PDF bytes for uniform processing |
| **Output Enums** | [`mineru/utils/enum_class.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/enum_class.py) | Defines `MakeMode` constants for output format selection |

## Summary

- **MinerU's Python API** centers on two functions: synchronous `do_parse` and asynchronous `aio_do_parse`, both defined in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py).
- **Input handling** is managed by `read_fn`, which normalizes images to PDF bytes and supports automatic format detection.
- **Three backends** are available: `pipeline` (fully local), `vlm-*` (remote Vision-Language Model), and `hybrid-*` (combined approach), selected via the `backend` parameter.
- **Page ranges** can be sliced efficiently using `start_page_id` and `end_page_id`, which invoke `pypdfium2` via `_prepare_pdf_bytes` to reduce memory footprint.
- **Output artifacts** include Markdown, middle-JSON, content-list JSON, and optional visualization PDFs, controlled by boolean flags like `f_dump_md` and `f_draw_layout_bbox`.

## Frequently Asked Questions

### How do I parse an image file instead of a PDF using MinerU's Python API?

Use the `read_fn` utility from [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py) to handle image inputs. This function automatically detects image file extensions (PNG, JPG, etc.) and converts them to PDF byte streams using the conversion utilities in [`mineru/utils/pdf_image_tools.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_image_tools.py) before passing them to `do_parse` or `aio_do_parse`.

### What is the difference between the pipeline, VLM, and hybrid backends in MinerU?

The **pipeline** backend runs entirely locally using OCR, formula detection, and table structure models defined in [`mineru/backend/pipeline/pipeline_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_analyze.py). The **VLM** backend sends document images to a remote Vision-Language Model server via `vlm_doc_analyze` in [`mineru/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/vlm/vlm_analyze.py). The **hybrid** backend combines both approaches, using local OCR for text and remote VLM for layout analysis via `aio_hybrid_doc_analyze` in [`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py).

### How can I process only specific pages of a large PDF to save memory?

Pass the `start_page_id` and `end_page_id` parameters to `do_parse` or `aio_do_parse`. These trigger `_prepare_pdf_bytes` in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py), which uses `pypdfium2` (via `convert_pdf_bytes_to_bytes_by_pypdfium2`) to slice the PDF before analysis. This prevents loading the entire document into memory and significantly reduces processing time for large files.

### Which output formats does MinerU's Python API generate, and how do I control them?

The API generates four primary output types: **Markdown** (`*.md`), **middle-JSON** (`*_middle.json`), **content-list JSON** (`*_content_list.json`), and **model-output JSON** (`*_model.json`). Additionally, you can enable visualization PDFs showing layout and span bounding boxes. Control these via boolean flags in `do_parse` or `aio_do_parse`: `f_dump_md`, `f_dump_middle_json`, `f_dump_content_list`, `f_draw_layout_bbox`, and `f_draw_span_bbox`. These flags are processed by `_process_output` in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py).