# How to Convert PDF to JSON Using MinerU: A Complete Technical Guide

> Learn to convert PDF to JSON with MinerU. This technical guide details the six-stage pipeline for structured data extraction from PDFs using MagicModel inference. Explore the opendatalab MinerU repository.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: how-to-guide
- Published: 2026-02-23

---

**MinerU converts PDF documents to structured JSON through a six-stage pipeline that reads PDF bytes via `read_fn` in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py), preprocesses them with `convert_pdf_bytes_to_bytes_by_pypdfium2`, runs MagicModel inference, and constructs a middle-JSON representation via `result_to_middle_json` in [`mineru/backend/pipeline/model_json_to_middle_json.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/model_json_to_middle_json.py).**

MinerU is an open-source PDF parsing toolkit that extracts structured content from documents. When you convert PDF to JSON using MinerU, the tool generates a "middle-JSON" format containing layout information, text blocks, images, tables, and metadata. This structured output enables downstream applications to consume PDF content programmatically without handling raw binary data.

## The MinerU PDF to JSON Pipeline Architecture

The conversion process follows a deterministic pipeline implemented across several core modules. Understanding these stages helps you customize the extraction for specific document types.

### File Intake and Preprocessing

The entry point begins in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py), where the `read_fn` function handles PDF loading【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L32-L44】. This utility accepts both PDF files and images (converting images to PDF format internally).

Once loaded, the `convert_pdf_bytes_to_bytes_by_pypdfium2` function normalizes the raw bytes to eliminate corrupted pages using the pypdfium2 library【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L54-L80】. This preprocessing ensures downstream models receive clean input.

### Model Inference and Analysis

After preprocessing, MinerU routes the document through the selected backend. In **pipeline** mode (the default), the `do_parse` function calls `pipeline_doc_analyze`, which executes the **MagicModel** on each page. This produces a `model_list` containing detected layout elements, text regions, and bounding boxes.

For **VLM** or **Hybrid** modes, alternative backends process the content, but the subsequent JSON construction phase remains identical.

### Middle-JSON Construction

The critical transformation occurs in [`mineru/backend/pipeline/model_json_to_middle_json.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/model_json_to_middle_json.py). The `result_to_middle_json` function receives the `model_list` from inference, extracted page images, the original `PdfDocument` object, and an image writer instance【/cache/repos/github.com/opendatalab/MinerU/master/mineru/backend/pipeline/model_json_to_middle_json.py#L76-L89】.

This function assembles a dictionary with a `"pdf_info"` key containing per-page arrays of layout blocks, text spans, bounding boxes, images, tables, and mathematical equations【/cache/repos/github.com/opendatalab/MinerU/master/mineru/backend/pipeline/model_json_to_middle_json.py#L118-L130】. The result is a JSON-serializable Python object.

Finally, the `_process_output` function in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py) persists the middle-JSON to disk when `f_dump_middle_json` is `True` (default), using the naming pattern `<pdf_name>_middle.json`【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L56-L60】.

## Converting PDF to JSON via Command Line

The simplest way to convert PDF to JSON using MinerU is through the CLI. Ensure you have installed MinerU and configured your model weights, then execute:

```bash
mineru parse \
    -p /path/to/document.pdf \
    -o ./output \
    --backend pipeline \
    --method auto \
    --return_middle_json True

```

This command processes `document.pdf` and creates [`./output/document/document_middle.json`](https://github.com/opendatalab/MinerU/blob/main/./output/document/document_middle.json). The `--return_middle_json True` flag ensures the middle-JSON file is written to disk alongside other output formats.

## Converting PDF to JSON via Python API

For programmatic integration, import the core utilities from `mineru.cli.common` to process PDFs within your Python application.

### Basic Pipeline Usage

```python
from mineru.cli.common import read_fn, do_parse
from pathlib import Path

pdf_path = Path("research_paper.pdf")
pdf_bytes = read_fn(pdf_path)  # Loads PDF or converts image to PDF

do_parse(
    output_dir="./results",
    pdf_file_names=[pdf_path.stem],
    pdf_bytes_list=[pdf_bytes],
    p_lang_list=["en"],           # Language hint: "en", "ch", etc.

    backend="pipeline",
    parse_method="auto",
    f_dump_middle_json=True,      # Required to output JSON

)

# Generates: ./results/research_paper/research_paper_middle.json

```

### Direct JSON Builder Access

For custom pipelines where you already have model predictions, call `result_to_middle_json` directly:

```python
from mineru.backend.pipeline.model_json_to_middle_json import result_to_middle_json
from mineru.data.data_reader_writer.filebase import FileBasedDataWriter
from mineru.utils.pdf_reader import pdf_to_images
from pypdfium2 import PdfDocument

# Assume model_list contains MagicModel outputs and pdf_bytes is your source

pdf_doc = PdfDocument(pdf_bytes)
images = pdf_to_images(pdf_bytes)  # List of PIL.Image objects

# Writer handles image extraction to disk

image_writer = FileBasedDataWriter("./extracted_images")

middle_json = result_to_middle_json(
    model_list,
    images,
    pdf_doc,
    image_writer,
    lang="en",
    ocr_enable=False,
    formula_enabled=True,
)

# Serialize to JSON file

import json
with open("output_middle.json", "w", encoding="utf-8") as f:
    json.dump(middle_json, f, ensure_ascii=False, indent=4)

```

## Key Source Files for PDF to JSON Conversion

The following modules implement the core logic for converting PDF to JSON using MinerU:

| File | Purpose | Key Functions |
|------|---------|---------------|
| [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py) | CLI utilities and pipeline orchestration | `read_fn`, `do_parse`, `convert_pdf_bytes_to_bytes_by_pypdfium2`, `_process_output` |
| [`mineru/backend/pipeline/model_json_to_middle_json.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/model_json_to_middle_json.py) | Middle-JSON generation from model outputs | `result_to_middle_json` |
| [`mineru/backend/pipeline/pipeline_middle_json_mkcontent.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_middle_json_mkcontent.py) | Post-processing middle-JSON into other formats | Content rendering utilities |
| [`mineru/utils/pdf_reader.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_reader.py) | PDF to image conversion for OCR and analysis | `pdf_to_images` |
| [`mineru/data/data_reader_writer/filebase.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/data/data_reader_writer/filebase.py) | File I/O abstraction for extracted assets | `FileBasedDataWriter` |
| [`mineru/cli/fast_api.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/fast_api.py) | HTTP API endpoint serving JSON output | `/file_parse` endpoint |

## Summary

- MinerU generates a **middle-JSON** format that captures document structure, layout elements, and metadata during PDF processing.
- The conversion pipeline in [`mineru/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/common.py) handles file intake, preprocessing with pypdfium2, and model inference before JSON construction.
- The `result_to_middle_json` function in [`mineru/backend/pipeline/model_json_to_middle_json.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/model_json_to_middle_json.py) assembles the final JSON structure containing `pdf_info` with per-page layout blocks, text, images, and equations.
- You can convert PDF to JSON using MinerU via the CLI with `--return_middle_json True` or programmatically using the Python API by setting `f_dump_middle_json=True` in `do_parse`.
- For custom workflows, import `result_to_middle_json` directly to build JSON objects from existing model predictions.

## Frequently Asked Questions

### What is the middle-JSON format in MinerU?

The middle-JSON format is an intermediate representation generated during PDF processing that contains structured document information. It includes a `pdf_info` key with per-page arrays of layout blocks, text spans, bounding boxes, images, tables, and mathematical equations. This format serves as the foundation for downstream conversions to Markdown or other output formats.

### How do I enable JSON output when using the MinerU CLI?

To generate JSON output via the command line, append the `--return_middle_json True` flag to your `mineru parse` command. By default, MinerU writes the middle-JSON file to the output directory using the naming pattern `<pdf_name>_middle.json`. You can also set `f_dump_middle_json=True` when using the Python API.

### Can I convert PDF to JSON programmatically without using the CLI?

Yes, you can convert PDF to JSON programmatically by importing functions from `mineru.cli.common`. Use `read_fn` to load PDF bytes and `do_parse` with `f_dump_middle_json=True` to execute the pipeline. For advanced use cases, you can directly call `result_to_middle_json` from `mineru.backend.pipeline.model_json_to_middle_json` to build JSON objects from custom model outputs.

### Where does MinerU handle image extraction during JSON conversion?

Image extraction occurs within the `result_to_middle_json` function in [`mineru/backend/pipeline/model_json_to_middle_json.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/model_json_to_middle_json.py). This function receives an image writer instance (typically `FileBasedDataWriter` from [`mineru/data/data_reader_writer/filebase.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/data/data_reader_writer/filebase.py)) that handles persisting extracted images to disk. The image paths are then referenced within the middle-JSON structure under the respective layout blocks.