How to Convert PDF to JSON Using MinerU: A Complete Technical Guide

MinerU converts PDF documents to structured JSON through a six-stage pipeline that reads PDF bytes via read_fn in mineru/cli/common.py, preprocesses them with convert_pdf_bytes_to_bytes_by_pypdfium2, runs MagicModel inference, and constructs a middle-JSON representation via result_to_middle_json in mineru/backend/pipeline/model_json_to_middle_json.py.

MinerU is an open-source PDF parsing toolkit that extracts structured content from documents. When you convert PDF to JSON using MinerU, the tool generates a "middle-JSON" format containing layout information, text blocks, images, tables, and metadata. This structured output enables downstream applications to consume PDF content programmatically without handling raw binary data.

The MinerU PDF to JSON Pipeline Architecture

The conversion process follows a deterministic pipeline implemented across several core modules. Understanding these stages helps you customize the extraction for specific document types.

File Intake and Preprocessing

The entry point begins in mineru/cli/common.py, where the read_fn function handles PDF loading【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L32-L44】. This utility accepts both PDF files and images (converting images to PDF format internally).

Once loaded, the convert_pdf_bytes_to_bytes_by_pypdfium2 function normalizes the raw bytes to eliminate corrupted pages using the pypdfium2 library【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L54-L80】. This preprocessing ensures downstream models receive clean input.

Model Inference and Analysis

After preprocessing, MinerU routes the document through the selected backend. In pipeline mode (the default), the do_parse function calls pipeline_doc_analyze, which executes the MagicModel on each page. This produces a model_list containing detected layout elements, text regions, and bounding boxes.

For VLM or Hybrid modes, alternative backends process the content, but the subsequent JSON construction phase remains identical.

Middle-JSON Construction

The critical transformation occurs in mineru/backend/pipeline/model_json_to_middle_json.py. The result_to_middle_json function receives the model_list from inference, extracted page images, the original PdfDocument object, and an image writer instance【/cache/repos/github.com/opendatalab/MinerU/master/mineru/backend/pipeline/model_json_to_middle_json.py#L76-L89】.

This function assembles a dictionary with a "pdf_info" key containing per-page arrays of layout blocks, text spans, bounding boxes, images, tables, and mathematical equations【/cache/repos/github.com/opendatalab/MinerU/master/mineru/backend/pipeline/model_json_to_middle_json.py#L118-L130】. The result is a JSON-serializable Python object.

Finally, the _process_output function in mineru/cli/common.py persists the middle-JSON to disk when f_dump_middle_json is True (default), using the naming pattern <pdf_name>_middle.json【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L56-L60】.

Converting PDF to JSON via Command Line

The simplest way to convert PDF to JSON using MinerU is through the CLI. Ensure you have installed MinerU and configured your model weights, then execute:

mineru parse \
    -p /path/to/document.pdf \
    -o ./output \
    --backend pipeline \
    --method auto \
    --return_middle_json True

This command processes document.pdf and creates ./output/document/document_middle.json. The --return_middle_json True flag ensures the middle-JSON file is written to disk alongside other output formats.

Converting PDF to JSON via Python API

For programmatic integration, import the core utilities from mineru.cli.common to process PDFs within your Python application.

Basic Pipeline Usage

from mineru.cli.common import read_fn, do_parse
from pathlib import Path

pdf_path = Path("research_paper.pdf")
pdf_bytes = read_fn(pdf_path)  # Loads PDF or converts image to PDF

do_parse(
    output_dir="./results",
    pdf_file_names=[pdf_path.stem],
    pdf_bytes_list=[pdf_bytes],
    p_lang_list=["en"],           # Language hint: "en", "ch", etc.

    backend="pipeline",
    parse_method="auto",
    f_dump_middle_json=True,      # Required to output JSON

)

# Generates: ./results/research_paper/research_paper_middle.json

Direct JSON Builder Access

For custom pipelines where you already have model predictions, call result_to_middle_json directly:

from mineru.backend.pipeline.model_json_to_middle_json import result_to_middle_json
from mineru.data.data_reader_writer.filebase import FileBasedDataWriter
from mineru.utils.pdf_reader import pdf_to_images
from pypdfium2 import PdfDocument

# Assume model_list contains MagicModel outputs and pdf_bytes is your source

pdf_doc = PdfDocument(pdf_bytes)
images = pdf_to_images(pdf_bytes)  # List of PIL.Image objects

# Writer handles image extraction to disk

image_writer = FileBasedDataWriter("./extracted_images")

middle_json = result_to_middle_json(
    model_list,
    images,
    pdf_doc,
    image_writer,
    lang="en",
    ocr_enable=False,
    formula_enabled=True,
)

# Serialize to JSON file

import json
with open("output_middle.json", "w", encoding="utf-8") as f:
    json.dump(middle_json, f, ensure_ascii=False, indent=4)

Key Source Files for PDF to JSON Conversion

The following modules implement the core logic for converting PDF to JSON using MinerU:

File Purpose Key Functions
mineru/cli/common.py CLI utilities and pipeline orchestration read_fn, do_parse, convert_pdf_bytes_to_bytes_by_pypdfium2, _process_output
mineru/backend/pipeline/model_json_to_middle_json.py Middle-JSON generation from model outputs result_to_middle_json
mineru/backend/pipeline/pipeline_middle_json_mkcontent.py Post-processing middle-JSON into other formats Content rendering utilities
mineru/utils/pdf_reader.py PDF to image conversion for OCR and analysis pdf_to_images
mineru/data/data_reader_writer/filebase.py File I/O abstraction for extracted assets FileBasedDataWriter
mineru/cli/fast_api.py HTTP API endpoint serving JSON output /file_parse endpoint

Summary

  • MinerU generates a middle-JSON format that captures document structure, layout elements, and metadata during PDF processing.
  • The conversion pipeline in mineru/cli/common.py handles file intake, preprocessing with pypdfium2, and model inference before JSON construction.
  • The result_to_middle_json function in mineru/backend/pipeline/model_json_to_middle_json.py assembles the final JSON structure containing pdf_info with per-page layout blocks, text, images, and equations.
  • You can convert PDF to JSON using MinerU via the CLI with --return_middle_json True or programmatically using the Python API by setting f_dump_middle_json=True in do_parse.
  • For custom workflows, import result_to_middle_json directly to build JSON objects from existing model predictions.

Frequently Asked Questions

What is the middle-JSON format in MinerU?

The middle-JSON format is an intermediate representation generated during PDF processing that contains structured document information. It includes a pdf_info key with per-page arrays of layout blocks, text spans, bounding boxes, images, tables, and mathematical equations. This format serves as the foundation for downstream conversions to Markdown or other output formats.

How do I enable JSON output when using the MinerU CLI?

To generate JSON output via the command line, append the --return_middle_json True flag to your mineru parse command. By default, MinerU writes the middle-JSON file to the output directory using the naming pattern <pdf_name>_middle.json. You can also set f_dump_middle_json=True when using the Python API.

Can I convert PDF to JSON programmatically without using the CLI?

Yes, you can convert PDF to JSON programmatically by importing functions from mineru.cli.common. Use read_fn to load PDF bytes and do_parse with f_dump_middle_json=True to execute the pipeline. For advanced use cases, you can directly call result_to_middle_json from mineru.backend.pipeline.model_json_to_middle_json to build JSON objects from custom model outputs.

Where does MinerU handle image extraction during JSON conversion?

Image extraction occurs within the result_to_middle_json function in mineru/backend/pipeline/model_json_to_middle_json.py. This function receives an image writer instance (typically FileBasedDataWriter from mineru/data/data_reader_writer/filebase.py) that handles persisting extracted images to disk. The image paths are then referenced within the middle-JSON structure under the respective layout blocks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →