How to Convert PDF to JSON Using MinerU: A Complete Technical Guide
MinerU converts PDF documents to structured JSON through a six-stage pipeline that reads PDF bytes via read_fn in mineru/cli/common.py, preprocesses them with convert_pdf_bytes_to_bytes_by_pypdfium2, runs MagicModel inference, and constructs a middle-JSON representation via result_to_middle_json in mineru/backend/pipeline/model_json_to_middle_json.py.
MinerU is an open-source PDF parsing toolkit that extracts structured content from documents. When you convert PDF to JSON using MinerU, the tool generates a "middle-JSON" format containing layout information, text blocks, images, tables, and metadata. This structured output enables downstream applications to consume PDF content programmatically without handling raw binary data.
The MinerU PDF to JSON Pipeline Architecture
The conversion process follows a deterministic pipeline implemented across several core modules. Understanding these stages helps you customize the extraction for specific document types.
File Intake and Preprocessing
The entry point begins in mineru/cli/common.py, where the read_fn function handles PDF loading【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L32-L44】. This utility accepts both PDF files and images (converting images to PDF format internally).
Once loaded, the convert_pdf_bytes_to_bytes_by_pypdfium2 function normalizes the raw bytes to eliminate corrupted pages using the pypdfium2 library【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L54-L80】. This preprocessing ensures downstream models receive clean input.
Model Inference and Analysis
After preprocessing, MinerU routes the document through the selected backend. In pipeline mode (the default), the do_parse function calls pipeline_doc_analyze, which executes the MagicModel on each page. This produces a model_list containing detected layout elements, text regions, and bounding boxes.
For VLM or Hybrid modes, alternative backends process the content, but the subsequent JSON construction phase remains identical.
Middle-JSON Construction
The critical transformation occurs in mineru/backend/pipeline/model_json_to_middle_json.py. The result_to_middle_json function receives the model_list from inference, extracted page images, the original PdfDocument object, and an image writer instance【/cache/repos/github.com/opendatalab/MinerU/master/mineru/backend/pipeline/model_json_to_middle_json.py#L76-L89】.
This function assembles a dictionary with a "pdf_info" key containing per-page arrays of layout blocks, text spans, bounding boxes, images, tables, and mathematical equations【/cache/repos/github.com/opendatalab/MinerU/master/mineru/backend/pipeline/model_json_to_middle_json.py#L118-L130】. The result is a JSON-serializable Python object.
Finally, the _process_output function in mineru/cli/common.py persists the middle-JSON to disk when f_dump_middle_json is True (default), using the naming pattern <pdf_name>_middle.json【/cache/repos/github.com/opendatalab/MinerU/master/mineru/cli/common.py#L56-L60】.
Converting PDF to JSON via Command Line
The simplest way to convert PDF to JSON using MinerU is through the CLI. Ensure you have installed MinerU and configured your model weights, then execute:
mineru parse \
-p /path/to/document.pdf \
-o ./output \
--backend pipeline \
--method auto \
--return_middle_json True
This command processes document.pdf and creates ./output/document/document_middle.json. The --return_middle_json True flag ensures the middle-JSON file is written to disk alongside other output formats.
Converting PDF to JSON via Python API
For programmatic integration, import the core utilities from mineru.cli.common to process PDFs within your Python application.
Basic Pipeline Usage
from mineru.cli.common import read_fn, do_parse
from pathlib import Path
pdf_path = Path("research_paper.pdf")
pdf_bytes = read_fn(pdf_path) # Loads PDF or converts image to PDF
do_parse(
output_dir="./results",
pdf_file_names=[pdf_path.stem],
pdf_bytes_list=[pdf_bytes],
p_lang_list=["en"], # Language hint: "en", "ch", etc.
backend="pipeline",
parse_method="auto",
f_dump_middle_json=True, # Required to output JSON
)
# Generates: ./results/research_paper/research_paper_middle.json
Direct JSON Builder Access
For custom pipelines where you already have model predictions, call result_to_middle_json directly:
from mineru.backend.pipeline.model_json_to_middle_json import result_to_middle_json
from mineru.data.data_reader_writer.filebase import FileBasedDataWriter
from mineru.utils.pdf_reader import pdf_to_images
from pypdfium2 import PdfDocument
# Assume model_list contains MagicModel outputs and pdf_bytes is your source
pdf_doc = PdfDocument(pdf_bytes)
images = pdf_to_images(pdf_bytes) # List of PIL.Image objects
# Writer handles image extraction to disk
image_writer = FileBasedDataWriter("./extracted_images")
middle_json = result_to_middle_json(
model_list,
images,
pdf_doc,
image_writer,
lang="en",
ocr_enable=False,
formula_enabled=True,
)
# Serialize to JSON file
import json
with open("output_middle.json", "w", encoding="utf-8") as f:
json.dump(middle_json, f, ensure_ascii=False, indent=4)
Key Source Files for PDF to JSON Conversion
The following modules implement the core logic for converting PDF to JSON using MinerU:
| File | Purpose | Key Functions |
|---|---|---|
mineru/cli/common.py |
CLI utilities and pipeline orchestration | read_fn, do_parse, convert_pdf_bytes_to_bytes_by_pypdfium2, _process_output |
mineru/backend/pipeline/model_json_to_middle_json.py |
Middle-JSON generation from model outputs | result_to_middle_json |
mineru/backend/pipeline/pipeline_middle_json_mkcontent.py |
Post-processing middle-JSON into other formats | Content rendering utilities |
mineru/utils/pdf_reader.py |
PDF to image conversion for OCR and analysis | pdf_to_images |
mineru/data/data_reader_writer/filebase.py |
File I/O abstraction for extracted assets | FileBasedDataWriter |
mineru/cli/fast_api.py |
HTTP API endpoint serving JSON output | /file_parse endpoint |
Summary
- MinerU generates a middle-JSON format that captures document structure, layout elements, and metadata during PDF processing.
- The conversion pipeline in
mineru/cli/common.pyhandles file intake, preprocessing with pypdfium2, and model inference before JSON construction. - The
result_to_middle_jsonfunction inmineru/backend/pipeline/model_json_to_middle_json.pyassembles the final JSON structure containingpdf_infowith per-page layout blocks, text, images, and equations. - You can convert PDF to JSON using MinerU via the CLI with
--return_middle_json Trueor programmatically using the Python API by settingf_dump_middle_json=Trueindo_parse. - For custom workflows, import
result_to_middle_jsondirectly to build JSON objects from existing model predictions.
Frequently Asked Questions
What is the middle-JSON format in MinerU?
The middle-JSON format is an intermediate representation generated during PDF processing that contains structured document information. It includes a pdf_info key with per-page arrays of layout blocks, text spans, bounding boxes, images, tables, and mathematical equations. This format serves as the foundation for downstream conversions to Markdown or other output formats.
How do I enable JSON output when using the MinerU CLI?
To generate JSON output via the command line, append the --return_middle_json True flag to your mineru parse command. By default, MinerU writes the middle-JSON file to the output directory using the naming pattern <pdf_name>_middle.json. You can also set f_dump_middle_json=True when using the Python API.
Can I convert PDF to JSON programmatically without using the CLI?
Yes, you can convert PDF to JSON programmatically by importing functions from mineru.cli.common. Use read_fn to load PDF bytes and do_parse with f_dump_middle_json=True to execute the pipeline. For advanced use cases, you can directly call result_to_middle_json from mineru.backend.pipeline.model_json_to_middle_json to build JSON objects from custom model outputs.
Where does MinerU handle image extraction during JSON conversion?
Image extraction occurs within the result_to_middle_json function in mineru/backend/pipeline/model_json_to_middle_json.py. This function receives an image writer instance (typically FileBasedDataWriter from mineru/data/data_reader_writer/filebase.py) that handles persisting extracted images to disk. The image paths are then referenced within the middle-JSON structure under the respective layout blocks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →