Complete Guide to Output Formats Supported by MinerU
MinerU supports seven distinct output formats including Multimodal Markdown, NLP-focused Markdown, reading-order-sorted JSON, intermediate layout JSON, raw model output, extracted images, and ZIP archives.
MinerU is an open-source document parsing toolkit developed by OpenDataLab that converts PDFs into structured, machine-readable representations. Understanding the output formats supported by MinerU is essential for selecting the right export type for your workflow, whether you need human-readable documentation or structured data for NLP pipelines.
Overview of MinerU Export Formats
MinerU provides flexible export options through three primary interfaces: the FastAPI server (mineru/cli/fast_api.py), the command-line interface (mineru/cli/common.py), and direct Python API calls using the MakeMode enum. Each format serves specific use cases ranging from visual preservation to data extraction.
Detailed Breakdown of Each Output Format
Multimodal Markdown (mm_markdown)
The Multimodal Markdown format (MakeMode.MM_MD) produces full-featured Markdown that preserves the original document layout. This format embeds images as base64-encoded strings or URLs, renders mathematical formulas as LaTeX, and includes tables in HTML format. It is ideal for human-readable reports and documentation.
To request this format via FastAPI, use --return_md True or programmatically set f_make_md_mode=MakeMode.MM_MD.
NLP-Focused Markdown (nlp_markdown)
The NLP-Focused Markdown format (MakeMode.NLP_MD) generates cleaner Markdown optimized for downstream language model processing. This variant removes visual clutter and focuses on textual content and simple tables, making it suitable for RAG (Retrieval-Augmented Generation) pipelines and text analysis workflows.
Set f_make_md_mode=MakeMode.NLP_MD in your Python code to enable this format.
Reading-Order-Sorted JSON (content_list)
The Reading-Order-Sorted JSON format (content_list or content_list_v2) exports structured JSON where each element represents a logical document block—such as paragraphs, tables, images, or equations—sorted by visual reading order. The v2 variant adds richer type annotations and metadata.
Request this via CLI with --return-content-list or via FastAPI with --return_content_list True. The underlying implementation resides in mineru/backend/pipeline/pipeline_middle_json_mkcontent.py and mineru/backend/vlm/vlm_middle_json_mkcontent.py.
Intermediate Middle JSON
The Intermediate Middle JSON format provides detailed JSON produced after layout analysis but before final Markdown conversion. This output contains bounding boxes, span information, OCR results, and layout metadata—useful for debugging, custom post-processing, or extracting specific geometric data.
Enable this output using --return_middle_json True in the CLI or FastAPI.
Model-Output JSON
The Model-Output JSON format exposes the raw output from the underlying vision-language model, including token scores, logits, and intermediate representations. This format is primarily intended for research, model debugging, or troubleshooting parsing behavior.
Request this via --return_model_output True.
Extracted Images
The Extracted Images format saves all raster images detected during parsing as individual JPEG files. This is useful when you need to process or analyze document images separately from the text content.
Enable this with --return_images True.
ZIP Archive
The ZIP Archive format packages any combination of the above outputs into a single downloadable .zip file. This is convenient for batch processing or when you need to transfer multiple output types together.
Use --response_format_zip True in FastAPI to receive a ZIP archive.
How to Request Specific Output Formats
Using the FastAPI Endpoint
The FastAPI server defined in mineru/cli/fast_api.py accepts boolean flags for each output type. For example, to request both Markdown and content list in a ZIP archive:
curl -X POST "http://localhost:8000/parse" \
-F "files=@sample.pdf" \
-F "return_md=true" \
-F "return_content_list=true" \
-F "response_format_zip=true"
The response returns a ZIP file containing sample.md and sample_content_list.json. The flag definitions are located at lines 177-186 in mineru/cli/fast_api.py.
Using the Command-Line Interface
The CLI implementation in mineru/cli/common.py supports similar flags. To generate Markdown and JSON outputs:
mineru \
--file sample.pdf \
--backend pipeline \
--return-md \
--return-content-list \
--output-dir ./out
Results are written to ./out/<uuid>/sample/ as separate files. The argument parsing occurs at lines 143-150 in mineru/cli/common.py.
Programmatic Configuration with MakeMode
For Python applications, the MakeMode enum in mineru/utils/enum_class.py provides explicit control over output generation. The vlm_union_make function accepts a MakeMode parameter:
from mineru.utils.enum_class import MakeMode
# Generate multimodal markdown
md_output = vlm_union_make(pdf_info, MakeMode.MM_MD, image_dir)
# Generate content list JSON
json_output = vlm_union_make(pdf_info, MakeMode.CONTENT_LIST, image_dir)
Available modes include MM_MD, NLP_MD, CONTENT_LIST, and CONTENT_LIST_V2, as defined in mineru/utils/enum_class.py at lines 86-90.
Key Source Files and Implementation Details
Understanding the codebase helps when customizing outputs:
mineru/utils/enum_class.py: Defines theMakeModeenum that controls markdown generation strategies.mineru/cli/fast_api.py: Implements the FastAPI endpoint withreturn_*boolean flags (lines 177-186).mineru/cli/common.py: Handles CLI argument parsing for output format selection (lines 143-150).mineru/backend/pipeline/pipeline_middle_json_mkcontent.py: Converts middle JSON to final output in pipeline mode.mineru/backend/vlm/vlm_middle_json_mkcontent.py: Handles content generation for VLM backend.
Summary
- MinerU supports seven primary output formats: Multimodal Markdown, NLP Markdown, reading-order JSON, intermediate JSON, model output JSON, extracted images, and ZIP archives.
- Three interfaces control output: FastAPI endpoint (
mineru/cli/fast_api.py), CLI (mineru/cli/common.py), and Python API usingMakeModeenum. - Markdown variants serve different purposes:
MM_MDpreserves visual layout with images and LaTeX, whileNLP_MDoptimizes text for language model processing. - JSON outputs provide structured data:
content_listdelivers reading-order sorted blocks, whilemiddle_jsonoffers pre-conversion layout analysis data.
Frequently Asked Questions
What is the difference between Multimodal Markdown and NLP Markdown in MinerU?
Multimodal Markdown (MakeMode.MM_MD) preserves the full visual context of the original document, embedding images as base64 or URLs, rendering equations as LaTeX, and maintaining complex table layouts in HTML. NLP Markdown (MakeMode.NLP_MD) strips visual elements to produce clean, text-focused Markdown optimized for downstream natural language processing tasks and RAG pipelines.
How do I extract structured JSON instead of Markdown from MinerU?
Request the content_list format using --return-content-list in the CLI, --return_content_list True in the FastAPI call, or MakeMode.CONTENT_LIST (or CONTENT_LIST_V2 for enhanced annotations) in Python. This outputs a JSON array where each object represents a logical document block sorted by visual reading order.
Can MinerU output raw layout analysis data for custom processing?
Yes, enable the Intermediate Middle JSON format using --return_middle_json True. This outputs the detailed JSON representation produced after layout analysis but before Markdown conversion, containing bounding boxes, span-level OCR results, and geometric metadata. This format is processed by mineru/backend/pipeline/pipeline_middle_json_mkcontent.py and mineru/backend/vlm/vlm_middle_json_mkcontent.py.
Is it possible to get all output formats in a single request?
Yes, use the ZIP Archive option by setting --response_format_zip True in FastAPI or combining multiple --return-* flags in the CLI. This packages all requested formats into a single downloadable .zip file, making it convenient for batch processing workflows that require both human-readable Markdown and structured JSON data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →