Complete Guide to Output Formats Supported by MinerU

MinerU supports seven distinct output formats including Multimodal Markdown, NLP-focused Markdown, reading-order-sorted JSON, intermediate layout JSON, raw model output, extracted images, and ZIP archives.

MinerU is an open-source document parsing toolkit developed by OpenDataLab that converts PDFs into structured, machine-readable representations. Understanding the output formats supported by MinerU is essential for selecting the right export type for your workflow, whether you need human-readable documentation or structured data for NLP pipelines.

Overview of MinerU Export Formats

MinerU provides flexible export options through three primary interfaces: the FastAPI server (mineru/cli/fast_api.py), the command-line interface (mineru/cli/common.py), and direct Python API calls using the MakeMode enum. Each format serves specific use cases ranging from visual preservation to data extraction.

Detailed Breakdown of Each Output Format

Multimodal Markdown (mm_markdown)

The Multimodal Markdown format (MakeMode.MM_MD) produces full-featured Markdown that preserves the original document layout. This format embeds images as base64-encoded strings or URLs, renders mathematical formulas as LaTeX, and includes tables in HTML format. It is ideal for human-readable reports and documentation.

To request this format via FastAPI, use --return_md True or programmatically set f_make_md_mode=MakeMode.MM_MD.

NLP-Focused Markdown (nlp_markdown)

The NLP-Focused Markdown format (MakeMode.NLP_MD) generates cleaner Markdown optimized for downstream language model processing. This variant removes visual clutter and focuses on textual content and simple tables, making it suitable for RAG (Retrieval-Augmented Generation) pipelines and text analysis workflows.

Set f_make_md_mode=MakeMode.NLP_MD in your Python code to enable this format.

Reading-Order-Sorted JSON (content_list)

The Reading-Order-Sorted JSON format (content_list or content_list_v2) exports structured JSON where each element represents a logical document block—such as paragraphs, tables, images, or equations—sorted by visual reading order. The v2 variant adds richer type annotations and metadata.

Request this via CLI with --return-content-list or via FastAPI with --return_content_list True. The underlying implementation resides in mineru/backend/pipeline/pipeline_middle_json_mkcontent.py and mineru/backend/vlm/vlm_middle_json_mkcontent.py.

Intermediate Middle JSON

The Intermediate Middle JSON format provides detailed JSON produced after layout analysis but before final Markdown conversion. This output contains bounding boxes, span information, OCR results, and layout metadata—useful for debugging, custom post-processing, or extracting specific geometric data.

Enable this output using --return_middle_json True in the CLI or FastAPI.

Model-Output JSON

The Model-Output JSON format exposes the raw output from the underlying vision-language model, including token scores, logits, and intermediate representations. This format is primarily intended for research, model debugging, or troubleshooting parsing behavior.

Request this via --return_model_output True.

Extracted Images

The Extracted Images format saves all raster images detected during parsing as individual JPEG files. This is useful when you need to process or analyze document images separately from the text content.

Enable this with --return_images True.

ZIP Archive

The ZIP Archive format packages any combination of the above outputs into a single downloadable .zip file. This is convenient for batch processing or when you need to transfer multiple output types together.

Use --response_format_zip True in FastAPI to receive a ZIP archive.

How to Request Specific Output Formats

Using the FastAPI Endpoint

The FastAPI server defined in mineru/cli/fast_api.py accepts boolean flags for each output type. For example, to request both Markdown and content list in a ZIP archive:

curl -X POST "http://localhost:8000/parse" \
  -F "files=@sample.pdf" \
  -F "return_md=true" \
  -F "return_content_list=true" \
  -F "response_format_zip=true"

The response returns a ZIP file containing sample.md and sample_content_list.json. The flag definitions are located at lines 177-186 in mineru/cli/fast_api.py.

Using the Command-Line Interface

The CLI implementation in mineru/cli/common.py supports similar flags. To generate Markdown and JSON outputs:

mineru \
  --file sample.pdf \
  --backend pipeline \
  --return-md \
  --return-content-list \
  --output-dir ./out

Results are written to ./out/<uuid>/sample/ as separate files. The argument parsing occurs at lines 143-150 in mineru/cli/common.py.

Programmatic Configuration with MakeMode

For Python applications, the MakeMode enum in mineru/utils/enum_class.py provides explicit control over output generation. The vlm_union_make function accepts a MakeMode parameter:

from mineru.utils.enum_class import MakeMode

# Generate multimodal markdown

md_output = vlm_union_make(pdf_info, MakeMode.MM_MD, image_dir)

# Generate content list JSON

json_output = vlm_union_make(pdf_info, MakeMode.CONTENT_LIST, image_dir)

Available modes include MM_MD, NLP_MD, CONTENT_LIST, and CONTENT_LIST_V2, as defined in mineru/utils/enum_class.py at lines 86-90.

Key Source Files and Implementation Details

Understanding the codebase helps when customizing outputs:

Summary

  • MinerU supports seven primary output formats: Multimodal Markdown, NLP Markdown, reading-order JSON, intermediate JSON, model output JSON, extracted images, and ZIP archives.
  • Three interfaces control output: FastAPI endpoint (mineru/cli/fast_api.py), CLI (mineru/cli/common.py), and Python API using MakeMode enum.
  • Markdown variants serve different purposes: MM_MD preserves visual layout with images and LaTeX, while NLP_MD optimizes text for language model processing.
  • JSON outputs provide structured data: content_list delivers reading-order sorted blocks, while middle_json offers pre-conversion layout analysis data.

Frequently Asked Questions

What is the difference between Multimodal Markdown and NLP Markdown in MinerU?

Multimodal Markdown (MakeMode.MM_MD) preserves the full visual context of the original document, embedding images as base64 or URLs, rendering equations as LaTeX, and maintaining complex table layouts in HTML. NLP Markdown (MakeMode.NLP_MD) strips visual elements to produce clean, text-focused Markdown optimized for downstream natural language processing tasks and RAG pipelines.

How do I extract structured JSON instead of Markdown from MinerU?

Request the content_list format using --return-content-list in the CLI, --return_content_list True in the FastAPI call, or MakeMode.CONTENT_LIST (or CONTENT_LIST_V2 for enhanced annotations) in Python. This outputs a JSON array where each object represents a logical document block sorted by visual reading order.

Can MinerU output raw layout analysis data for custom processing?

Yes, enable the Intermediate Middle JSON format using --return_middle_json True. This outputs the detailed JSON representation produced after layout analysis but before Markdown conversion, containing bounding boxes, span-level OCR results, and geometric metadata. This format is processed by mineru/backend/pipeline/pipeline_middle_json_mkcontent.py and mineru/backend/vlm/vlm_middle_json_mkcontent.py.

Is it possible to get all output formats in a single request?

Yes, use the ZIP Archive option by setting --response_format_zip True in FastAPI or combining multiple --return-* flags in the CLI. This packages all requested formats into a single downloadable .zip file, making it convenient for batch processing workflows that require both human-readable Markdown and structured JSON data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →