How to Convert PDF to Markdown Using MinerU: A Complete Technical Guide

MinerU converts PDFs and image-based documents into clean Markdown through a multi-stage pipeline that extracts page images, runs OCR or vision-language models, builds an intermediate "middle-JSON" representation, and finally renders Markdown output.

The open-source opendatalab/MinerU repository provides a robust toolkit to convert PDF to Markdown using MinerU through multiple interfaces: a command-line tool, a Python API, and a Gradio web UI. The system supports three distinct backends—traditional OCR pipelines, vision-language models (VLMs), and a hybrid approach—allowing users to balance speed and accuracy based on document complexity.

Understanding the MinerU PDF to Markdown Pipeline

When you convert PDF to Markdown using MinerU, the library executes a seven-stage pipeline that transforms raw bytes into structured text. Each stage is implemented in specific source modules within the repository:

  1. File Ingestion – The input path is read as raw bytes via read_fn in mineru/cli/common.py, used by both the Gradio UI and demo scripts.

  2. PDF-to-Image Conversion – For backends requiring images, each page is rasterized using pypdfium2 through pdf_to_images or pdf_to_images_b64strs in mineru/utils/pdf_reader.py.

  3. Backend Selection – The system selects between pipeline, vlm, or hybrid backends based on the --backend flag or UI dropdown, resolved by get_vlm_engine in mineru/utils/engine_utils.py.

  4. Document Analysis –

  5. Middle-JSON Construction – Raw outputs are normalized into a uniform middle_json format using pipeline_result_to_middle_json or vlm_middle_json_mkcontent.union_make.

  6. Markdown Rendering – The middle-JSON is converted to Markdown via pipeline_union_make in mineru/backend/pipeline/pipeline_middle_json_mkcontent.py or vlm_union_make in mineru/backend/vlm/vlm_middle_json_mkcontent.py, handling image embedding, table formatting, and character escaping.

  7. Output Handling – Markdown files are written to the output directory alongside optional artifacts like layout PDFs, image directories, and the middle-JSON representation.

Conversion Methods: CLI, Python API, and Gradio UI

You can convert PDF to Markdown using MinerU through three primary interfaces, each suited to different workflows.

Command Line Interface

The CLI provides the fastest way to process single files or directories:

minerU -p mypaper.pdf -o ./out -b hybrid-auto-engine -l ch

Parameters explained:

  • -p: Input PDF or image path (or directory containing multiple files).
  • -o: Destination folder for Markdown output, layout PDFs, and archives.
  • -b: Backend selection (pipeline, vlm-auto-engine, hybrid-auto-engine).
  • -l: Language hint for OCR (e.g., ch for Chinese, en for English).

After execution, the output directory contains:

  • mypaper.md: The final Markdown file.
  • mypaper_layout.pdf: Visual layout preview.
  • mypaper_images/: Extracted page images.
  • mypaper_middle.json: Intermediate representation.
  • mypaper.zip: Bundled archive of all artifacts.

Python API for Batch Processing

For programmatic workflows, import the parse_doc helper from the demo module:

from mineru.demo import parse_doc
from pathlib import Path

# Collect all PDFs in a directory

pdf_dir = Path("./pdfs")
pdf_paths = list(pdf_dir.glob("*.pdf"))
output = Path("./markdown_results")

# Process batch with hybrid backend

parse_doc(
    path_list=pdf_paths,
    output_dir=output,
    backend="hybrid-auto-engine",
    method="auto",
    language="en"
)

print(f"Markdown files written to {output}")

The parse_doc function iterates over each path, calls do_parse (defined in demo/demo.py), and handles the backend-specific analysis and Markdown generation via _process_output.

Programmatic Gradio UI Access

You can also invoke the Gradio interface programmatically using the to_markdown async function:

import asyncio
from mineru.cli.gradio_app import to_markdown

async def convert_pdf():
    md_content, raw_md, zip_path, layout_pdf = await to_markdown(
        file_path="research_paper.pdf",
        end_pages=5,              # Process first 5 pages only

        is_ocr=False,             # Auto-detect text vs. OCR

        formula_enable=True,      # Enable formula detection

        table_enable=True,        # Enable table extraction

        language="en",
        backend="hybrid-auto-engine",
        url=None                  # For remote VLM servers

    )
    
    print("Conversion complete!")
    print(f"ZIP archive: {zip_path}")
    print(f"Markdown preview:\n{md_content[:1000]}")

# Execute

asyncio.run(convert_pdf())

This approach leverages the same parse_pdf function used by the web interface, ensuring consistent output while allowing integration into automated workflows.

Backend Options: Pipeline, VLM, and Hybrid

When you convert PDF to Markdown using MinerU, you can choose from three analysis backends, each implemented in separate modules:

Pipeline Backend (mineru/backend/pipeline/)

VLM Backend (mineru/backend/vlm/)

  • Employs vision-language models (local or via OpenAI-compatible APIs).
  • Ideal for complex layouts, figures, and documents requiring semantic understanding.
  • Key function: vlm_doc_analyze in vlm_analyze.py.
  • Engine resolution: get_vlm_engine in mineru/utils/engine_utils.py maps names like qwen2 or glm4 to implementations.
  • Markdown generation: vlm_union_make in vlm_middle_json_mkcontent.py.

Hybrid Backend (mineru/backend/hybrid/)

  • Combines VLM high-level understanding with OCR low-level text accuracy.
  • Recommended for scanned documents or mixed content types.
  • Key function: hybrid_doc_analyze in hybrid_analyze.py.
  • Merges outputs from both pipeline and VLM stages before middle-JSON construction.

Key Source Files and Functions

Understanding the codebase helps when customizing or debugging the conversion process:

File Key Function Purpose
mineru/cli/gradio_app.py to_markdown Gradio UI entry point; orchestrates PDF conversion, ZIP packaging, and Base64 image embedding.
mineru/cli/common.py read_fn Loads raw bytes from PDF or image files.
mineru/utils/pdf_reader.py pdf_to_images Renders PDF pages to PIL images using pypdfium2.
mineru/utils/engine_utils.py get_vlm_engine Resolves backend engine names to concrete implementations.
mineru/backend/pipeline/pipeline_analyze.py pipeline_doc_analyze OCR and layout analysis for text-based documents.
mineru/backend/vlm/vlm_analyze.py vlm_doc_analyze Vision-language model inference for complex layouts.
mineru/backend/hybrid/hybrid_analyze.py hybrid_doc_analyze Combines VLM and OCR results.
mineru/backend/pipeline/pipeline_middle_json_mkcontent.py pipeline_union_make Renders Markdown from pipeline middle-JSON.
mineru/backend/vlm/vlm_middle_json_mkcontent.py vlm_union_make Renders Markdown from VLM middle-JSON.
demo/demo.py do_parse, _process_output Script-friendly entry points demonstrating the full conversion flow.

Summary

To convert PDF to Markdown using MinerU effectively:

  • Choose the right backend: Use pipeline for clean text documents, vlm for complex layouts with figures, and hybrid for scanned or mixed-content PDFs.
  • Use the CLI for quick conversions: The minerU command provides immediate results with flags for backend selection and language hints.
  • Leverage the Python API for automation: Import parse_doc from mineru.demo for batch processing entire directories.
  • Understand the pipeline stages: The conversion flows from file ingestion → image rasterization → backend analysis → middle-JSON → Markdown rendering, with specific functions in mineru/cli/gradio_app.py and mineru/backend/*/union_make handling the critical transformation steps.

Frequently Asked Questions

What is the difference between the Pipeline and VLM backends in MinerU?

The Pipeline backend uses traditional OCR engines and layout analysis models to extract text, tables, and formulas, making it faster and more accurate for clean, text-based PDFs. The VLM backend employs vision-language models (like Qwen2 or GLM4) to understand complex visual layouts, figures, and semantic structure, which is better for documents with heavy formatting, images, or handwritten content but requires more computational resources.

How do I handle scanned PDFs or image-based documents?

For scanned PDFs or image-based documents, use the Hybrid backend (-b hybrid-auto-engine in CLI or backend="hybrid-auto-engine" in Python). This backend combines the high-level layout understanding of VLMs with the low-level text accuracy of OCR. The hybrid_doc_analyze function in mineru/backend/hybrid/hybrid_analyze.py merges outputs from both approaches before constructing the middle-JSON representation.

Can I convert multiple PDFs in batch using MinerU?

Yes, batch conversion is supported through the Python API. Import parse_doc from mineru.demo and pass a list of Path objects to the path_list parameter. The function iterates over each file, calls do_parse (defined in demo/demo.py), and writes individual Markdown files to the specified output directory. This approach maintains consistent backend settings across the entire batch.

Where does the actual Markdown generation happen in the source code?

The final Markdown rendering occurs in the union_make functions within the backend-specific modules. For the Pipeline backend, pipeline_union_make in mineru/backend/pipeline/pipeline_middle_json_mkcontent.py walks the middle-JSON schema and writes Markdown text. For the VLM backend, vlm_union_make in mineru/backend/vlm/vlm_middle_json_mkcontent.py performs the same function for VLM-generated structures. Both functions handle image embedding, table formatting, and Markdown character escaping.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →