How to Convert PDF to Markdown Using MinerU: A Complete Technical Guide
MinerU converts PDFs and image-based documents into clean Markdown through a multi-stage pipeline that extracts page images, runs OCR or vision-language models, builds an intermediate "middle-JSON" representation, and finally renders Markdown output.
The open-source opendatalab/MinerU repository provides a robust toolkit to convert PDF to Markdown using MinerU through multiple interfaces: a command-line tool, a Python API, and a Gradio web UI. The system supports three distinct backends—traditional OCR pipelines, vision-language models (VLMs), and a hybrid approach—allowing users to balance speed and accuracy based on document complexity.
Understanding the MinerU PDF to Markdown Pipeline
When you convert PDF to Markdown using MinerU, the library executes a seven-stage pipeline that transforms raw bytes into structured text. Each stage is implemented in specific source modules within the repository:
-
File Ingestion – The input path is read as raw bytes via
read_fninmineru/cli/common.py, used by both the Gradio UI and demo scripts. -
PDF-to-Image Conversion – For backends requiring images, each page is rasterized using pypdfium2 through
pdf_to_imagesorpdf_to_images_b64strsinmineru/utils/pdf_reader.py. -
Backend Selection – The system selects between
pipeline,vlm, orhybridbackends based on the--backendflag or UI dropdown, resolved byget_vlm_engineinmineru/utils/engine_utils.py. -
Document Analysis –
- Pipeline:
pipeline_doc_analyzeinmineru/backend/pipeline/pipeline_analyze.pyextracts text blocks, tables, and formulas. - VLM:
vlm_doc_analyzeinmineru/backend/vlm/vlm_analyze.pysends images to the model. - Hybrid:
hybrid_doc_analyzeinmineru/backend/hybrid/hybrid_analyze.pymerges VLM output with OCR results.
- Pipeline:
-
Middle-JSON Construction – Raw outputs are normalized into a uniform
middle_jsonformat usingpipeline_result_to_middle_jsonorvlm_middle_json_mkcontent.union_make. -
Markdown Rendering – The middle-JSON is converted to Markdown via
pipeline_union_makeinmineru/backend/pipeline/pipeline_middle_json_mkcontent.pyorvlm_union_makeinmineru/backend/vlm/vlm_middle_json_mkcontent.py, handling image embedding, table formatting, and character escaping. -
Output Handling – Markdown files are written to the output directory alongside optional artifacts like layout PDFs, image directories, and the middle-JSON representation.
Conversion Methods: CLI, Python API, and Gradio UI
You can convert PDF to Markdown using MinerU through three primary interfaces, each suited to different workflows.
Command Line Interface
The CLI provides the fastest way to process single files or directories:
minerU -p mypaper.pdf -o ./out -b hybrid-auto-engine -l ch
Parameters explained:
-p: Input PDF or image path (or directory containing multiple files).-o: Destination folder for Markdown output, layout PDFs, and archives.-b: Backend selection (pipeline,vlm-auto-engine,hybrid-auto-engine).-l: Language hint for OCR (e.g.,chfor Chinese,enfor English).
After execution, the output directory contains:
mypaper.md: The final Markdown file.mypaper_layout.pdf: Visual layout preview.mypaper_images/: Extracted page images.mypaper_middle.json: Intermediate representation.mypaper.zip: Bundled archive of all artifacts.
Python API for Batch Processing
For programmatic workflows, import the parse_doc helper from the demo module:
from mineru.demo import parse_doc
from pathlib import Path
# Collect all PDFs in a directory
pdf_dir = Path("./pdfs")
pdf_paths = list(pdf_dir.glob("*.pdf"))
output = Path("./markdown_results")
# Process batch with hybrid backend
parse_doc(
path_list=pdf_paths,
output_dir=output,
backend="hybrid-auto-engine",
method="auto",
language="en"
)
print(f"Markdown files written to {output}")
The parse_doc function iterates over each path, calls do_parse (defined in demo/demo.py), and handles the backend-specific analysis and Markdown generation via _process_output.
Programmatic Gradio UI Access
You can also invoke the Gradio interface programmatically using the to_markdown async function:
import asyncio
from mineru.cli.gradio_app import to_markdown
async def convert_pdf():
md_content, raw_md, zip_path, layout_pdf = await to_markdown(
file_path="research_paper.pdf",
end_pages=5, # Process first 5 pages only
is_ocr=False, # Auto-detect text vs. OCR
formula_enable=True, # Enable formula detection
table_enable=True, # Enable table extraction
language="en",
backend="hybrid-auto-engine",
url=None # For remote VLM servers
)
print("Conversion complete!")
print(f"ZIP archive: {zip_path}")
print(f"Markdown preview:\n{md_content[:1000]}")
# Execute
asyncio.run(convert_pdf())
This approach leverages the same parse_pdf function used by the web interface, ensuring consistent output while allowing integration into automated workflows.
Backend Options: Pipeline, VLM, and Hybrid
When you convert PDF to Markdown using MinerU, you can choose from three analysis backends, each implemented in separate modules:
Pipeline Backend (mineru/backend/pipeline/)
- Uses traditional OCR and layout analysis models.
- Best for text-heavy documents with standard layouts.
- Key function:
pipeline_doc_analyzeinpipeline_analyze.py. - Markdown generation:
pipeline_union_makeinpipeline_middle_json_mkcontent.py.
VLM Backend (mineru/backend/vlm/)
- Employs vision-language models (local or via OpenAI-compatible APIs).
- Ideal for complex layouts, figures, and documents requiring semantic understanding.
- Key function:
vlm_doc_analyzeinvlm_analyze.py. - Engine resolution:
get_vlm_engineinmineru/utils/engine_utils.pymaps names likeqwen2orglm4to implementations. - Markdown generation:
vlm_union_makeinvlm_middle_json_mkcontent.py.
Hybrid Backend (mineru/backend/hybrid/)
- Combines VLM high-level understanding with OCR low-level text accuracy.
- Recommended for scanned documents or mixed content types.
- Key function:
hybrid_doc_analyzeinhybrid_analyze.py. - Merges outputs from both pipeline and VLM stages before middle-JSON construction.
Key Source Files and Functions
Understanding the codebase helps when customizing or debugging the conversion process:
| File | Key Function | Purpose |
|---|---|---|
mineru/cli/gradio_app.py |
to_markdown |
Gradio UI entry point; orchestrates PDF conversion, ZIP packaging, and Base64 image embedding. |
mineru/cli/common.py |
read_fn |
Loads raw bytes from PDF or image files. |
mineru/utils/pdf_reader.py |
pdf_to_images |
Renders PDF pages to PIL images using pypdfium2. |
mineru/utils/engine_utils.py |
get_vlm_engine |
Resolves backend engine names to concrete implementations. |
mineru/backend/pipeline/pipeline_analyze.py |
pipeline_doc_analyze |
OCR and layout analysis for text-based documents. |
mineru/backend/vlm/vlm_analyze.py |
vlm_doc_analyze |
Vision-language model inference for complex layouts. |
mineru/backend/hybrid/hybrid_analyze.py |
hybrid_doc_analyze |
Combines VLM and OCR results. |
mineru/backend/pipeline/pipeline_middle_json_mkcontent.py |
pipeline_union_make |
Renders Markdown from pipeline middle-JSON. |
mineru/backend/vlm/vlm_middle_json_mkcontent.py |
vlm_union_make |
Renders Markdown from VLM middle-JSON. |
demo/demo.py |
do_parse, _process_output |
Script-friendly entry points demonstrating the full conversion flow. |
Summary
To convert PDF to Markdown using MinerU effectively:
- Choose the right backend: Use
pipelinefor clean text documents,vlmfor complex layouts with figures, andhybridfor scanned or mixed-content PDFs. - Use the CLI for quick conversions: The
minerUcommand provides immediate results with flags for backend selection and language hints. - Leverage the Python API for automation: Import
parse_docfrommineru.demofor batch processing entire directories. - Understand the pipeline stages: The conversion flows from file ingestion → image rasterization → backend analysis → middle-JSON → Markdown rendering, with specific functions in
mineru/cli/gradio_app.pyandmineru/backend/*/union_makehandling the critical transformation steps.
Frequently Asked Questions
What is the difference between the Pipeline and VLM backends in MinerU?
The Pipeline backend uses traditional OCR engines and layout analysis models to extract text, tables, and formulas, making it faster and more accurate for clean, text-based PDFs. The VLM backend employs vision-language models (like Qwen2 or GLM4) to understand complex visual layouts, figures, and semantic structure, which is better for documents with heavy formatting, images, or handwritten content but requires more computational resources.
How do I handle scanned PDFs or image-based documents?
For scanned PDFs or image-based documents, use the Hybrid backend (-b hybrid-auto-engine in CLI or backend="hybrid-auto-engine" in Python). This backend combines the high-level layout understanding of VLMs with the low-level text accuracy of OCR. The hybrid_doc_analyze function in mineru/backend/hybrid/hybrid_analyze.py merges outputs from both approaches before constructing the middle-JSON representation.
Can I convert multiple PDFs in batch using MinerU?
Yes, batch conversion is supported through the Python API. Import parse_doc from mineru.demo and pass a list of Path objects to the path_list parameter. The function iterates over each file, calls do_parse (defined in demo/demo.py), and writes individual Markdown files to the specified output directory. This approach maintains consistent backend settings across the entire batch.
Where does the actual Markdown generation happen in the source code?
The final Markdown rendering occurs in the union_make functions within the backend-specific modules. For the Pipeline backend, pipeline_union_make in mineru/backend/pipeline/pipeline_middle_json_mkcontent.py walks the middle-JSON schema and writes Markdown text. For the VLM backend, vlm_union_make in mineru/backend/vlm/vlm_middle_json_mkcontent.py performs the same function for VLM-generated structures. Both functions handle image embedding, table formatting, and Markdown character escaping.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →