How to Utilize MinerU's VLM Backend for High-Accuracy Parsing
To achieve high-accuracy PDF parsing with MinerU, configure the Vision-Language Model (VLM) backend using environment variables for formula and table extraction, select the appropriate inference engine (vLLM, transformers, or LMDeploy), and leverage the MagicModel post-processor to convert raw spans into structured Content-List V2 output.
The MinerU VLM backend serves as the core engine in the opendatalab/MinerU repository that transforms PDF pages into richly-structured middle-JSON representations. By understanding the three-layer pipeline—model initialization, inference, and post-processing—you can optimize the system for maximum extraction fidelity across complex documents containing equations, tables, and multilingual text.
Understanding the VLM Backend Architecture
The parsing pipeline consists of three distinct logical layers that process input from raw bytes to structured output.
Model Initialization Layer
The entry point for any VLM operation begins in minerU/backend/vlm/vlm_analyze.py through the ModelSingleton.get_model method. This singleton pattern guarantees that the same backend and model combination is instantiated only once per process, eliminating redundant heavy model loads.
MinerU supports five distinct backend engines:
- transformers: Loads Qwen2-VL via
Qwen2VLForConditionalGeneration.from_pretrained, automatically selecting the correct dtype key based on your transformers version (lines 69-73 invlm_analyze.py) - vllm-engine: Creates a synchronous
vllm.LLMobject - vllm-async-engine: Creates an asynchronous
vllm.AsyncLLMinstance for concurrent requests - lmdeploy-engine: Initializes a
VLAsyncEnginewith automatic backend selection (PyTorch or Turbomind) viaset_lmdeploy_backend(lines 59-78) - mlx-engine: Exclusive to macOS 13.5+ Apple Silicon, guarded by
is_mac_os_version_supported
The initialization layer in minerU/backend/vlm/utils.py provides critical helper functions including set_default_gpu_memory_utilization, which returns 0.7 for low-VRAM cards (<8 GiB) and 0.5 otherwise (lines 82-89), and mod_kwargs_by_device_type for hardware-specific compilation configuration.
Inference Layer
Once initialized, the system processes page-wise image lists through doc_analyze or aio_doc_analyze in vlm_analyze.py. These functions encapsulate calls to MinerUClient.batch_two_step_extract, which feeds images to the selected VLM and returns raw spans containing text, inline equations, display equations, images, and tables with bounding box coordinates.
If CUDA is present and VLLM_USE_V1=1 is set (default), the system automatically injects MinerULogitsProcessor via enable_custom_logits_processors to improve token-level decisions for Chinese, Japanese, Korean, and hyphenated English text.
Post-Processing Layer
Raw VLM outputs undergo sophisticated transformation through two key components:
-
MagicModel (
minerU/backend/vlm/vlm_magic_model.py): TheMagicModelclass (lines 13-183) normalizes coordinates, splits inline formulas, fixes hyphenation, and groups spans into hierarchical lines and blocks. It handles complex layout cases including tables with misplaced captions, nested lists, and algorithm-style code throughfix_two_layer_blocksandfix_list_blocks. -
Middle-JSON Creation (
minerU/backend/vlm/vlm_middle_json_mkcontent.py): Theunion_makefunction orchestrates conversion to your target format, selecting fromMakeMode.MM_MD,MakeMode.NLP_MD,MakeMode.CONTENT_LIST, orMakeMode.CONTENT_LIST_V2. This process stitches processed blocks, adds page indices, normalizes bounding boxes to a 0-1000 scale, and optionally injects original image paths.
Configuration for Maximum Accuracy
Achieving high-accuracy parsing requires specific environment configurations that control how the VLM handles complex document elements.
Enable Formula and Table Recognition
Set these environment variables before initialization to preserve semantic structure:
export MINERU_VLM_FORMULA_ENABLE=true # Keeps LaTeX as display equations
export MINERU_VLM_TABLE_ENABLE=true # Emits HTML tables when recognized
When MINERU_VLM_FORMULA_ENABLE is true, INTERLINE_EQUATION spans remain as LaTeX rather than converting to images. Similarly, MINERU_VLM_TABLE_ENABLE allows mk_blocks_to_markdown (lines 41-62 in vlm_middle_json_mkcontent.py) to emit HTML <table> tags instead of image fallbacks.
GPU Memory and Performance Tuning
Control resource allocation to prevent out-of-memory crashes while maintaining throughput:
# Override auto-detection (0.7 for <8GB, 0.5 otherwise)
export VLLM_GPU_MEMORY_UTILIZATION=0.7
# Enable custom logits processing for CJK and hyphenation
export VLLM_USE_V1=1
For specialized hardware, export device-specific configuration:
export MINERU_VLLM_DEVICE=kxpu # or corex
export MINERU_VLLM_DEVICE_CONFIG='{"block_size":128,"dtype":"float16"}'
The mod_kwargs_by_device_type function (lines 71-89 in utils.py) reads these variables to inject appropriate --block-size and --dtype arguments into the vLLM compilation pipeline.
Implementing the Parsing Pipeline
Starting a VLM Server
For production deployments, launch a dedicated inference server using the built-in CLI:
# vLLM backend (default)
mineru vlm_server --engine vllm
# LMDeploy backend for alternative acceleration
mineru vlm_server --engine lmdeploy
The CLI entry point in minerU/cli/vlm_server.py forwards to minerU/model/vlm/vllm_server.py or minerU/model/vlm/lmdeploy_server.py, automatically injecting GPU memory and logits settings.
Synchronous Python API
For batch processing or simple scripts, use the high-level doc_analyze function:
import os
from mineru.backend.vlm.vlm_analyze import doc_analyze
# Configure for maximum accuracy
os.environ["MINERU_VLM_FORMULA_ENABLE"] = "true"
os.environ["MINERU_VLM_TABLE_ENABLE"] = "true"
os.environ["VLLM_USE_V1"] = "1"
# Load PDF bytes
with open("research_paper.pdf", "rb") as f:
pdf_bytes = f.read()
# Execute parsing with vLLM backend
middle_json, raw_results = doc_analyze(
pdf_bytes,
image_writer=None, # Optional: DataWriter for extracted images
backend="vllm-engine", # Options: transformers, vllm-engine, lmdeploy-engine
model_path="/models/qwen2vl", # Optional: auto-downloads if omitted
gpu_memory_utilization="0.7", # Override auto-detected value
)
# Access structured content
print(middle_json[0]["blocks"][0]["type"]) # e.g., "TITLE"
print(middle_json[0]["blocks"][0]["content"]) # Extracted text or LaTeX
The ModelSingleton caches the vLLM LLM object, making subsequent calls instant. The function returns Content-List V2 format by default through union_make.
Asynchronous Python API
For web services or high-concurrency applications, use the async interface:
import asyncio
from mineru.backend.vlm.vlm_analyze import aio_doc_analyze
async def parse_document(pdf_path: str):
with open(pdf_path, "rb") as f:
pdf_bytes = f.read()
middle_json, _ = await aio_doc_analyze(
pdf_bytes,
backend="vllm-async-engine",
model_path="/models/qwen2vl",
)
return middle_json
# Execute
result = asyncio.run(parse_document("document.pdf"))
This path creates a vllm.AsyncLLM instance (lines 24-55 in vlm_analyze.py) capable of serving concurrent requests while applying the same high-accuracy post-processing pipeline.
Device-Specific Compilation
For non-standard GPUs like Kunlun XPU, configure compilation parameters:
export MINERU_VLLM_DEVICE=kxpu
export MINERU_VLLM_DEVICE_CONFIG='{"block_size":128,"dtype":"float16"}'
mineru vlm_server --engine vllm
The mod_kwargs_by_device_type function aligns the vLLM compilation pipeline with hardware-specific block sizes, optimizing both speed and accuracy for specialized accelerators.
Summary
- MinerU's VLM backend uses a three-layer architecture: model initialization (
ModelSingleton), inference (doc_analyze/aio_doc_analyze), and post-processing (MagicModelandunion_make). - High-accuracy configuration requires setting
MINERU_VLM_FORMULA_ENABLE=trueandMINERU_VLM_TABLE_ENABLE=trueto preserve LaTeX and HTML table structures. - GPU optimization relies on
set_default_gpu_memory_utilization(0.7 for <8GB VRAM, 0.5 otherwise) and optionalMINERU_VLLM_DEVICEconfiguration for custom hardware. - Backend options include
transformers,vllm-engine,vllm-async-engine,lmdeploy-engine, andmlx-engine, each loaded viaminerU/backend/vlm/vlm_analyze.py. - Output formats are controlled through
MakeMode, withCONTENT_LIST_V2providing the most structured representation for downstream applications.
Frequently Asked Questions
How do I choose between the vLLM and transformers backend?
Use the vLLM backend (vllm-engine or vllm-async-engine) for production workloads requiring high throughput and concurrent request handling, as implemented in minerU/backend/vlm/vlm_analyze.py lines 24-55. The transformers backend is ideal for quick prototyping or environments where vLLM dependencies cannot be installed, loading Qwen2-VL directly via Qwen2VLForConditionalGeneration.from_pretrained.
Why are my mathematical formulas being extracted as images instead of LaTeX?
Ensure you have set the environment variable MINERU_VLM_FORMULA_ENABLE=true before initializing the model. When this flag is disabled, the union_make function in minerU/backend/vlm/vlm_middle_json_mkcontent.py treats INTERLINE_EQUATION spans as images rather than preserving the LaTeX source code (lines 14-15).
Can I run MinerU's VLM backend on Apple Silicon Macs?
Yes, but only on macOS 13.5 or later using the mlx-engine backend. The code explicitly checks is_mac_os_version_supported before allowing MLX initialization. Note that this engine is not available for Linux or Windows environments, where you should use vllm-engine or transformers instead.
What causes out-of-memory errors during PDF parsing and how do I fix them?
OOM errors typically occur when GPU memory utilization exceeds available VRAM. The system auto-detects card capacity in minerU/backend/vlm/utils.py (lines 82-89), setting utilization to 0.7 for cards with less than 8 GiB and 0.5 otherwise. You can manually override this by passing gpu_memory_utilization="0.6" to doc_analyze or setting the VLLM_GPU_MEMORY_UTILIZATION environment variable.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →