How to Utilize MinerU's VLM Backend for High-Accuracy Parsing

To achieve high-accuracy PDF parsing with MinerU, configure the Vision-Language Model (VLM) backend using environment variables for formula and table extraction, select the appropriate inference engine (vLLM, transformers, or LMDeploy), and leverage the MagicModel post-processor to convert raw spans into structured Content-List V2 output.

The MinerU VLM backend serves as the core engine in the opendatalab/MinerU repository that transforms PDF pages into richly-structured middle-JSON representations. By understanding the three-layer pipeline—model initialization, inference, and post-processing—you can optimize the system for maximum extraction fidelity across complex documents containing equations, tables, and multilingual text.

Understanding the VLM Backend Architecture

The parsing pipeline consists of three distinct logical layers that process input from raw bytes to structured output.

Model Initialization Layer

The entry point for any VLM operation begins in minerU/backend/vlm/vlm_analyze.py through the ModelSingleton.get_model method. This singleton pattern guarantees that the same backend and model combination is instantiated only once per process, eliminating redundant heavy model loads.

MinerU supports five distinct backend engines:

  • transformers: Loads Qwen2-VL via Qwen2VLForConditionalGeneration.from_pretrained, automatically selecting the correct dtype key based on your transformers version (lines 69-73 in vlm_analyze.py)
  • vllm-engine: Creates a synchronous vllm.LLM object
  • vllm-async-engine: Creates an asynchronous vllm.AsyncLLM instance for concurrent requests
  • lmdeploy-engine: Initializes a VLAsyncEngine with automatic backend selection (PyTorch or Turbomind) via set_lmdeploy_backend (lines 59-78)
  • mlx-engine: Exclusive to macOS 13.5+ Apple Silicon, guarded by is_mac_os_version_supported

The initialization layer in minerU/backend/vlm/utils.py provides critical helper functions including set_default_gpu_memory_utilization, which returns 0.7 for low-VRAM cards (<8 GiB) and 0.5 otherwise (lines 82-89), and mod_kwargs_by_device_type for hardware-specific compilation configuration.

Inference Layer

Once initialized, the system processes page-wise image lists through doc_analyze or aio_doc_analyze in vlm_analyze.py. These functions encapsulate calls to MinerUClient.batch_two_step_extract, which feeds images to the selected VLM and returns raw spans containing text, inline equations, display equations, images, and tables with bounding box coordinates.

If CUDA is present and VLLM_USE_V1=1 is set (default), the system automatically injects MinerULogitsProcessor via enable_custom_logits_processors to improve token-level decisions for Chinese, Japanese, Korean, and hyphenated English text.

Post-Processing Layer

Raw VLM outputs undergo sophisticated transformation through two key components:

  1. MagicModel (minerU/backend/vlm/vlm_magic_model.py): The MagicModel class (lines 13-183) normalizes coordinates, splits inline formulas, fixes hyphenation, and groups spans into hierarchical lines and blocks. It handles complex layout cases including tables with misplaced captions, nested lists, and algorithm-style code through fix_two_layer_blocks and fix_list_blocks.

  2. Middle-JSON Creation (minerU/backend/vlm/vlm_middle_json_mkcontent.py): The union_make function orchestrates conversion to your target format, selecting from MakeMode.MM_MD, MakeMode.NLP_MD, MakeMode.CONTENT_LIST, or MakeMode.CONTENT_LIST_V2. This process stitches processed blocks, adds page indices, normalizes bounding boxes to a 0-1000 scale, and optionally injects original image paths.

Configuration for Maximum Accuracy

Achieving high-accuracy parsing requires specific environment configurations that control how the VLM handles complex document elements.

Enable Formula and Table Recognition

Set these environment variables before initialization to preserve semantic structure:

export MINERU_VLM_FORMULA_ENABLE=true   # Keeps LaTeX as display equations

export MINERU_VLM_TABLE_ENABLE=true     # Emits HTML tables when recognized

When MINERU_VLM_FORMULA_ENABLE is true, INTERLINE_EQUATION spans remain as LaTeX rather than converting to images. Similarly, MINERU_VLM_TABLE_ENABLE allows mk_blocks_to_markdown (lines 41-62 in vlm_middle_json_mkcontent.py) to emit HTML <table> tags instead of image fallbacks.

GPU Memory and Performance Tuning

Control resource allocation to prevent out-of-memory crashes while maintaining throughput:


# Override auto-detection (0.7 for <8GB, 0.5 otherwise)

export VLLM_GPU_MEMORY_UTILIZATION=0.7

# Enable custom logits processing for CJK and hyphenation

export VLLM_USE_V1=1

For specialized hardware, export device-specific configuration:

export MINERU_VLLM_DEVICE=kxpu  # or corex

export MINERU_VLLM_DEVICE_CONFIG='{"block_size":128,"dtype":"float16"}'

The mod_kwargs_by_device_type function (lines 71-89 in utils.py) reads these variables to inject appropriate --block-size and --dtype arguments into the vLLM compilation pipeline.

Implementing the Parsing Pipeline

Starting a VLM Server

For production deployments, launch a dedicated inference server using the built-in CLI:


# vLLM backend (default)

mineru vlm_server --engine vllm

# LMDeploy backend for alternative acceleration

mineru vlm_server --engine lmdeploy

The CLI entry point in minerU/cli/vlm_server.py forwards to minerU/model/vlm/vllm_server.py or minerU/model/vlm/lmdeploy_server.py, automatically injecting GPU memory and logits settings.

Synchronous Python API

For batch processing or simple scripts, use the high-level doc_analyze function:

import os
from mineru.backend.vlm.vlm_analyze import doc_analyze

# Configure for maximum accuracy

os.environ["MINERU_VLM_FORMULA_ENABLE"] = "true"
os.environ["MINERU_VLM_TABLE_ENABLE"] = "true"
os.environ["VLLM_USE_V1"] = "1"

# Load PDF bytes

with open("research_paper.pdf", "rb") as f:
    pdf_bytes = f.read()

# Execute parsing with vLLM backend

middle_json, raw_results = doc_analyze(
    pdf_bytes,
    image_writer=None,                    # Optional: DataWriter for extracted images

    backend="vllm-engine",               # Options: transformers, vllm-engine, lmdeploy-engine

    model_path="/models/qwen2vl",        # Optional: auto-downloads if omitted

    gpu_memory_utilization="0.7",        # Override auto-detected value

)

# Access structured content

print(middle_json[0]["blocks"][0]["type"])  # e.g., "TITLE"

print(middle_json[0]["blocks"][0]["content"])  # Extracted text or LaTeX

The ModelSingleton caches the vLLM LLM object, making subsequent calls instant. The function returns Content-List V2 format by default through union_make.

Asynchronous Python API

For web services or high-concurrency applications, use the async interface:

import asyncio
from mineru.backend.vlm.vlm_analyze import aio_doc_analyze

async def parse_document(pdf_path: str):
    with open(pdf_path, "rb") as f:
        pdf_bytes = f.read()
    
    middle_json, _ = await aio_doc_analyze(
        pdf_bytes,
        backend="vllm-async-engine",
        model_path="/models/qwen2vl",
    )
    return middle_json

# Execute

result = asyncio.run(parse_document("document.pdf"))

This path creates a vllm.AsyncLLM instance (lines 24-55 in vlm_analyze.py) capable of serving concurrent requests while applying the same high-accuracy post-processing pipeline.

Device-Specific Compilation

For non-standard GPUs like Kunlun XPU, configure compilation parameters:

export MINERU_VLLM_DEVICE=kxpu
export MINERU_VLLM_DEVICE_CONFIG='{"block_size":128,"dtype":"float16"}'
mineru vlm_server --engine vllm

The mod_kwargs_by_device_type function aligns the vLLM compilation pipeline with hardware-specific block sizes, optimizing both speed and accuracy for specialized accelerators.

Summary

  • MinerU's VLM backend uses a three-layer architecture: model initialization (ModelSingleton), inference (doc_analyze/aio_doc_analyze), and post-processing (MagicModel and union_make).
  • High-accuracy configuration requires setting MINERU_VLM_FORMULA_ENABLE=true and MINERU_VLM_TABLE_ENABLE=true to preserve LaTeX and HTML table structures.
  • GPU optimization relies on set_default_gpu_memory_utilization (0.7 for <8GB VRAM, 0.5 otherwise) and optional MINERU_VLLM_DEVICE configuration for custom hardware.
  • Backend options include transformers, vllm-engine, vllm-async-engine, lmdeploy-engine, and mlx-engine, each loaded via minerU/backend/vlm/vlm_analyze.py.
  • Output formats are controlled through MakeMode, with CONTENT_LIST_V2 providing the most structured representation for downstream applications.

Frequently Asked Questions

How do I choose between the vLLM and transformers backend?

Use the vLLM backend (vllm-engine or vllm-async-engine) for production workloads requiring high throughput and concurrent request handling, as implemented in minerU/backend/vlm/vlm_analyze.py lines 24-55. The transformers backend is ideal for quick prototyping or environments where vLLM dependencies cannot be installed, loading Qwen2-VL directly via Qwen2VLForConditionalGeneration.from_pretrained.

Why are my mathematical formulas being extracted as images instead of LaTeX?

Ensure you have set the environment variable MINERU_VLM_FORMULA_ENABLE=true before initializing the model. When this flag is disabled, the union_make function in minerU/backend/vlm/vlm_middle_json_mkcontent.py treats INTERLINE_EQUATION spans as images rather than preserving the LaTeX source code (lines 14-15).

Can I run MinerU's VLM backend on Apple Silicon Macs?

Yes, but only on macOS 13.5 or later using the mlx-engine backend. The code explicitly checks is_mac_os_version_supported before allowing MLX initialization. Note that this engine is not available for Linux or Windows environments, where you should use vllm-engine or transformers instead.

What causes out-of-memory errors during PDF parsing and how do I fix them?

OOM errors typically occur when GPU memory utilization exceeds available VRAM. The system auto-detects card capacity in minerU/backend/vlm/utils.py (lines 82-89), setting utilization to 0.7 for cards with less than 8 GiB and 0.5 otherwise. You can manually override this by passing gpu_memory_utilization="0.6" to doc_analyze or setting the VLLM_GPU_MEMORY_UTILIZATION environment variable.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →