# How to Utilize MinerU's VLM Backend for High-Accuracy Parsing

> Master PDF parsing with MinerU's VLM backend. Configure environment variables, choose an inference engine, and use MagicModel for accurate formula and table extraction, delivering structured Content-List V2 output.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: how-to-guide
- Published: 2026-02-23

---

**To achieve high-accuracy PDF parsing with MinerU, configure the Vision-Language Model (VLM) backend using environment variables for formula and table extraction, select the appropriate inference engine (vLLM, transformers, or LMDeploy), and leverage the MagicModel post-processor to convert raw spans into structured Content-List V2 output.**

The **MinerU VLM backend** serves as the core engine in the opendatalab/MinerU repository that transforms PDF pages into richly-structured middle-JSON representations. By understanding the three-layer pipeline—model initialization, inference, and post-processing—you can optimize the system for maximum extraction fidelity across complex documents containing equations, tables, and multilingual text.

## Understanding the VLM Backend Architecture

The parsing pipeline consists of three distinct logical layers that process input from raw bytes to structured output.

### Model Initialization Layer

The entry point for any VLM operation begins in [`minerU/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_analyze.py) through the `ModelSingleton.get_model` method. This **singleton pattern** guarantees that the same backend and model combination is instantiated only once per process, eliminating redundant heavy model loads.

MinerU supports five distinct backend engines:

- **transformers**: Loads Qwen2-VL via `Qwen2VLForConditionalGeneration.from_pretrained`, automatically selecting the correct dtype key based on your transformers version (lines 69-73 in [`vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/vlm_analyze.py))
- **vllm-engine**: Creates a synchronous `vllm.LLM` object
- **vllm-async-engine**: Creates an asynchronous `vllm.AsyncLLM` instance for concurrent requests
- **lmdeploy-engine**: Initializes a `VLAsyncEngine` with automatic backend selection (PyTorch or Turbomind) via `set_lmdeploy_backend` (lines 59-78)
- **mlx-engine**: Exclusive to macOS 13.5+ Apple Silicon, guarded by `is_mac_os_version_supported`

The initialization layer in [`minerU/backend/vlm/utils.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/utils.py) provides critical helper functions including `set_default_gpu_memory_utilization`, which returns **0.7** for low-VRAM cards (<8 GiB) and **0.5** otherwise (lines 82-89), and `mod_kwargs_by_device_type` for hardware-specific compilation configuration.

### Inference Layer

Once initialized, the system processes page-wise image lists through `doc_analyze` or `aio_doc_analyze` in [`vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/vlm_analyze.py). These functions encapsulate calls to `MinerUClient.batch_two_step_extract`, which feeds images to the selected VLM and returns **raw spans** containing text, inline equations, display equations, images, and tables with bounding box coordinates.

If CUDA is present and `VLLM_USE_V1=1` is set (default), the system automatically injects `MinerULogitsProcessor` via `enable_custom_logits_processors` to improve token-level decisions for Chinese, Japanese, Korean, and hyphenated English text.

### Post-Processing Layer

Raw VLM outputs undergo sophisticated transformation through two key components:

1. **MagicModel** ([`minerU/backend/vlm/vlm_magic_model.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_magic_model.py)): The `MagicModel` class (lines 13-183) normalizes coordinates, splits inline formulas, fixes hyphenation, and groups spans into hierarchical lines and blocks. It handles complex layout cases including tables with misplaced captions, nested lists, and algorithm-style code through `fix_two_layer_blocks` and `fix_list_blocks`.

2. **Middle-JSON Creation** ([`minerU/backend/vlm/vlm_middle_json_mkcontent.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_middle_json_mkcontent.py)): The `union_make` function orchestrates conversion to your target format, selecting from `MakeMode.MM_MD`, `MakeMode.NLP_MD`, `MakeMode.CONTENT_LIST`, or `MakeMode.CONTENT_LIST_V2`. This process stitches processed blocks, adds page indices, normalizes bounding boxes to a 0-1000 scale, and optionally injects original image paths.

## Configuration for Maximum Accuracy

Achieving high-accuracy parsing requires specific environment configurations that control how the VLM handles complex document elements.

### Enable Formula and Table Recognition

Set these environment variables before initialization to preserve semantic structure:

```bash
export MINERU_VLM_FORMULA_ENABLE=true   # Keeps LaTeX as display equations

export MINERU_VLM_TABLE_ENABLE=true     # Emits HTML tables when recognized

```

When `MINERU_VLM_FORMULA_ENABLE` is true, `INTERLINE_EQUATION` spans remain as LaTeX rather than converting to images. Similarly, `MINERU_VLM_TABLE_ENABLE` allows `mk_blocks_to_markdown` (lines 41-62 in [`vlm_middle_json_mkcontent.py`](https://github.com/opendatalab/MinerU/blob/main/vlm_middle_json_mkcontent.py)) to emit HTML `<table>` tags instead of image fallbacks.

### GPU Memory and Performance Tuning

Control resource allocation to prevent out-of-memory crashes while maintaining throughput:

```bash

# Override auto-detection (0.7 for <8GB, 0.5 otherwise)

export VLLM_GPU_MEMORY_UTILIZATION=0.7

# Enable custom logits processing for CJK and hyphenation

export VLLM_USE_V1=1

```

For specialized hardware, export device-specific configuration:

```bash
export MINERU_VLLM_DEVICE=kxpu  # or corex

export MINERU_VLLM_DEVICE_CONFIG='{"block_size":128,"dtype":"float16"}'

```

The `mod_kwargs_by_device_type` function (lines 71-89 in [`utils.py`](https://github.com/opendatalab/MinerU/blob/main/utils.py)) reads these variables to inject appropriate `--block-size` and `--dtype` arguments into the vLLM compilation pipeline.

## Implementing the Parsing Pipeline

### Starting a VLM Server

For production deployments, launch a dedicated inference server using the built-in CLI:

```bash

# vLLM backend (default)

mineru vlm_server --engine vllm

# LMDeploy backend for alternative acceleration

mineru vlm_server --engine lmdeploy

```

The CLI entry point in [`minerU/cli/vlm_server.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/vlm_server.py) forwards to [`minerU/model/vlm/vllm_server.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/model/vlm/vllm_server.py) or [`minerU/model/vlm/lmdeploy_server.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/model/vlm/lmdeploy_server.py), automatically injecting GPU memory and logits settings.

### Synchronous Python API

For batch processing or simple scripts, use the high-level `doc_analyze` function:

```python
import os
from mineru.backend.vlm.vlm_analyze import doc_analyze

# Configure for maximum accuracy

os.environ["MINERU_VLM_FORMULA_ENABLE"] = "true"
os.environ["MINERU_VLM_TABLE_ENABLE"] = "true"
os.environ["VLLM_USE_V1"] = "1"

# Load PDF bytes

with open("research_paper.pdf", "rb") as f:
    pdf_bytes = f.read()

# Execute parsing with vLLM backend

middle_json, raw_results = doc_analyze(
    pdf_bytes,
    image_writer=None,                    # Optional: DataWriter for extracted images

    backend="vllm-engine",               # Options: transformers, vllm-engine, lmdeploy-engine

    model_path="/models/qwen2vl",        # Optional: auto-downloads if omitted

    gpu_memory_utilization="0.7",        # Override auto-detected value

)

# Access structured content

print(middle_json[0]["blocks"][0]["type"])  # e.g., "TITLE"

print(middle_json[0]["blocks"][0]["content"])  # Extracted text or LaTeX

```

The `ModelSingleton` caches the vLLM `LLM` object, making subsequent calls instant. The function returns Content-List V2 format by default through `union_make`.

### Asynchronous Python API

For web services or high-concurrency applications, use the async interface:

```python
import asyncio
from mineru.backend.vlm.vlm_analyze import aio_doc_analyze

async def parse_document(pdf_path: str):
    with open(pdf_path, "rb") as f:
        pdf_bytes = f.read()
    
    middle_json, _ = await aio_doc_analyze(
        pdf_bytes,
        backend="vllm-async-engine",
        model_path="/models/qwen2vl",
    )
    return middle_json

# Execute

result = asyncio.run(parse_document("document.pdf"))

```

This path creates a `vllm.AsyncLLM` instance (lines 24-55 in [`vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/vlm_analyze.py)) capable of serving concurrent requests while applying the same high-accuracy post-processing pipeline.

### Device-Specific Compilation

For non-standard GPUs like Kunlun XPU, configure compilation parameters:

```bash
export MINERU_VLLM_DEVICE=kxpu
export MINERU_VLLM_DEVICE_CONFIG='{"block_size":128,"dtype":"float16"}'
mineru vlm_server --engine vllm

```

The `mod_kwargs_by_device_type` function aligns the vLLM compilation pipeline with hardware-specific block sizes, optimizing both speed and accuracy for specialized accelerators.

## Summary

- **MinerU's VLM backend** uses a three-layer architecture: model initialization (`ModelSingleton`), inference (`doc_analyze`/`aio_doc_analyze`), and post-processing (`MagicModel` and `union_make`).
- **High-accuracy configuration** requires setting `MINERU_VLM_FORMULA_ENABLE=true` and `MINERU_VLM_TABLE_ENABLE=true` to preserve LaTeX and HTML table structures.
- **GPU optimization** relies on `set_default_gpu_memory_utilization` (0.7 for <8GB VRAM, 0.5 otherwise) and optional `MINERU_VLLM_DEVICE` configuration for custom hardware.
- **Backend options** include `transformers`, `vllm-engine`, `vllm-async-engine`, `lmdeploy-engine`, and `mlx-engine`, each loaded via [`minerU/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_analyze.py).
- **Output formats** are controlled through `MakeMode`, with `CONTENT_LIST_V2` providing the most structured representation for downstream applications.

## Frequently Asked Questions

### How do I choose between the vLLM and transformers backend?

**Use the vLLM backend (`vllm-engine` or `vllm-async-engine`) for production workloads requiring high throughput and concurrent request handling**, as implemented in [`minerU/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_analyze.py) lines 24-55. The transformers backend is ideal for quick prototyping or environments where vLLM dependencies cannot be installed, loading Qwen2-VL directly via `Qwen2VLForConditionalGeneration.from_pretrained`.

### Why are my mathematical formulas being extracted as images instead of LaTeX?

**Ensure you have set the environment variable `MINERU_VLM_FORMULA_ENABLE=true` before initializing the model.** When this flag is disabled, the `union_make` function in [`minerU/backend/vlm/vlm_middle_json_mkcontent.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_middle_json_mkcontent.py) treats `INTERLINE_EQUATION` spans as images rather than preserving the LaTeX source code (lines 14-15).

### Can I run MinerU's VLM backend on Apple Silicon Macs?

**Yes, but only on macOS 13.5 or later using the `mlx-engine` backend.** The code explicitly checks `is_mac_os_version_supported` before allowing MLX initialization. Note that this engine is not available for Linux or Windows environments, where you should use `vllm-engine` or `transformers` instead.

### What causes out-of-memory errors during PDF parsing and how do I fix them?

**OOM errors typically occur when GPU memory utilization exceeds available VRAM.** The system auto-detects card capacity in [`minerU/backend/vlm/utils.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/utils.py) (lines 82-89), setting utilization to 0.7 for cards with less than 8 GiB and 0.5 otherwise. You can manually override this by passing `gpu_memory_utilization="0.6"` to `doc_analyze` or setting the `VLLM_GPU_MEMORY_UTILIZATION` environment variable.