How to Choose Between MinerU's Pipeline, VLM, and Hybrid Backends: A Complete Guide
Use the pipeline backend for CPU-only or low-VRAM environments (6GB+), the VLM backend for maximum accuracy on high-end GPUs (8GB+ VRAM), and the hybrid backend to combine VLM-grade layout detection with proven OCR and formula extraction while allowing remote server offload.
MinerU supports three distinct parsing architectures—pipeline, VLM, and hybrid—each optimized for different hardware constraints and accuracy requirements. Choosing the right backend determines whether your document extraction runs on a laptop CPU, a single high-end GPU, or distributed across edge clients and central servers. This guide explains how to choose between MinerU's pipeline, VLM, and hybrid backends based on the opendatalab/MinerU source code and practical deployment scenarios.
Architecture Overview
MinerU's three backends represent fundamentally different approaches to document parsing, ranging from classical computer vision cascades to modern vision-language models.
Pipeline Backend
The pipeline backend implements a classic multi-model cascade: layout detection → OCR → formula recognition → table extraction. In mineru/backend/pipeline/pipeline_analyze.py, each stage uses a specialized lightweight model.
- No hallucinations: Layout derives from dedicated detection models rather than generative outputs
- Full feature support: Complete access to table merging, formula detection, and multilingual OCR post-processing
- Minimal VRAM: Runs on CPU-only machines or GPUs with as little as 6GB VRAM
VLM Backend
The VLM (Vision-Language Model) backend processes documents in a single forward pass using a large vision-language model. The implementation in mineru/backend/vlm/vlm_analyze.py handles layout analysis, OCR, formula extraction, and table recognition simultaneously.
- Highest accuracy: Superior performance on complex multi-column layouts and dense scientific papers
- Single-pass inference: Eliminates cascade overhead, reducing latency for individual documents
- Hardware requirement: Requires a GPU with ≥8GB VRAM using vLLM, LMDeploy, or MLX engines
Hybrid Backend
The hybrid backend combines VLM layout detection with the pipeline's specialized OCR, formula, and table modules. Orchestrated in mineru/backend/hybrid/hybrid_analyze.py, this approach runs the VLM to identify layout blocks, then crops regions for processing by the pipeline's proven sub-modules.
- Balanced precision: VLM-grade layout quality with pipeline stability for text recognition
- Flexible deployment: Supports local GPU execution (
hybrid-auto-engine) or remote VLM servers (hybrid-http-client) - VRAM efficiency: Runs on ≥6GB GPUs locally, or shifts VLM load to remote servers while keeping clients lightweight
How Backend Selection Works Internally
When you specify a backend via CLI or API, MinerU follows a strict decision flow implemented across several core files:
-
Backend name parsing: The CLI extracts the backend string (e.g.,
-b hybrid-auto-engine) -
Engine automatic selection: For backends ending with
-auto-engine, theget_vlm_engine()function inmineru/utils/engine_utils.pyselects the optimal inference engine:- vLLM on Linux
- LMDeploy on Windows
- MLX on macOS
- Falls back to Transformers if specialized engines are unavailable
-
Hybrid execution path: If the backend starts with
hybrid, MinerU first runs the VLM model to obtain layout blocks, then feeds cropped image regions to the pipeline's OCR, formula, and table models -
VLM-only path: For
vlm-auto-engineorvlm-http-client, the entire document processes through the VLM model inmineru/backend/vlm/vlm_analyze.py -
Pipeline path: When using
pipeline, the classic cascade executes viamineru/backend/pipeline/pipeline_analyze.py
The system respects several environment variables during selection:
MINERU_HYBRID_BATCH_RATIO: Controls batch size for OCR/formula models in hybrid mode to fit various GPU sizesMINERU_FORCE_VLM_OCR_ENABLE: Forces VLM-based OCR even on low-VRAM devicesMINERU_HYBRID_FORCE_PIPELINE_ENABLE: Disables VLM OCR in hybrid mode for debugging purposes
Decision Matrix: When to Use Each Backend
Match your deployment scenario to the appropriate backend using these specific criteria:
| Situation | Recommended Backend | Rationale |
|---|---|---|
| CPU-only or GPU <6GB VRAM | pipeline |
No VLM engine can be loaded; lightweight models fit the memory budget |
| Maximum layout accuracy required and GPU ≥8GB available | vlm-auto-engine or vlm-http-client |
VLM yields superior layout detection on complex documents |
| VLM layout quality + proven OCR/formula needed with GPU ≥6GB | hybrid-auto-engine |
Combines VLM layout precision with pipeline stability |
| Edge devices processing via remote service | hybrid-http-client |
Client stays lightweight while delegating VLM work to a central server |
| Fastest end-to-end speed on single GPU | vlm-auto-engine |
Single forward pass eliminates cascade overhead |
| Deterministic, hallucination-free results | pipeline |
Layout model trained specifically for document structure without generative errors |
Practical Implementation Examples
CLI Backend Selection
Select your backend using the -b or --backend flag:
# Pure CPU pipeline (default for low-resource machines)
mineru -p mydoc.pdf -o out_dir -b pipeline
# VLM backend with automatic engine selection (requires 8GB+ GPU)
mineru -p mydoc.pdf -o out_dir -b vlm-auto-engine
# Hybrid local engine (6GB+ GPU)
mineru -p mydoc.pdf -o out_dir -b hybrid-auto-engine
# Hybrid remote client (edge device talking to central server)
mineru -p mydoc.pdf -o out_dir -b hybrid-http-client -u http://192.168.1.10:30000
Python API Integration
Explicitly set the backend when initializing the MinerU class:
from mineru import MinerU
# CPU-friendly pipeline
miner = MinerU(backend="pipeline")
# High-accuracy VLM (requires GPU)
# miner = MinerU(backend="vlm-auto-engine")
# Balanced hybrid approach
# miner = MinerU(backend="hybrid-auto-engine")
# Parse document
mid_json, results, vlm_ocr = miner.parse(pdf_path="mydoc.pdf")
Environment Variable Configuration
Configure defaults without CLI flags:
export MINERU_BACKEND=hybrid-auto-engine # Default backend for all commands
export MINERU_HYBRID_BATCH_RATIO=4 # Reduce VRAM usage on 6GB cards
export MINERU_FORCE_VLM_OCR_ENABLE=true # Force VLM OCR even on low-VRAM
mineru -p mydoc.pdf -o out_dir # Uses environment settings
Remote VLM Server Deployment
For distributed deployments, run a central VLM server and lightweight clients:
# On GPU-rich server node
mineru-openai-server --engine vllm --model MinerU2.5 --host 0.0.0.0 --port 30000
# On edge client machine
mineru -p mydoc.pdf -o out_dir -b hybrid-http-client -u http://<server_ip>:30000
Summary
- Pipeline backend: Choose for CPU-only environments, deterministic output, and minimal VRAM requirements (6GB+ if using GPU acceleration)
- VLM backend: Select when you have 8GB+ VRAM and require maximum layout accuracy on complex documents, accepting higher resource consumption for single-pass inference
- Hybrid backend: Use to balance VLM-grade layout detection with proven OCR and formula extraction, enabling both local 6GB GPU execution and remote server offload via HTTP client mode
- Engine selection: The
get_vlm_engine()function inmineru/utils/engine_utils.pyautomatically selects vLLM, LMDeploy, or MLX based on your operating system when using-auto-enginevariants
Frequently Asked Questions
What is the main difference between MinerU's pipeline and VLM backends?
The pipeline backend uses a cascade of specialized models (layout → OCR → formula → table) executed sequentially in mineru/backend/pipeline/pipeline_analyze.py, providing deterministic results without generative hallucinations. The VLM backend processes the entire document through a single vision-language model in mineru/backend/vlm/vlm_analyze.py, offering higher accuracy on complex layouts but requiring 8GB+ VRAM and producing generative outputs.
Can I run the MinerU VLM backend on a CPU-only machine?
No. The VLM backend requires a GPU with at least 8GB VRAM to load the vision-language model using vLLM, LMDeploy, or MLX engines. For CPU-only deployments, use the pipeline backend, or use hybrid-http-client mode to process documents on a CPU while offloading VLM inference to a remote GPU server.
How does the hybrid backend reduce VRAM usage compared to pure VLM mode?
The hybrid backend in mineru/backend/hybrid/hybrid_analyze.py uses the VLM only for layout detection, then processes cropped regions through the pipeline's lightweight OCR and formula models. This reduces VRAM requirements to 6GB for local execution. Additionally, hybrid-http-client mode allows multiple lightweight CPU clients to share a single VLM server, eliminating per-client GPU memory entirely.
When should I use hybrid-http-client versus hybrid-auto-engine?
Use hybrid-auto-engine when running MinerU on a local GPU with 6GB+ VRAM, allowing automatic engine selection via mineru/utils/engine_utils.py. Choose hybrid-http-client when deploying to edge devices without GPUs or when consolidating GPU resources—this mode sends layout detection requests to a remote OpenAI-compatible server (e.g., mineru-openai-server) while running OCR and formula extraction locally.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →