How to Choose Between MinerU's Pipeline, VLM, and Hybrid Backends: A Complete Guide

Use the pipeline backend for CPU-only or low-VRAM environments (6GB+), the VLM backend for maximum accuracy on high-end GPUs (8GB+ VRAM), and the hybrid backend to combine VLM-grade layout detection with proven OCR and formula extraction while allowing remote server offload.

MinerU supports three distinct parsing architectures—pipeline, VLM, and hybrid—each optimized for different hardware constraints and accuracy requirements. Choosing the right backend determines whether your document extraction runs on a laptop CPU, a single high-end GPU, or distributed across edge clients and central servers. This guide explains how to choose between MinerU's pipeline, VLM, and hybrid backends based on the opendatalab/MinerU source code and practical deployment scenarios.

Architecture Overview

MinerU's three backends represent fundamentally different approaches to document parsing, ranging from classical computer vision cascades to modern vision-language models.

Pipeline Backend

The pipeline backend implements a classic multi-model cascade: layout detection → OCR → formula recognition → table extraction. In mineru/backend/pipeline/pipeline_analyze.py, each stage uses a specialized lightweight model.

  • No hallucinations: Layout derives from dedicated detection models rather than generative outputs
  • Full feature support: Complete access to table merging, formula detection, and multilingual OCR post-processing
  • Minimal VRAM: Runs on CPU-only machines or GPUs with as little as 6GB VRAM

VLM Backend

The VLM (Vision-Language Model) backend processes documents in a single forward pass using a large vision-language model. The implementation in mineru/backend/vlm/vlm_analyze.py handles layout analysis, OCR, formula extraction, and table recognition simultaneously.

  • Highest accuracy: Superior performance on complex multi-column layouts and dense scientific papers
  • Single-pass inference: Eliminates cascade overhead, reducing latency for individual documents
  • Hardware requirement: Requires a GPU with ≥8GB VRAM using vLLM, LMDeploy, or MLX engines

Hybrid Backend

The hybrid backend combines VLM layout detection with the pipeline's specialized OCR, formula, and table modules. Orchestrated in mineru/backend/hybrid/hybrid_analyze.py, this approach runs the VLM to identify layout blocks, then crops regions for processing by the pipeline's proven sub-modules.

  • Balanced precision: VLM-grade layout quality with pipeline stability for text recognition
  • Flexible deployment: Supports local GPU execution (hybrid-auto-engine) or remote VLM servers (hybrid-http-client)
  • VRAM efficiency: Runs on ≥6GB GPUs locally, or shifts VLM load to remote servers while keeping clients lightweight

How Backend Selection Works Internally

When you specify a backend via CLI or API, MinerU follows a strict decision flow implemented across several core files:

  1. Backend name parsing: The CLI extracts the backend string (e.g., -b hybrid-auto-engine)

  2. Engine automatic selection: For backends ending with -auto-engine, the get_vlm_engine() function in mineru/utils/engine_utils.py selects the optimal inference engine:

    • vLLM on Linux
    • LMDeploy on Windows
    • MLX on macOS
    • Falls back to Transformers if specialized engines are unavailable
  3. Hybrid execution path: If the backend starts with hybrid, MinerU first runs the VLM model to obtain layout blocks, then feeds cropped image regions to the pipeline's OCR, formula, and table models

  4. VLM-only path: For vlm-auto-engine or vlm-http-client, the entire document processes through the VLM model in mineru/backend/vlm/vlm_analyze.py

  5. Pipeline path: When using pipeline, the classic cascade executes via mineru/backend/pipeline/pipeline_analyze.py

The system respects several environment variables during selection:

  • MINERU_HYBRID_BATCH_RATIO: Controls batch size for OCR/formula models in hybrid mode to fit various GPU sizes
  • MINERU_FORCE_VLM_OCR_ENABLE: Forces VLM-based OCR even on low-VRAM devices
  • MINERU_HYBRID_FORCE_PIPELINE_ENABLE: Disables VLM OCR in hybrid mode for debugging purposes

Decision Matrix: When to Use Each Backend

Match your deployment scenario to the appropriate backend using these specific criteria:

Situation Recommended Backend Rationale
CPU-only or GPU <6GB VRAM pipeline No VLM engine can be loaded; lightweight models fit the memory budget
Maximum layout accuracy required and GPU ≥8GB available vlm-auto-engine or vlm-http-client VLM yields superior layout detection on complex documents
VLM layout quality + proven OCR/formula needed with GPU ≥6GB hybrid-auto-engine Combines VLM layout precision with pipeline stability
Edge devices processing via remote service hybrid-http-client Client stays lightweight while delegating VLM work to a central server
Fastest end-to-end speed on single GPU vlm-auto-engine Single forward pass eliminates cascade overhead
Deterministic, hallucination-free results pipeline Layout model trained specifically for document structure without generative errors

Practical Implementation Examples

CLI Backend Selection

Select your backend using the -b or --backend flag:


# Pure CPU pipeline (default for low-resource machines)

mineru -p mydoc.pdf -o out_dir -b pipeline

# VLM backend with automatic engine selection (requires 8GB+ GPU)

mineru -p mydoc.pdf -o out_dir -b vlm-auto-engine

# Hybrid local engine (6GB+ GPU)

mineru -p mydoc.pdf -o out_dir -b hybrid-auto-engine

# Hybrid remote client (edge device talking to central server)

mineru -p mydoc.pdf -o out_dir -b hybrid-http-client -u http://192.168.1.10:30000

Python API Integration

Explicitly set the backend when initializing the MinerU class:

from mineru import MinerU

# CPU-friendly pipeline

miner = MinerU(backend="pipeline")

# High-accuracy VLM (requires GPU)

# miner = MinerU(backend="vlm-auto-engine")

# Balanced hybrid approach

# miner = MinerU(backend="hybrid-auto-engine")

# Parse document

mid_json, results, vlm_ocr = miner.parse(pdf_path="mydoc.pdf")

Environment Variable Configuration

Configure defaults without CLI flags:

export MINERU_BACKEND=hybrid-auto-engine        # Default backend for all commands

export MINERU_HYBRID_BATCH_RATIO=4             # Reduce VRAM usage on 6GB cards

export MINERU_FORCE_VLM_OCR_ENABLE=true        # Force VLM OCR even on low-VRAM

mineru -p mydoc.pdf -o out_dir                  # Uses environment settings

Remote VLM Server Deployment

For distributed deployments, run a central VLM server and lightweight clients:


# On GPU-rich server node

mineru-openai-server --engine vllm --model MinerU2.5 --host 0.0.0.0 --port 30000

# On edge client machine

mineru -p mydoc.pdf -o out_dir -b hybrid-http-client -u http://<server_ip>:30000

Summary

  • Pipeline backend: Choose for CPU-only environments, deterministic output, and minimal VRAM requirements (6GB+ if using GPU acceleration)
  • VLM backend: Select when you have 8GB+ VRAM and require maximum layout accuracy on complex documents, accepting higher resource consumption for single-pass inference
  • Hybrid backend: Use to balance VLM-grade layout detection with proven OCR and formula extraction, enabling both local 6GB GPU execution and remote server offload via HTTP client mode
  • Engine selection: The get_vlm_engine() function in mineru/utils/engine_utils.py automatically selects vLLM, LMDeploy, or MLX based on your operating system when using -auto-engine variants

Frequently Asked Questions

What is the main difference between MinerU's pipeline and VLM backends?

The pipeline backend uses a cascade of specialized models (layout → OCR → formula → table) executed sequentially in mineru/backend/pipeline/pipeline_analyze.py, providing deterministic results without generative hallucinations. The VLM backend processes the entire document through a single vision-language model in mineru/backend/vlm/vlm_analyze.py, offering higher accuracy on complex layouts but requiring 8GB+ VRAM and producing generative outputs.

Can I run the MinerU VLM backend on a CPU-only machine?

No. The VLM backend requires a GPU with at least 8GB VRAM to load the vision-language model using vLLM, LMDeploy, or MLX engines. For CPU-only deployments, use the pipeline backend, or use hybrid-http-client mode to process documents on a CPU while offloading VLM inference to a remote GPU server.

How does the hybrid backend reduce VRAM usage compared to pure VLM mode?

The hybrid backend in mineru/backend/hybrid/hybrid_analyze.py uses the VLM only for layout detection, then processes cropped regions through the pipeline's lightweight OCR and formula models. This reduces VRAM requirements to 6GB for local execution. Additionally, hybrid-http-client mode allows multiple lightweight CPU clients to share a single VLM server, eliminating per-client GPU memory entirely.

When should I use hybrid-http-client versus hybrid-auto-engine?

Use hybrid-auto-engine when running MinerU on a local GPU with 6GB+ VRAM, allowing automatic engine selection via mineru/utils/engine_utils.py. Choose hybrid-http-client when deploying to edge devices without GPUs or when consolidating GPU resources—this mode sends layout detection requests to a remote OpenAI-compatible server (e.g., mineru-openai-server) while running OCR and formula extraction locally.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →