MinerU Backend Source Files: Complete Guide to VLM, Pipeline, and Hybrid Locations
The key MinerU backend source files are located in the mineru/backend/ directory, with three main sub-packages for VLM (vlm/), Pipeline (pipeline/), and Hybrid (hybrid/) processing, plus shared utilities in utils.py.
The opendatalab/MinerU repository implements three distinct backend pipelines for document analysis. Understanding the MinerU backend source files is essential for customizing extraction workflows, debugging inference issues, or extending the toolkit with new vision-language models.
MinerU Backend Architecture Overview
All backend logic resides under the top-level package mineru.backend. The architecture separates concerns into three specialized pipelines:
- VLM Backend: Vision-language model inference for end-to-end document understanding
- Pipeline Backend: Classic OCR-first approach using layout detection, OCR, and formula detection
- Hybrid Backend: Combines VLM extraction with OCR fallback for complex documents
Each sub-package follows a consistent pattern with an analyze module, magic model wrapper, and output conversion utilities. Shared functionality for device selection and memory management lives in mineru/backend/utils.py.
VLM Backend Source Files
The VLM backend handles vision-language-model inference, model-singleton management, and conversion of raw VLM output to the intermediate JSON format.
Core Analysis Module (vlm_analyze.py)
The primary entry point resides in mineru/backend/vlm/vlm_analyze.py. This module implements the doc_analyze function and maintains a MinerUClient singleton to avoid reloading models across multiple PDF pages.
Key responsibilities include:
- Image loading and preprocessing for VLM consumption
- Managing the
MinerUClientinstance (lines 22-41) - Coordinating with
vlm_magic_model.pyfor engine-specific inference
Model Engine Wrapper (vlm_magic_model.py)
Located at mineru/backend/vlm/vlm_magic_model.py, this module abstracts the underlying VLM engines. It supports multiple backends including transformers, vllm-engine, and lmdeploy-engine.
The magic model pattern allows the analyze module to call a unified interface while the wrapper handles engine-specific tensor operations and batching strategies.
Output Conversion (model_output_to_middle_json.py)
The file mineru/backend/vlm/model_output_to_middle_json.py transforms raw VLM detections into MinerU's standard intermediate JSON schema. This normalization step ensures downstream tools receive consistent data regardless of which VLM engine generated the predictions.
Pipeline Backend Source Files
The Pipeline backend provides the classic OCR-first processing chain: layout detection → OCR → formula detection.
Entry Point (pipeline_analyze.py)
mineru/backend/pipeline/pipeline_analyze.py serves as the main interface. It implements the doc_analyze function and manages the MineruPipelineModel singleton (lines 18-41).
This module coordinates the three-stage pipeline:
- Layout detection to identify text regions
- OCR processing on detected regions
- Formula detection for mathematical expressions
Model Initialization (model_init.py)
The mineru/backend/pipeline/model_init.py module centralizes loading for all three model types: layout detection, OCR, and formula recognition. It handles device selection (CPU vs. GPU) and ensures models are loaded only once per process.
Batch Processing (batch_analyze.py)
Located at mineru/backend/pipeline/batch_analyze.py, this utility optimizes throughput by processing OCR and formula detection in batches. It dynamically adjusts batch size based on available VRAM to prevent out-of-memory errors on large documents.
Model Stack (pipeline_magic_model.py)
mineru/backend/pipeline/pipeline_magic_model.py constructs the MineruPipelineModel instance that encapsulates the layout, OCR, and formula models. This abstraction allows pipeline_analyze.py to call a single predict method while the magic model handles the multi-stage inference logic.
Hybrid Backend Source Files
The Hybrid backend combines VLM extraction with OCR fallback, masking image and table regions before running formula detection and merging results.
Orchestration Logic (hybrid_analyze.py)
mineru/backend/hybrid/hybrid_analyze.py contains the core orchestration logic (lines 84-150). This module:
- Runs initial VLM extraction to identify text and structural elements
- Masks image and table regions to prevent double-processing
- Falls back to OCR for regions where VLM confidence is low
- Coordinates formula detection on the merged content
Combined Model Stack (hybrid_magic_model.py)
The mineru/backend/hybrid/hybrid_magic_model.py module instantiates the combined model stack, managing both the VLM components and the OCR pipeline models within a single interface. This allows the hybrid analyzer to switch between extraction strategies without reloading weights.
Result Merging (hybrid_model_output_to_middle_json.py)
Located at mineru/backend/hybrid/hybrid_model_output_to_middle_json.py, this module merges VLM and OCR results into the unified JSON schema. It resolves conflicts between the two extraction paths and ensures the final output maintains consistent bounding boxes and text ordering.
Shared Backend Utilities
All three backends rely on common functionality in mineru/backend/utils.py. This module provides:
- Device Detection: Automatically selects between
cpu,cuda,mps, andnpubackends based on available hardware - VRAM Management: The
get_vramfunction queries available GPU memory to set safe default batch sizes - Configuration Helpers:
set_default_gpu_memory_utilizationandset_default_batch_sizeallow runtime adjustment of memory limits without modifying config files
These utilities ensure consistent behavior across VLM, Pipeline, and Hybrid backends when handling GPU resources.
Practical Usage Examples
Using the VLM Backend
from mineru.backend.vlm import vlm_analyze
# pdf_bytes: raw PDF content (e.g. read from a file)
middle_json, results = vlm_analyze.doc_analyze(
pdf_bytes,
image_writer=None, # optional DataWriter for image output
backend="transformers", # or "vllm-engine", "lmdeploy-engine"
model_path="/path/to/qwen2vl", # optional; auto-downloaded if omitted
)
Using the Pipeline Backend
from mineru.backend.pipeline import pipeline_analyze
middle_json, results, _ = pipeline_analyze.doc_analyze(
pdf_bytes,
image_writer=None,
backend="transformers", # layout + OCR are always transformers-based
parse_method="auto", # auto-detect OCR need
)
Using the Hybrid Backend
from mineru.backend.hybrid import hybrid_analyze
middle_json, results, vlm_ocr_enabled = hybrid_analyze.doc_analyze(
pdf_bytes,
image_writer=None,
backend="transformers", # VLM backend
parse_method="auto",
language="ch", # language code for OCR models
inline_formula_enable=True,
)
All three calls return a middle JSON structure that the rest of MinerU consumes for downstream processing.
Summary
- MinerU backend source files are organized under
mineru/backend/with three distinct sub-packages:vlm/,pipeline/, andhybrid/. - VLM backend files in
mineru/backend/vlm/handle vision-language model inference throughvlm_analyze.py, engine abstraction invlm_magic_model.py, and JSON conversion inmodel_output_to_middle_json.py. - Pipeline backend files in
mineru/backend/pipeline/implement the OCR-first workflow viapipeline_analyze.py, model initialization inmodel_init.py, and batch optimization inbatch_analyze.py. - Hybrid backend files in
mineru/backend/hybrid/combine both approaches usinghybrid_analyze.pyfor orchestration andhybrid_model_output_to_middle_json.pyfor result merging. - Shared utilities in
mineru/backend/utils.pymanage device selection, VRAM queries, and batch-size heuristics across all backends.
Frequently Asked Questions
Where are the MinerU backend source files located?
The MinerU backend source files are located in the mineru/backend/ directory of the repository. This top-level package contains three sub-directories: vlm/ for vision-language model processing, pipeline/ for traditional OCR pipelines, and hybrid/ for combined approaches, along with a shared utils.py file.
What is the difference between the VLM and Pipeline backends?
The VLM backend uses vision-language models (like Qwen2-VL) for end-to-end document understanding through files like vlm_analyze.py and vlm_magic_model.py. The Pipeline backend follows a classic multi-stage approach using pipeline_analyze.py, running layout detection first, then OCR, then formula detection through the MineruPipelineModel defined in pipeline_magic_model.py.
How does the Hybrid backend combine VLM and Pipeline approaches?
The Hybrid backend, orchestrated in hybrid_analyze.py, first runs VLM extraction to identify text and structural elements, then masks image and table regions to prevent duplicate processing. It falls back to OCR (from the Pipeline backend) for low-confidence regions and runs formula detection before merging all results via hybrid_model_output_to_middle_json.py into a unified JSON format.
What shared utilities are available across all MinerU backends?
All backends use mineru/backend/utils.py for common operations including automatic device detection (CPU, CUDA, MPS, NPU), VRAM querying via get_vram, and batch size management through set_default_batch_size and set_default_gpu_memory_utilization. These utilities ensure consistent GPU resource handling whether using VLM, Pipeline, or Hybrid backends.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →