MinerU Backend Source Files: Complete Guide to VLM, Pipeline, and Hybrid Locations

The key MinerU backend source files are located in the mineru/backend/ directory, with three main sub-packages for VLM (vlm/), Pipeline (pipeline/), and Hybrid (hybrid/) processing, plus shared utilities in utils.py.

The opendatalab/MinerU repository implements three distinct backend pipelines for document analysis. Understanding the MinerU backend source files is essential for customizing extraction workflows, debugging inference issues, or extending the toolkit with new vision-language models.

MinerU Backend Architecture Overview

All backend logic resides under the top-level package mineru.backend. The architecture separates concerns into three specialized pipelines:

  • VLM Backend: Vision-language model inference for end-to-end document understanding
  • Pipeline Backend: Classic OCR-first approach using layout detection, OCR, and formula detection
  • Hybrid Backend: Combines VLM extraction with OCR fallback for complex documents

Each sub-package follows a consistent pattern with an analyze module, magic model wrapper, and output conversion utilities. Shared functionality for device selection and memory management lives in mineru/backend/utils.py.

VLM Backend Source Files

The VLM backend handles vision-language-model inference, model-singleton management, and conversion of raw VLM output to the intermediate JSON format.

Core Analysis Module (vlm_analyze.py)

The primary entry point resides in mineru/backend/vlm/vlm_analyze.py. This module implements the doc_analyze function and maintains a MinerUClient singleton to avoid reloading models across multiple PDF pages.

Key responsibilities include:

  • Image loading and preprocessing for VLM consumption
  • Managing the MinerUClient instance (lines 22-41)
  • Coordinating with vlm_magic_model.py for engine-specific inference

Model Engine Wrapper (vlm_magic_model.py)

Located at mineru/backend/vlm/vlm_magic_model.py, this module abstracts the underlying VLM engines. It supports multiple backends including transformers, vllm-engine, and lmdeploy-engine.

The magic model pattern allows the analyze module to call a unified interface while the wrapper handles engine-specific tensor operations and batching strategies.

Output Conversion (model_output_to_middle_json.py)

The file mineru/backend/vlm/model_output_to_middle_json.py transforms raw VLM detections into MinerU's standard intermediate JSON schema. This normalization step ensures downstream tools receive consistent data regardless of which VLM engine generated the predictions.

Pipeline Backend Source Files

The Pipeline backend provides the classic OCR-first processing chain: layout detection → OCR → formula detection.

Entry Point (pipeline_analyze.py)

mineru/backend/pipeline/pipeline_analyze.py serves as the main interface. It implements the doc_analyze function and manages the MineruPipelineModel singleton (lines 18-41).

This module coordinates the three-stage pipeline:

  1. Layout detection to identify text regions
  2. OCR processing on detected regions
  3. Formula detection for mathematical expressions

Model Initialization (model_init.py)

The mineru/backend/pipeline/model_init.py module centralizes loading for all three model types: layout detection, OCR, and formula recognition. It handles device selection (CPU vs. GPU) and ensures models are loaded only once per process.

Batch Processing (batch_analyze.py)

Located at mineru/backend/pipeline/batch_analyze.py, this utility optimizes throughput by processing OCR and formula detection in batches. It dynamically adjusts batch size based on available VRAM to prevent out-of-memory errors on large documents.

Model Stack (pipeline_magic_model.py)

mineru/backend/pipeline/pipeline_magic_model.py constructs the MineruPipelineModel instance that encapsulates the layout, OCR, and formula models. This abstraction allows pipeline_analyze.py to call a single predict method while the magic model handles the multi-stage inference logic.

Hybrid Backend Source Files

The Hybrid backend combines VLM extraction with OCR fallback, masking image and table regions before running formula detection and merging results.

Orchestration Logic (hybrid_analyze.py)

mineru/backend/hybrid/hybrid_analyze.py contains the core orchestration logic (lines 84-150). This module:

  • Runs initial VLM extraction to identify text and structural elements
  • Masks image and table regions to prevent double-processing
  • Falls back to OCR for regions where VLM confidence is low
  • Coordinates formula detection on the merged content

Combined Model Stack (hybrid_magic_model.py)

The mineru/backend/hybrid/hybrid_magic_model.py module instantiates the combined model stack, managing both the VLM components and the OCR pipeline models within a single interface. This allows the hybrid analyzer to switch between extraction strategies without reloading weights.

Result Merging (hybrid_model_output_to_middle_json.py)

Located at mineru/backend/hybrid/hybrid_model_output_to_middle_json.py, this module merges VLM and OCR results into the unified JSON schema. It resolves conflicts between the two extraction paths and ensures the final output maintains consistent bounding boxes and text ordering.

Shared Backend Utilities

All three backends rely on common functionality in mineru/backend/utils.py. This module provides:

  • Device Detection: Automatically selects between cpu, cuda, mps, and npu backends based on available hardware
  • VRAM Management: The get_vram function queries available GPU memory to set safe default batch sizes
  • Configuration Helpers: set_default_gpu_memory_utilization and set_default_batch_size allow runtime adjustment of memory limits without modifying config files

These utilities ensure consistent behavior across VLM, Pipeline, and Hybrid backends when handling GPU resources.

Practical Usage Examples

Using the VLM Backend

from mineru.backend.vlm import vlm_analyze

# pdf_bytes: raw PDF content (e.g. read from a file)

middle_json, results = vlm_analyze.doc_analyze(
    pdf_bytes,
    image_writer=None,          # optional DataWriter for image output

    backend="transformers",    # or "vllm-engine", "lmdeploy-engine"

    model_path="/path/to/qwen2vl",  # optional; auto-downloaded if omitted

)

Using the Pipeline Backend

from mineru.backend.pipeline import pipeline_analyze

middle_json, results, _ = pipeline_analyze.doc_analyze(
    pdf_bytes,
    image_writer=None,
    backend="transformers",    # layout + OCR are always transformers-based

    parse_method="auto",      # auto-detect OCR need

)

Using the Hybrid Backend

from mineru.backend.hybrid import hybrid_analyze

middle_json, results, vlm_ocr_enabled = hybrid_analyze.doc_analyze(
    pdf_bytes,
    image_writer=None,
    backend="transformers",    # VLM backend

    parse_method="auto",
    language="ch",            # language code for OCR models

    inline_formula_enable=True,
)

All three calls return a middle JSON structure that the rest of MinerU consumes for downstream processing.

Summary

Frequently Asked Questions

Where are the MinerU backend source files located?

The MinerU backend source files are located in the mineru/backend/ directory of the repository. This top-level package contains three sub-directories: vlm/ for vision-language model processing, pipeline/ for traditional OCR pipelines, and hybrid/ for combined approaches, along with a shared utils.py file.

What is the difference between the VLM and Pipeline backends?

The VLM backend uses vision-language models (like Qwen2-VL) for end-to-end document understanding through files like vlm_analyze.py and vlm_magic_model.py. The Pipeline backend follows a classic multi-stage approach using pipeline_analyze.py, running layout detection first, then OCR, then formula detection through the MineruPipelineModel defined in pipeline_magic_model.py.

How does the Hybrid backend combine VLM and Pipeline approaches?

The Hybrid backend, orchestrated in hybrid_analyze.py, first runs VLM extraction to identify text and structural elements, then masks image and table regions to prevent duplicate processing. It falls back to OCR (from the Pipeline backend) for low-confidence regions and runs formula detection before merging all results via hybrid_model_output_to_middle_json.py into a unified JSON format.

What shared utilities are available across all MinerU backends?

All backends use mineru/backend/utils.py for common operations including automatic device detection (CPU, CUDA, MPS, NPU), VRAM querying via get_vram, and batch size management through set_default_batch_size and set_default_gpu_memory_utilization. These utilities ensure consistent GPU resource handling whether using VLM, Pipeline, or Hybrid backends.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →