# How MinerU's Hybrid Backend Combines Pipeline and VLM Advantages for PDF Parsing

> Discover how MinerU's hybrid backend leverages VLM and OCR pipelines to effectively parse PDFs, merging VLM layout understanding with OCR accuracy for unified JSON output.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: architecture
- Published: 2026-02-22

---

**MinerU's hybrid backend automatically routes documents through either a pure Vision-Language Model (VLM) pipeline, a traditional OCR/formula pipeline, or a combination of both, merging VLM layout understanding with specialized OCR accuracy to produce a unified middle-JSON output.**

The `opendatalab/MinerU` project implements an intelligent **MinerU hybrid backend** that eliminates the trade-off between speed and accuracy in document parsing. By analyzing each PDF's characteristics at runtime, the system dynamically selects the optimal processing strategy—leveraging the holistic visual understanding of VLMs for complex layouts while falling back to high-precision OCR and formula recognition engines for text-heavy content.

## What Is the MinerU Hybrid Backend?

The **MinerU hybrid backend** is an intelligent routing system implemented in [[`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/hybrid/hybrid_analyze.py). Rather than forcing users to choose between a pure VLM approach or a traditional OCR pipeline, the backend analyzes document properties at runtime to determine the optimal processing strategy.

This architecture allows MinerU to combine the **global visual understanding** of Vision-Language Models (accurate block ordering, layout detection, visual context) with the **specialized precision** of traditional OCR and formula recognition engines (fine-grained text extraction, complex equation handling).

## Three-Step Decision Logic in Hybrid Analysis

The hybrid backend's decision engine operates through three distinct phases defined in the `doc_analyze` function:

### Step 1: Detect OCR Requirements with `ocr_classify`

First, the system determines whether the PDF requires OCR processing at all. The `ocr_classify` function in [[`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/utils/pdf_classify.py) analyzes the PDF bytes to check for searchable text layers using the underlying `classify` utility.

If the PDF contains embedded text, the system avoids unnecessary OCR overhead, routing the document through text extraction paths instead of image-based recognition.

### Step 2: Determine VLM OCR Enablement

Next, the `_should_enable_vlm_ocr` function evaluates whether to use the VLM for OCR tasks or fall back to traditional pipeline OCR. This decision considers:

- **Environment variables**: `MINERU_FORCE_VLM_OCR_ENABLE` forces VLM OCR on, while `MINERU_HYBRID_FORCE_PIPELINE_ENABLE` forces it off
- **Language settings**: VLM OCR is optimized for Chinese (`ch`) and English (`en`) content
- **Formula detection**: The `inline_formula_enable` flag indicates whether mathematical content is present

This logic resides in [[`mineru/backend/hybrid/utils.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/utils.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/hybrid/utils.py).

### Step 3: Execute the Appropriate Backend Path

Based on the previous steps, the system selects one of three execution paths:

**Pure VLM Path**: When VLM OCR is enabled, the system calls `predictor.batch_two_step_extract` from [[`mineru/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/vlm/vlm_analyze.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/vlm/vlm_analyze.py), completely bypassing the formula detection, formula recognition, and OCR pipeline stages.

**Hybrid Pipeline Path**: When VLM OCR is disabled but layout analysis is active, the system constructs a hybrid pipeline using the `HybridModelSingleton`:
- **Formula detection**: `mfd_model.batch_predict` identifies mathematical formula regions
- **Formula recognition**: `mfr_model.batch_predict` converts formulas to LaTeX
- **OCR detection**: `ocr_det` runs either per-page or in batches controlled by the `enable_ocr_det_batch` flag
- **OCR recognition**: Executes when `_ocr_enable` is true

This approach preserves the VLM's superior layout extraction (blocks, captions, tables, images) while applying specialized OCR and formula models for text-heavy regions.

## How the Hybrid Backend Merges Pipeline and VLM Results

The true innovation of the MinerU hybrid backend lies in its result-merging architecture, implemented primarily in [`hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/hybrid_analyze.py) and supporting modules.

### Alignment and Processing

When running in hybrid mode, the system generates three parallel data structures:
1. **VLM layout blocks** (`results`): Contains detected images, tables, captions, and text blocks with visual coordinates
2. **Pipeline formula data** (`inline_formula_list`): Mathematical expressions detected by specialized MFD/MFR models
3. **Pipeline OCR results** (`ocr_res_list`): Text recognized by traditional OCR engines

The `_process_ocr_and_formulas` function aligns these streams by:
- Mapping OCR detections to VLM-generated blocks using coordinate overlap
- Masking image, table, and equation areas to prevent OCR interference with visual elements
- Normalizing coordinate systems between the VLM's layout predictions and the pipeline's text detections

### Final Assembly

The `result_to_middle_json` function in [[`mineru/backend/hybrid/hybrid_model_output_to_middle_json.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_model_output_to_middle_json.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/hybrid/hybrid_model_output_to_middle_json.py) assembles the unified output structure:

- **VLM-derived blocks** retain their types (images, tables, captions) and visual bounding boxes
- **Inline formulas** from the pipeline are inserted into appropriate text blocks
- **OCR text** is attached to spans within the VLM layout hierarchy
- **Post-processing**: Optional LLM-aided title re-ranking and cross-page table merging refine the final structure

This architecture ensures that the **holistic visual understanding** of VLMs (accurate reading order, visual context, layout hierarchy) is preserved while **specialized OCR and formula engines** handle fine-grained text extraction and complex equations where they excel.

## Practical Code Examples for MinerU Hybrid Backend Configuration

The following examples demonstrate how to control the hybrid backend's behavior using Python and environment variables.

### Automatic Mode: Let MinerU Decide

By default, the hybrid backend analyzes document properties and selects the optimal processing strategy:

```python
from mineru.backend.hybrid import doc_analyze

# Load PDF bytes from file or stream

with open("document.pdf", "rb") as f:
    pdf_bytes = f.read()

# Automatic backend selection

middle_json, vlm_results, vlm_ocr_used = doc_analyze(
    pdf_bytes,
    image_writer=None,               # Optional DataWriter for image assets

    backend="transformers",          # VLM backend type

    parse_method="auto",             # Auto-detect OCR necessity

    language="ch",                   # Language influences VLM OCR enablement

    inline_formula_enable=True,      # Enable inline formula detection

)

print(f"VLM OCR was used: {vlm_ocr_used}")

```

The `vlm_ocr_used` boolean indicates whether the pure VLM path was taken (`True`) or if the hybrid pipeline with traditional OCR was employed (`False`).

### Force Pipeline Mode: Disable VLM OCR

To bypass VLM OCR and use traditional pipeline OCR and formula recognition:

```bash
export MINERU_HYBRID_FORCE_PIPELINE_ENABLE=1

```

```python
from mineru.backend.hybrid import doc_analyze

middle_json, _, _ = doc_analyze(
    pdf_bytes, 
    backend="transformers"
)

```

In this mode, the VLM still provides layout analysis (block detection, reading order), but all text recognition runs through the specialized OCR and formula pipeline models.

### Force VLM Mode: Pure Vision-Language Processing

To force the system to use only the VLM for all tasks, skipping pipeline OCR entirely:

```bash
export MINERU_FORCE_VLM_OCR_ENABLE=1

```

```python
from mineru.backend.hybrid import doc_analyze

middle_json, _, _ = doc_analyze(
    pdf_bytes,
    backend="transformers"
)

```

This configuration routes execution directly to `predictor.batch_two_step_extract` in the VLM module, bypassing formula detection and OCR pipeline stages completely.

### Inspecting the Merged Output

The `middle_json` structure contains the unified results from both processing streams:

```python
import json

# Examine paragraph blocks

print(json.dumps(
    middle_json["pdf_info"][0]["para_blocks"], 
    indent=2, 
    ensure_ascii=False
))

```

Each block contains layout information from the VLM merged with text content from the pipeline:

```json
{
  "type": "text",
  "bbox": [100, 200, 500, 300],
  "lines": [
    {
      "spans": [
        {
          "type": "text", 
          "content": "Extracted OCR text merged with VLM layout", 
          "score": 0.96
        }
      ]
    }
  ]
}

```

Blocks originating from VLM layout detection (images, tables, captions) retain their original `type` attributes while OCR spans attach as child elements, preserving the visual hierarchy.

## Key Source Files in the Hybrid Backend Architecture

The hybrid backend implementation spans several critical modules:

| File | Role |
|------|------|
| [[`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/hybrid/hybrid_analyze.py) | Main entry point; decides VLM vs. pipeline, orchestrates OCR/formula pipelines, builds middle-JSON. |
| [[`mineru/backend/hybrid/hybrid_magic_model.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_magic_model.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/hybrid/hybrid_magic_model.py) | Post-processing of VLM blocks, merges OCR/formula spans, resolves overlapping spans, builds final block lists. |
| [[`mineru/backend/hybrid/hybrid_model_output_to_middle_json.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_model_output_to_middle_json.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/hybrid/hybrid_model_output_to_middle_json.py) | Converts VLM + pipeline results into unified middle-JSON; adds LLM-aided title ranking and cross-page table merging. |
| [[`mineru/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/vlm/vlm_analyze.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/vlm/vlm_analyze.py) | VLM model singleton and `doc_analyze` for the pure VLM path. |
| [[`mineru/backend/pipeline/pipeline_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/pipeline/pipeline_analyze.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/pipeline/pipeline_analyze.py) | Pure pipeline implementation used when the Hybrid backend decides to skip VLM. |
| [[`mineru/backend/hybrid/utils.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/utils.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/backend/hybrid/utils.py) | Helper utilities including `get_batch_ratio` and `_should_enable_vlm_ocr` that govern hybrid decisions. |
| [[`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/utils/pdf_classify.py) | Determines whether a PDF needs OCR via the `classify` function. |
| [[`mineru/utils/model_utils.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/model_utils.py)](https://github.com/opendatalab/MinerU/blob/master/mineru/utils/model_utils.py) | GPU-memory based batch-size heuristics including `get_vram` and `clean_memory`. |

These modules collectively implement the **MinerU hybrid backend** that intelligently combines the **global visual understanding** of VLMs with the **specialized, high-accuracy OCR and formula engines** of the classic pipeline.

## Summary

- The **MinerU hybrid backend** automatically selects between pure VLM processing, traditional OCR/formula pipelines, or a combination based on document characteristics analyzed at runtime.
- Decision logic in [`hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/hybrid_analyze.py) uses `ocr_classify` to detect OCR needs and `_should_enable_vlm_ocr` to evaluate language settings and formula requirements.
- Environment variables `MINERU_FORCE_VLM_OCR_ENABLE` and `MINERU_HYBRID_FORCE_PIPELINE_ENABLE` allow manual override of the automatic routing logic.
- The merging process aligns VLM layout blocks with pipeline OCR and formula results through `_process_ocr_and_formulas`, producing a unified middle-JSON via `result_to_middle_json`.
- This architecture preserves VLM advantages in layout understanding and reading order while leveraging pipeline strengths in fine-grained text extraction and complex formula recognition.

## Frequently Asked Questions

### What triggers MinerU to use VLM OCR instead of the traditional pipeline?

The hybrid backend evaluates three primary factors in [`hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/hybrid_analyze.py). First, `ocr_classify` checks if the PDF contains searchable text or requires image-based OCR. Then, `_should_enable_vlm_ocr` examines the document language (optimizing for Chinese and English) and whether inline formulas are detected. Finally, environment variables can force the decision: setting `MINERU_FORCE_VLM_OCR_ENABLE=1` mandates VLM OCR, while `MINERU_HYBRID_FORCE_PIPELINE_ENABLE=1` disables it in favor of traditional OCR engines.

### Can I force MinerU to use only the VLM backend without any pipeline components?

Yes, by setting the environment variable `MINERU_FORCE_VLM_OCR_ENABLE=1` before calling `doc_analyze`. When this flag is active, the hybrid backend routes execution directly to `predictor.batch_two_step_extract` in [`mineru/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/vlm/vlm_analyze.py), completely bypassing the formula detection, formula recognition, and OCR pipeline stages. This mode relies entirely on the VLM for text, layout, and formula extraction.

### How does the hybrid backend handle complex mathematical formulas?

In hybrid mode (the default when VLM OCR is disabled but layout analysis is active), the system processes formulas through specialized pipeline models while using the VLM for layout. The `mfd_model.batch_predict` function detects formula regions, and `mfr_model.batch_predict` converts formulas to LaTeX. These results are then merged with the VLM's layout blocks via `_process_ocr_and_formulas`, which aligns formula coordinates with the VLM's reading order and attaches recognized LaTeX to the appropriate layout blocks in the final middle-JSON output.

### Where does the final document structure get assembled in the MinerU codebase?

The unified document structure is assembled in [`mineru/backend/hybrid/hybrid_model_output_to_middle_json.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_model_output_to_middle_json.py) by the `result_to_middle_json` function. This module receives the VLM layout blocks, pipeline OCR results (`ocr_res_list`), and formula recognition data (`inline_formula_list`), then stitches them into a single **middle-JSON** format. The output includes VLM-derived blocks for images and tables, OCR text attached to appropriate spans, inline formulas integrated into text blocks, and optional post-processing such as LLM-aided title re-ranking and cross-page table merging.