# How MinerU Performs OCR on Scanned PDFs: A Deep Dive into the Pipeline

> Discover how MinerU performs OCR on scanned PDFs. Learn about its pipeline for text extraction, region detection, and layout integration. Explore the opendatalab/MinerU repository for details.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: deep-dive
- Published: 2026-02-23

---

**MinerU extracts text from scanned PDFs by rendering pages to high-resolution images, detecting text regions with a PyTorch-Paddle OCR model, and integrating the recognized text into the document's layout structure.**

MinerU is an open-source document parsing toolkit developed by OpenDataLab that handles complex PDF extraction tasks. When processing scanned documents that contain only images, MinerU activates its OCR pipeline to recover machine-readable text. This article examines the complete technical implementation, from initial classification through final text integration.

## OCR Pipeline Overview

The OCR process in MinerU follows a deterministic multi-stage pipeline. First, the system classifies whether the PDF requires OCR processing. Then it rasterizes pages into images, detects text regions using a deep learning model, refines bounding boxes, performs text recognition, and finally packages results into MinerU's internal annotation format.

This architecture separates concerns between layout analysis (handled by a Vision Language Model) and text recognition (handled by PaddleOCR), allowing each component to specialize while maintaining coordinate consistency across the pipeline.

## Step-by-Step OCR Process

### 1. Detecting OCR Requirements

Before processing, MinerU determines whether a PDF contains scannable text or is already text-based. The `ocr_classify` function in [`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py) examines the document using the `classify` method from [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py).

This classification step prevents unnecessary OCR processing on native PDFs, saving computational resources and preserving original text formatting.

### 2. Rendering PDF Pages to Images

When OCR is required, MinerU converts PDF pages into raster images using `pdf_to_images` in [`mineru/utils/pdf_reader.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_reader.py). This function leverages `pypdfium2` to render pages at a configurable DPI (typically 200 DPI), returning a list of `PIL.Image` objects.

The rasterization quality directly impacts OCR accuracy, as higher resolution preserves fine text details while maintaining reasonable processing speeds.

### 3. Preparing Text Region Crops

For each rendered page, the `ocr_det` function (in [`hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/hybrid_analyze.py)) extracts bounding boxes of text-type blocks identified by the Vision Language Model. Using `crop_img` from [`model_utils.py`](https://github.com/opendatalab/MinerU/blob/main/model_utils.py), it creates padded crops of these regions to feed into the OCR model.

This cropping strategy focuses computational effort on relevant areas while filtering out non-text elements like images or graphics.

### 4. Detecting Text Regions

Inside the `PytorchPaddleOCR` class in [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py), the `text_detector` predicts quadrilateral bounding boxes (`dt_boxes`) around text lines. The `ocr` method orchestrates this detection phase, returning geometric coordinates for all text instances in the cropped region.

The detection model handles multi-angle text and varying font sizes, producing precise localization even for densely packed document layouts.

### 5. Merging and Refining Bounding Boxes

Detected boxes undergo post-processing to improve readability. The `merge_det_boxes` function in [`mineru/utils/ocr_utils.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/ocr_utils.py) sorts and merges neighboring boxes to join fragmented text spans. If formula detection results (`mfd_res`) are available, `update_det_boxes` adjusts the text regions to avoid overlapping with mathematical expressions.

This refinement ensures that text lines remain coherent and properly separated from specialized content like equations.

### 6. Rotating and Cropping Regions

Each refined bounding box is normalized to a straight-axis image using `get_rotate_crop_image` in [`ocr_utils.py`](https://github.com/opendatalab/MinerU/blob/main/ocr_utils.py). This function rotates text regions to horizontal orientation, standardizing the input for the recognition phase regardless of original text angle.

Proper rotation correction significantly improves recognition accuracy for skewed or curved text common in scanned documents.

### 7. Recognizing Text Content

The `text_recognizer` in `PytorchPaddleOCR` processes the rotated crops, producing `(text, score)` pairs for each region. The `__call__` method in [`pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/pytorch_paddle.py) filters low-confidence results using `OcrConfidence` thresholds, discarding ambiguous predictions.

This recognition stage converts visual text into Unicode strings while maintaining confidence metrics for quality control.

### 8. Packaging OCR Results

Raw OCR output is converted to MinerU's internal format by `get_ocr_result_list` in [`ocr_utils.py`](https://github.com/opendatalab/MinerU/blob/main/ocr_utils.py). This function creates structured annotations with `category_id: 15` (text), polygon coordinates, confidence scores, and language tags.

Standardized packaging enables seamless integration with the layout analysis results from earlier pipeline stages.

### 9. Integrating with Layout Analysis

Finally, OCR results are normalized back to the original PDF coordinate space and merged with VLM-detected layout blocks. The `_process_ocr_and_formulas` function in [`hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/hybrid_analyze.py) orchestrates this integration, producing the middle-JSON representation that feeds into downstream document conversion.

This unified output preserves the document's structural hierarchy while replacing image regions with recognized text content.

## Key Implementation Files

| File | Role |
|------|------|
| [`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py) | Orchestrates OCR detection, cropping, and result integration |
| [`mineru/utils/pdf_reader.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_reader.py) | Renders PDF pages to images using pypdfium2 |
| [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py) | Wraps PaddleOCR detector and recognizer |
| [`mineru/utils/ocr_utils.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/ocr_utils.py) | Handles box merging, rotation, and result formatting |
| [`mineru/utils/pdf_image_tools.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_image_tools.py) | High-performance multi-process PDF-to-image conversion |
| [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py) | Determines OCR necessity via document classification |

## Practical Code Examples

The following snippets demonstrate how to interact with MinerU's OCR pipeline programmatically:

```python

# Determine whether OCR is required for a PDF

from mineru.backend.hybrid.hybrid_analyze import ocr_classify
need_ocr = ocr_classify(pdf_bytes, parse_method="auto")

```

```python

# Convert PDF pages to PIL images for OCR processing

from mineru.utils.pdf_reader import pdf_to_images
pages, _ = pdf_to_images(pdf_bytes, dpi=200)

```

```python

# Run the full OCR detection pipeline on rendered pages

from mineru.backend.hybrid.hybrid_analyze import _process_ocr_and_formulas

inline_formulas, ocr_res_list, pipeline = _process_ocr_and_formulas(
    images_pil_list=pages,
    results=vllm_results,          # Layout blocks from the VLM

    language="ch",                 # Target language for OCR

    inline_formula_enable=False,
    _ocr_enable=need_ocr,
    batch_radio=1,
)

```

Each entry in `ocr_res_list` contains structured text data:

```python

# Example OCR result entry

{
    'category_id': 15,           # Text category identifier

    'poly': [x1, y1, x2, y2, ...], # Polygon coordinates

    'score': 0.96,                # Recognition confidence

    'text': '识别文字',            # Recognized text content

    'lang': 'ch'                  # Language code

}

```

## Summary

- **Classification First**: MinerU uses `ocr_classify` to avoid unnecessary processing on native text PDFs.
- **Image Rendering**: The `pdf_to_images` function rasterizes pages at configurable DPI using `pypdfium2`.
- **Two-Stage OCR**: The pipeline separates detection (`text_detector` for bounding boxes) from recognition (`text_recognizer` for text content).
- **Box Refinement**: Functions like `merge_det_boxes` and `update_det_boxes` consolidate fragmented text and handle formula overlaps.
- **Standardized Output**: `get_ocr_result_list` converts raw OCR data into MinerU's internal JSON format with category IDs and confidence scores.

## Frequently Asked Questions

### How does MinerU decide whether to run OCR on a PDF?

MinerU analyzes the PDF structure using the `ocr_classify` function in [`mineru/backend/hybrid/hybrid_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/backend/hybrid/hybrid_analyze.py), which internally calls `classify` from [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py). This classification examines the document to determine if it contains scannable image content or existing text layers, ensuring OCR runs only when necessary.

### What OCR engine does MinerU use for text recognition?

MinerU implements a PyTorch-based wrapper around PaddleOCR, located in [`mineru/model/ocr/pytorch_paddle.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/model/ocr/pytorch_paddle.py). The `PytorchPaddleOCR` class orchestrates a two-stage process: first, `text_detector` locates text regions, then `text_recognizer` converts the cropped regions into Unicode strings with confidence scores.

### How does MinerU handle rotated or skewed text in scanned documents?

The pipeline includes geometric correction through `get_rotate_crop_image` in [`mineru/utils/ocr_utils.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/ocr_utils.py). After the detector identifies quadrilateral bounding boxes (`dt_boxes`), this function rotates and crops each region to horizontal orientation before passing it to the recognizer, ensuring accurate text extraction regardless of original page skew.