How MinerU Performs OCR on Scanned PDFs: A Deep Dive into the Pipeline
MinerU extracts text from scanned PDFs by rendering pages to high-resolution images, detecting text regions with a PyTorch-Paddle OCR model, and integrating the recognized text into the document's layout structure.
MinerU is an open-source document parsing toolkit developed by OpenDataLab that handles complex PDF extraction tasks. When processing scanned documents that contain only images, MinerU activates its OCR pipeline to recover machine-readable text. This article examines the complete technical implementation, from initial classification through final text integration.
OCR Pipeline Overview
The OCR process in MinerU follows a deterministic multi-stage pipeline. First, the system classifies whether the PDF requires OCR processing. Then it rasterizes pages into images, detects text regions using a deep learning model, refines bounding boxes, performs text recognition, and finally packages results into MinerU's internal annotation format.
This architecture separates concerns between layout analysis (handled by a Vision Language Model) and text recognition (handled by PaddleOCR), allowing each component to specialize while maintaining coordinate consistency across the pipeline.
Step-by-Step OCR Process
1. Detecting OCR Requirements
Before processing, MinerU determines whether a PDF contains scannable text or is already text-based. The ocr_classify function in mineru/backend/hybrid/hybrid_analyze.py examines the document using the classify method from mineru/utils/pdf_classify.py.
This classification step prevents unnecessary OCR processing on native PDFs, saving computational resources and preserving original text formatting.
2. Rendering PDF Pages to Images
When OCR is required, MinerU converts PDF pages into raster images using pdf_to_images in mineru/utils/pdf_reader.py. This function leverages pypdfium2 to render pages at a configurable DPI (typically 200 DPI), returning a list of PIL.Image objects.
The rasterization quality directly impacts OCR accuracy, as higher resolution preserves fine text details while maintaining reasonable processing speeds.
3. Preparing Text Region Crops
For each rendered page, the ocr_det function (in hybrid_analyze.py) extracts bounding boxes of text-type blocks identified by the Vision Language Model. Using crop_img from model_utils.py, it creates padded crops of these regions to feed into the OCR model.
This cropping strategy focuses computational effort on relevant areas while filtering out non-text elements like images or graphics.
4. Detecting Text Regions
Inside the PytorchPaddleOCR class in mineru/model/ocr/pytorch_paddle.py, the text_detector predicts quadrilateral bounding boxes (dt_boxes) around text lines. The ocr method orchestrates this detection phase, returning geometric coordinates for all text instances in the cropped region.
The detection model handles multi-angle text and varying font sizes, producing precise localization even for densely packed document layouts.
5. Merging and Refining Bounding Boxes
Detected boxes undergo post-processing to improve readability. The merge_det_boxes function in mineru/utils/ocr_utils.py sorts and merges neighboring boxes to join fragmented text spans. If formula detection results (mfd_res) are available, update_det_boxes adjusts the text regions to avoid overlapping with mathematical expressions.
This refinement ensures that text lines remain coherent and properly separated from specialized content like equations.
6. Rotating and Cropping Regions
Each refined bounding box is normalized to a straight-axis image using get_rotate_crop_image in ocr_utils.py. This function rotates text regions to horizontal orientation, standardizing the input for the recognition phase regardless of original text angle.
Proper rotation correction significantly improves recognition accuracy for skewed or curved text common in scanned documents.
7. Recognizing Text Content
The text_recognizer in PytorchPaddleOCR processes the rotated crops, producing (text, score) pairs for each region. The __call__ method in pytorch_paddle.py filters low-confidence results using OcrConfidence thresholds, discarding ambiguous predictions.
This recognition stage converts visual text into Unicode strings while maintaining confidence metrics for quality control.
8. Packaging OCR Results
Raw OCR output is converted to MinerU's internal format by get_ocr_result_list in ocr_utils.py. This function creates structured annotations with category_id: 15 (text), polygon coordinates, confidence scores, and language tags.
Standardized packaging enables seamless integration with the layout analysis results from earlier pipeline stages.
9. Integrating with Layout Analysis
Finally, OCR results are normalized back to the original PDF coordinate space and merged with VLM-detected layout blocks. The _process_ocr_and_formulas function in hybrid_analyze.py orchestrates this integration, producing the middle-JSON representation that feeds into downstream document conversion.
This unified output preserves the document's structural hierarchy while replacing image regions with recognized text content.
Key Implementation Files
| File | Role |
|---|---|
mineru/backend/hybrid/hybrid_analyze.py |
Orchestrates OCR detection, cropping, and result integration |
mineru/utils/pdf_reader.py |
Renders PDF pages to images using pypdfium2 |
mineru/model/ocr/pytorch_paddle.py |
Wraps PaddleOCR detector and recognizer |
mineru/utils/ocr_utils.py |
Handles box merging, rotation, and result formatting |
mineru/utils/pdf_image_tools.py |
High-performance multi-process PDF-to-image conversion |
mineru/utils/pdf_classify.py |
Determines OCR necessity via document classification |
Practical Code Examples
The following snippets demonstrate how to interact with MinerU's OCR pipeline programmatically:
# Determine whether OCR is required for a PDF
from mineru.backend.hybrid.hybrid_analyze import ocr_classify
need_ocr = ocr_classify(pdf_bytes, parse_method="auto")
# Convert PDF pages to PIL images for OCR processing
from mineru.utils.pdf_reader import pdf_to_images
pages, _ = pdf_to_images(pdf_bytes, dpi=200)
# Run the full OCR detection pipeline on rendered pages
from mineru.backend.hybrid.hybrid_analyze import _process_ocr_and_formulas
inline_formulas, ocr_res_list, pipeline = _process_ocr_and_formulas(
images_pil_list=pages,
results=vllm_results, # Layout blocks from the VLM
language="ch", # Target language for OCR
inline_formula_enable=False,
_ocr_enable=need_ocr,
batch_radio=1,
)
Each entry in ocr_res_list contains structured text data:
# Example OCR result entry
{
'category_id': 15, # Text category identifier
'poly': [x1, y1, x2, y2, ...], # Polygon coordinates
'score': 0.96, # Recognition confidence
'text': '识别文字', # Recognized text content
'lang': 'ch' # Language code
}
Summary
- Classification First: MinerU uses
ocr_classifyto avoid unnecessary processing on native text PDFs. - Image Rendering: The
pdf_to_imagesfunction rasterizes pages at configurable DPI usingpypdfium2. - Two-Stage OCR: The pipeline separates detection (
text_detectorfor bounding boxes) from recognition (text_recognizerfor text content). - Box Refinement: Functions like
merge_det_boxesandupdate_det_boxesconsolidate fragmented text and handle formula overlaps. - Standardized Output:
get_ocr_result_listconverts raw OCR data into MinerU's internal JSON format with category IDs and confidence scores.
Frequently Asked Questions
How does MinerU decide whether to run OCR on a PDF?
MinerU analyzes the PDF structure using the ocr_classify function in mineru/backend/hybrid/hybrid_analyze.py, which internally calls classify from mineru/utils/pdf_classify.py. This classification examines the document to determine if it contains scannable image content or existing text layers, ensuring OCR runs only when necessary.
What OCR engine does MinerU use for text recognition?
MinerU implements a PyTorch-based wrapper around PaddleOCR, located in mineru/model/ocr/pytorch_paddle.py. The PytorchPaddleOCR class orchestrates a two-stage process: first, text_detector locates text regions, then text_recognizer converts the cropped regions into Unicode strings with confidence scores.
How does MinerU handle rotated or skewed text in scanned documents?
The pipeline includes geometric correction through get_rotate_crop_image in mineru/utils/ocr_utils.py. After the detector identifies quadrilateral bounding boxes (dt_boxes), this function rotates and crops each region to horizontal orientation before passing it to the recognizer, ensuring accurate text extraction regardless of original page skew.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →