PaddleOCR Data Preprocessing Best Practices: Configuration-Driven Pipeline Guide

PaddleOCR data preprocessing relies on a modular, YAML-driven pipeline where operators like DecodeImage, RecResizeImg, and MultiLabelEncode are instantiated via create_operators and applied sequentially through transform to ensure reproducible, memory-efficient data loading for both detection and recognition tasks.

The PaddlePaddle/PaddleOCR repository implements a flexible, configuration-based preprocessing system that separates data transformation logic from model implementation. Understanding these PaddleOCR data preprocessing patterns is essential for customizing training pipelines, debugging inference issues, and optimizing memory usage across different OCR tasks.

Core Architecture of PaddleOCR Data Preprocessing

Configuration-Driven Operator Chain

PaddleOCR defines preprocessing steps as an ordered list of operator dictionaries in YAML configuration files. In configs/rec/PP-OCRv5/multi_language/en_PP-OCRv5_mobile_rec.yaml, the transforms list declares the exact sequence: RecConAug → RecAug → MultiLabelEncode → KeepKeys. This declarative approach ensures experimental reproducibility—every training run uses identical transformations without code modification.

The create_operators and transform Utilities

The factory function create_operators in ppocr/data/imaug/__init__.py (lines 79-96) parses the YAML list, dynamically imports operator classes, and instantiates them with supplied parameters. The transform function (lines 68-76) then sequentially applies these operators to a data dictionary, mutating it in-place through each stage.

Operator Classes and Their Roles

Individual operators in ppocr/data/imaug/ implement a standard __call__(self, data) interface:

  • Image decoding: DecodeImage enforces BGR format and returns NumPy arrays with shape H×W×C
  • Geometric augmentation: RecConAug, RecAug, and ABINetRecAug provide training-time variability
  • Resizing logic: RecResizeImg (recognition), DetResizeForTest (detection), and ClsResizeImg (classification) handle dimension normalization
  • Label encoding: MultiLabelEncode, CTCLabelEncode, and NRTRLabelEncode convert text labels to model-ready tensors

Dataset wrappers like SimpleDataSet in ppocr/data/simple_dataset.py build the operator list once via create_operators and call transform for every sample during iteration.

Stage 1: Image Decoding with DecodeImage

Always begin pipelines with DecodeImage configured for BGR mode (img_mode: "BGR", channel_first: False). This guarantees downstream operators receive consistent NumPy arrays and eliminates color channel mismatches that corrupt feature extraction.

Stage 2: Probabilistic Augmentation for Training

For recognition training, combine geometric and photometric augmentations with controlled probabilities (0.4–0.5). Use RecConAug for context-based augmentation and RecAug (which internally applies BaseDataAugmentation) for color jittering and noise. These operators reside in ppocr/data/imaug/rec_img_aug.py and should be removed entirely during inference to ensure deterministic results.

Stage 3: Resizing and Padding Strategies

Recognition tasks require RecResizeImg with explicit image_shape targets (e.g., [3, 48, 320]). This operator computes valid_ratio, which CTC-based decoders require to ignore padded regions.

Detection tasks use DetResizeForTest from ppocr/data/imaug/operators.py. Choose between fixed image_shape or dynamic limit_side_len scaling. Set keep_ratio=True to preserve aspect ratios and prevent bounding box distortion, or False for fixed-size batching. Inference scripts like tools/infer/predict_det.py demonstrate runtime configuration of these parameters.

Stage 4: Label Encoding and Key Selection

Execute label encoding (MultiLabelEncode or CTCLabelEncode) after resizing because some encoders depend on final image dimensions. Conclude with KeepKeys to retain only necessary fields (image, label_ctc, label_gtc, valid_ratio), dropping metadata like polys to reduce memory pressure in dataloader workers.

Common Pitfalls in PaddleOCR Data Preprocessing

Avoid these frequent configuration errors:

  • Mismatched image_shape: Verify that YAML image_shape matches the model's expected input_shape to prevent runtime tensor mismatches in the backbone.
  • Missing valid_ratio: If CTC predictions are misaligned, ensure RecResizeImg precedes MultiLabelEncode in the transform list.
  • Augmentations at test time: Remove RecAug and RecConAug from inference configs (as done in tools/infer/predict_rec.py) to prevent accuracy degradation from random transformations.
  • Detection aspect ratio distortion: Set keep_ratio=True in DetResizeForTest when using dynamic scaling to maintain geometric accuracy.
  • GPU memory exhaustion: Reduce limit_side_len (e.g., to 736) or switch to fixed image_shape when processing high-resolution detection inputs on limited hardware.

Practical Implementation Examples

Building a Training Pipeline from Config

import yaml
from ppocr.data.imaug import create_operators, transform

# Load training config transforms

with open(
    "https://raw.githubusercontent.com/PaddlePaddle/PaddleOCR/main/configs/rec/PP-OCRv5/multi_language/en_PP-OCRv5_mobile_rec.yaml",
    "r",
) as f:
    cfg = yaml.safe_load(f)

transforms_cfg = cfg["Train"]["dataset"]["transforms"]
ops = create_operators(transforms_cfg)

# Apply to sample

sample = {"image": cv2.imread("sample.jpg")}
processed = transform(sample, ops)
print(processed.keys())

# dict_keys(['image', 'label_ctc', 'label_gtc', 'length', 'valid_ratio'])

Inference Preprocessing for Detection

import cv2
from ppocr.data.imaug import create_operators, transform

det_transform_cfg = [
    {"DetResizeForTest": {"image_shape": [640, 640], "keep_ratio": False}}
]

det_ops = create_operators(det_transform_cfg)

def preprocess_for_det(img_path):
    data = {"image": cv2.imread(img_path)}
    data = transform(data, det_ops)
    # data["shape"] contains original dimensions and scale factors

    return data

out = preprocess_for_det("document.jpg")
print(out["shape"])  # [1024, 768, 0.625, 0.625]

Custom Augmentation Chain

from ppocr.data.imaug import create_operators, transform

custom_cfg = [
    {"DecodeImage": {"img_mode": "BGR", "channel_first": False}},
    {"RecConAug": {"prob": 0.7, "image_shape": [48, 320, 3], "max_text_length": 25}},
    {"RecAug": {}},
    {"RecResizeImg": {"image_shape": [3, 48, 320], "padding": True}},
    {"MultiLabelEncode": {"gtc_encode": "NRTRLabelEncode"}},
    {"KeepKeys": {"keep_keys": ["image", "label_ctc", "label_gtc", "valid_ratio"]}}
]

ops = create_operators(custom_cfg)
sample = {"image": cv2.imread("handwritten.jpg")}
processed = transform(sample, ops)

Summary

  • PaddleOCR data preprocessing uses a YAML-driven operator chain instantiated by create_operators and executed via transform in ppocr/data/imaug/__init__.py.
  • Always order operators as: DecodeImage → Augmentation (training only) → Resizing (RecResizeImg or DetResizeForTest) → Label Encoding → KeepKeys.
  • Include valid_ratio for recognition tasks by using RecResizeImg before label encoding to support CTC decoders.
  • Disable stochastic augmentations during inference to ensure deterministic, reproducible results.
  • Match image_shape configurations between preprocessing YAMLs and model input requirements to prevent runtime errors.

Frequently Asked Questions

What is the correct order of operators in a PaddleOCR preprocessing pipeline?

The mandatory sequence starts with DecodeImage to normalize color channels, followed by optional training augmentations (RecAug, RecConAug), then resizing operators (RecResizeImg for recognition or DetResizeForTest for detection), followed by label encoders (MultiLabelEncode), and finally KeepKeys to filter the output dictionary. This ordering ensures that label encoders receive final image dimensions and that valid_ratio is computed before encoding.

How do I prevent out-of-memory errors during detection preprocessing?

Reduce the limit_side_len parameter in DetResizeForTest (located in ppocr/data/imaug/operators.py) to a moderate value such as 736, or switch from dynamic resizing to a fixed image_shape. This controls the maximum spatial dimensions of input tensors before they reach the GPU, preventing allocation failures on hardware with limited VRAM.

Why is my CTC decoder producing incorrect sequence lengths?

This occurs when the valid_ratio field is missing from the data dictionary. Ensure that RecResizeImg (which calculates and injects valid_ratio) appears before MultiLabelEncode or CTCLabelEncode in the YAML transform list. The encoder uses this ratio to distinguish between valid content and padded regions in variable-length sequences.

Should I use the same transforms for training and inference?

No. Training configs should include probabilistic augmentations like RecConAug and RecAug to improve generalization, while inference configs must exclude these stochastic operators entirely. Inference pipelines should use only deterministic operators: DecodeImage, resizing operators, and KeepKeys, ensuring that evaluation metrics reflect model performance without random variation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →