PaddleOCR Data Preprocessing Best Practices: Configuration-Driven Pipeline Guide
PaddleOCR data preprocessing relies on a modular, YAML-driven pipeline where operators like DecodeImage, RecResizeImg, and MultiLabelEncode are instantiated via create_operators and applied sequentially through transform to ensure reproducible, memory-efficient data loading for both detection and recognition tasks.
The PaddlePaddle/PaddleOCR repository implements a flexible, configuration-based preprocessing system that separates data transformation logic from model implementation. Understanding these PaddleOCR data preprocessing patterns is essential for customizing training pipelines, debugging inference issues, and optimizing memory usage across different OCR tasks.
Core Architecture of PaddleOCR Data Preprocessing
Configuration-Driven Operator Chain
PaddleOCR defines preprocessing steps as an ordered list of operator dictionaries in YAML configuration files. In configs/rec/PP-OCRv5/multi_language/en_PP-OCRv5_mobile_rec.yaml, the transforms list declares the exact sequence: RecConAug → RecAug → MultiLabelEncode → KeepKeys. This declarative approach ensures experimental reproducibility—every training run uses identical transformations without code modification.
The create_operators and transform Utilities
The factory function create_operators in ppocr/data/imaug/__init__.py (lines 79-96) parses the YAML list, dynamically imports operator classes, and instantiates them with supplied parameters. The transform function (lines 68-76) then sequentially applies these operators to a data dictionary, mutating it in-place through each stage.
Operator Classes and Their Roles
Individual operators in ppocr/data/imaug/ implement a standard __call__(self, data) interface:
- Image decoding:
DecodeImageenforces BGR format and returns NumPy arrays with shapeH×W×C - Geometric augmentation:
RecConAug,RecAug, andABINetRecAugprovide training-time variability - Resizing logic:
RecResizeImg(recognition),DetResizeForTest(detection), andClsResizeImg(classification) handle dimension normalization - Label encoding:
MultiLabelEncode,CTCLabelEncode, andNRTRLabelEncodeconvert text labels to model-ready tensors
Dataset wrappers like SimpleDataSet in ppocr/data/simple_dataset.py build the operator list once via create_operators and call transform for every sample during iteration.
Recommended Preprocessing Workflow
Stage 1: Image Decoding with DecodeImage
Always begin pipelines with DecodeImage configured for BGR mode (img_mode: "BGR", channel_first: False). This guarantees downstream operators receive consistent NumPy arrays and eliminates color channel mismatches that corrupt feature extraction.
Stage 2: Probabilistic Augmentation for Training
For recognition training, combine geometric and photometric augmentations with controlled probabilities (0.4–0.5). Use RecConAug for context-based augmentation and RecAug (which internally applies BaseDataAugmentation) for color jittering and noise. These operators reside in ppocr/data/imaug/rec_img_aug.py and should be removed entirely during inference to ensure deterministic results.
Stage 3: Resizing and Padding Strategies
Recognition tasks require RecResizeImg with explicit image_shape targets (e.g., [3, 48, 320]). This operator computes valid_ratio, which CTC-based decoders require to ignore padded regions.
Detection tasks use DetResizeForTest from ppocr/data/imaug/operators.py. Choose between fixed image_shape or dynamic limit_side_len scaling. Set keep_ratio=True to preserve aspect ratios and prevent bounding box distortion, or False for fixed-size batching. Inference scripts like tools/infer/predict_det.py demonstrate runtime configuration of these parameters.
Stage 4: Label Encoding and Key Selection
Execute label encoding (MultiLabelEncode or CTCLabelEncode) after resizing because some encoders depend on final image dimensions. Conclude with KeepKeys to retain only necessary fields (image, label_ctc, label_gtc, valid_ratio), dropping metadata like polys to reduce memory pressure in dataloader workers.
Common Pitfalls in PaddleOCR Data Preprocessing
Avoid these frequent configuration errors:
- Mismatched
image_shape: Verify that YAMLimage_shapematches the model's expectedinput_shapeto prevent runtime tensor mismatches in the backbone. - Missing
valid_ratio: If CTC predictions are misaligned, ensureRecResizeImgprecedesMultiLabelEncodein the transform list. - Augmentations at test time: Remove
RecAugandRecConAugfrom inference configs (as done intools/infer/predict_rec.py) to prevent accuracy degradation from random transformations. - Detection aspect ratio distortion: Set
keep_ratio=TrueinDetResizeForTestwhen using dynamic scaling to maintain geometric accuracy. - GPU memory exhaustion: Reduce
limit_side_len(e.g., to 736) or switch to fixedimage_shapewhen processing high-resolution detection inputs on limited hardware.
Practical Implementation Examples
Building a Training Pipeline from Config
import yaml
from ppocr.data.imaug import create_operators, transform
# Load training config transforms
with open(
"https://raw.githubusercontent.com/PaddlePaddle/PaddleOCR/main/configs/rec/PP-OCRv5/multi_language/en_PP-OCRv5_mobile_rec.yaml",
"r",
) as f:
cfg = yaml.safe_load(f)
transforms_cfg = cfg["Train"]["dataset"]["transforms"]
ops = create_operators(transforms_cfg)
# Apply to sample
sample = {"image": cv2.imread("sample.jpg")}
processed = transform(sample, ops)
print(processed.keys())
# dict_keys(['image', 'label_ctc', 'label_gtc', 'length', 'valid_ratio'])
Inference Preprocessing for Detection
import cv2
from ppocr.data.imaug import create_operators, transform
det_transform_cfg = [
{"DetResizeForTest": {"image_shape": [640, 640], "keep_ratio": False}}
]
det_ops = create_operators(det_transform_cfg)
def preprocess_for_det(img_path):
data = {"image": cv2.imread(img_path)}
data = transform(data, det_ops)
# data["shape"] contains original dimensions and scale factors
return data
out = preprocess_for_det("document.jpg")
print(out["shape"]) # [1024, 768, 0.625, 0.625]
Custom Augmentation Chain
from ppocr.data.imaug import create_operators, transform
custom_cfg = [
{"DecodeImage": {"img_mode": "BGR", "channel_first": False}},
{"RecConAug": {"prob": 0.7, "image_shape": [48, 320, 3], "max_text_length": 25}},
{"RecAug": {}},
{"RecResizeImg": {"image_shape": [3, 48, 320], "padding": True}},
{"MultiLabelEncode": {"gtc_encode": "NRTRLabelEncode"}},
{"KeepKeys": {"keep_keys": ["image", "label_ctc", "label_gtc", "valid_ratio"]}}
]
ops = create_operators(custom_cfg)
sample = {"image": cv2.imread("handwritten.jpg")}
processed = transform(sample, ops)
Summary
- PaddleOCR data preprocessing uses a YAML-driven operator chain instantiated by
create_operatorsand executed viatransforminppocr/data/imaug/__init__.py. - Always order operators as:
DecodeImage→ Augmentation (training only) → Resizing (RecResizeImgorDetResizeForTest) → Label Encoding →KeepKeys. - Include
valid_ratiofor recognition tasks by usingRecResizeImgbefore label encoding to support CTC decoders. - Disable stochastic augmentations during inference to ensure deterministic, reproducible results.
- Match
image_shapeconfigurations between preprocessing YAMLs and model input requirements to prevent runtime errors.
Frequently Asked Questions
What is the correct order of operators in a PaddleOCR preprocessing pipeline?
The mandatory sequence starts with DecodeImage to normalize color channels, followed by optional training augmentations (RecAug, RecConAug), then resizing operators (RecResizeImg for recognition or DetResizeForTest for detection), followed by label encoders (MultiLabelEncode), and finally KeepKeys to filter the output dictionary. This ordering ensures that label encoders receive final image dimensions and that valid_ratio is computed before encoding.
How do I prevent out-of-memory errors during detection preprocessing?
Reduce the limit_side_len parameter in DetResizeForTest (located in ppocr/data/imaug/operators.py) to a moderate value such as 736, or switch from dynamic resizing to a fixed image_shape. This controls the maximum spatial dimensions of input tensors before they reach the GPU, preventing allocation failures on hardware with limited VRAM.
Why is my CTC decoder producing incorrect sequence lengths?
This occurs when the valid_ratio field is missing from the data dictionary. Ensure that RecResizeImg (which calculates and injects valid_ratio) appears before MultiLabelEncode or CTCLabelEncode in the YAML transform list. The encoder uses this ratio to distinguish between valid content and padded regions in variable-length sequences.
Should I use the same transforms for training and inference?
No. Training configs should include probabilistic augmentations like RecConAug and RecAug to improve generalization, while inference configs must exclude these stochastic operators entirely. Inference pipelines should use only deterministic operators: DecodeImage, resizing operators, and KeepKeys, ensuring that evaluation metrics reflect model performance without random variation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →