# PaddleOCR Data Preprocessing Best Practices: Configuration-Driven Pipeline Guide

> Master PaddleOCR data preprocessing with our guide to its configuration-driven pipeline. Learn best practices for reproducible and memory-efficient data loading.

- Repository: [PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)
- Tags: best-practices
- Published: 2026-03-03

---

**PaddleOCR data preprocessing relies on a modular, YAML-driven pipeline where operators like `DecodeImage`, `RecResizeImg`, and `MultiLabelEncode` are instantiated via `create_operators` and applied sequentially through `transform` to ensure reproducible, memory-efficient data loading for both detection and recognition tasks.**

The PaddlePaddle/PaddleOCR repository implements a flexible, configuration-based preprocessing system that separates data transformation logic from model implementation. Understanding these **PaddleOCR data preprocessing** patterns is essential for customizing training pipelines, debugging inference issues, and optimizing memory usage across different OCR tasks.

## Core Architecture of PaddleOCR Data Preprocessing

### Configuration-Driven Operator Chain

PaddleOCR defines preprocessing steps as an ordered list of operator dictionaries in YAML configuration files. In [`configs/rec/PP-OCRv5/multi_language/en_PP-OCRv5_mobile_rec.yaml`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/configs/rec/PP-OCRv5/multi_language/en_PP-OCRv5_mobile_rec.yaml), the `transforms` list declares the exact sequence: `RecConAug` → `RecAug` → `MultiLabelEncode` → `KeepKeys`. This declarative approach ensures experimental reproducibility—every training run uses identical transformations without code modification.

### The `create_operators` and `transform` Utilities

The factory function `create_operators` in [`ppocr/data/imaug/__init__.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/ppocr/data/imaug/__init__.py) (lines 79-96) parses the YAML list, dynamically imports operator classes, and instantiates them with supplied parameters. The `transform` function (lines 68-76) then sequentially applies these operators to a data dictionary, mutating it in-place through each stage.

### Operator Classes and Their Roles

Individual operators in `ppocr/data/imaug/` implement a standard `__call__(self, data)` interface:

- **Image decoding**: `DecodeImage` enforces BGR format and returns NumPy arrays with shape `H×W×C`
- **Geometric augmentation**: `RecConAug`, `RecAug`, and `ABINetRecAug` provide training-time variability
- **Resizing logic**: `RecResizeImg` (recognition), `DetResizeForTest` (detection), and `ClsResizeImg` (classification) handle dimension normalization
- **Label encoding**: `MultiLabelEncode`, `CTCLabelEncode`, and `NRTRLabelEncode` convert text labels to model-ready tensors

Dataset wrappers like `SimpleDataSet` in [`ppocr/data/simple_dataset.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/ppocr/data/simple_dataset.py) build the operator list once via `create_operators` and call `transform` for every sample during iteration.

## Recommended Preprocessing Workflow

### Stage 1: Image Decoding with `DecodeImage`

Always begin pipelines with `DecodeImage` configured for BGR mode (`img_mode: "BGR"`, `channel_first: False`). This guarantees downstream operators receive consistent NumPy arrays and eliminates color channel mismatches that corrupt feature extraction.

### Stage 2: Probabilistic Augmentation for Training

For recognition training, combine geometric and photometric augmentations with controlled probabilities (0.4–0.5). Use `RecConAug` for context-based augmentation and `RecAug` (which internally applies `BaseDataAugmentation`) for color jittering and noise. These operators reside in [`ppocr/data/imaug/rec_img_aug.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/ppocr/data/imaug/rec_img_aug.py) and should be removed entirely during inference to ensure deterministic results.

### Stage 3: Resizing and Padding Strategies

**Recognition tasks** require `RecResizeImg` with explicit `image_shape` targets (e.g., `[3, 48, 320]`). This operator computes `valid_ratio`, which CTC-based decoders require to ignore padded regions.

**Detection tasks** use `DetResizeForTest` from [`ppocr/data/imaug/operators.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/ppocr/data/imaug/operators.py). Choose between fixed `image_shape` or dynamic `limit_side_len` scaling. Set `keep_ratio=True` to preserve aspect ratios and prevent bounding box distortion, or `False` for fixed-size batching. Inference scripts like [`tools/infer/predict_det.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_det.py) demonstrate runtime configuration of these parameters.

### Stage 4: Label Encoding and Key Selection

Execute label encoding (`MultiLabelEncode` or `CTCLabelEncode`) **after** resizing because some encoders depend on final image dimensions. Conclude with `KeepKeys` to retain only necessary fields (`image`, `label_ctc`, `label_gtc`, `valid_ratio`), dropping metadata like `polys` to reduce memory pressure in dataloader workers.

## Common Pitfalls in PaddleOCR Data Preprocessing

Avoid these frequent configuration errors:

- **Mismatched `image_shape`**: Verify that YAML `image_shape` matches the model's expected `input_shape` to prevent runtime tensor mismatches in the backbone.
- **Missing `valid_ratio`**: If CTC predictions are misaligned, ensure `RecResizeImg` precedes `MultiLabelEncode` in the transform list.
- **Augmentations at test time**: Remove `RecAug` and `RecConAug` from inference configs (as done in [`tools/infer/predict_rec.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_rec.py)) to prevent accuracy degradation from random transformations.
- **Detection aspect ratio distortion**: Set `keep_ratio=True` in `DetResizeForTest` when using dynamic scaling to maintain geometric accuracy.
- **GPU memory exhaustion**: Reduce `limit_side_len` (e.g., to 736) or switch to fixed `image_shape` when processing high-resolution detection inputs on limited hardware.

## Practical Implementation Examples

### Building a Training Pipeline from Config

```python
import yaml
from ppocr.data.imaug import create_operators, transform

# Load training config transforms

with open(
    "https://raw.githubusercontent.com/PaddlePaddle/PaddleOCR/main/configs/rec/PP-OCRv5/multi_language/en_PP-OCRv5_mobile_rec.yaml",
    "r",
) as f:
    cfg = yaml.safe_load(f)

transforms_cfg = cfg["Train"]["dataset"]["transforms"]
ops = create_operators(transforms_cfg)

# Apply to sample

sample = {"image": cv2.imread("sample.jpg")}
processed = transform(sample, ops)
print(processed.keys())

# dict_keys(['image', 'label_ctc', 'label_gtc', 'length', 'valid_ratio'])

```

### Inference Preprocessing for Detection

```python
import cv2
from ppocr.data.imaug import create_operators, transform

det_transform_cfg = [
    {"DetResizeForTest": {"image_shape": [640, 640], "keep_ratio": False}}
]

det_ops = create_operators(det_transform_cfg)

def preprocess_for_det(img_path):
    data = {"image": cv2.imread(img_path)}
    data = transform(data, det_ops)
    # data["shape"] contains original dimensions and scale factors

    return data

out = preprocess_for_det("document.jpg")
print(out["shape"])  # [1024, 768, 0.625, 0.625]

```

### Custom Augmentation Chain

```python
from ppocr.data.imaug import create_operators, transform

custom_cfg = [
    {"DecodeImage": {"img_mode": "BGR", "channel_first": False}},
    {"RecConAug": {"prob": 0.7, "image_shape": [48, 320, 3], "max_text_length": 25}},
    {"RecAug": {}},
    {"RecResizeImg": {"image_shape": [3, 48, 320], "padding": True}},
    {"MultiLabelEncode": {"gtc_encode": "NRTRLabelEncode"}},
    {"KeepKeys": {"keep_keys": ["image", "label_ctc", "label_gtc", "valid_ratio"]}}
]

ops = create_operators(custom_cfg)
sample = {"image": cv2.imread("handwritten.jpg")}
processed = transform(sample, ops)

```

## Summary

- **PaddleOCR data preprocessing** uses a YAML-driven operator chain instantiated by `create_operators` and executed via `transform` in [`ppocr/data/imaug/__init__.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/ppocr/data/imaug/__init__.py).
- Always order operators as: `DecodeImage` → Augmentation (training only) → Resizing (`RecResizeImg` or `DetResizeForTest`) → Label Encoding → `KeepKeys`.
- Include `valid_ratio` for recognition tasks by using `RecResizeImg` before label encoding to support CTC decoders.
- Disable stochastic augmentations during inference to ensure deterministic, reproducible results.
- Match `image_shape` configurations between preprocessing YAMLs and model input requirements to prevent runtime errors.

## Frequently Asked Questions

### What is the correct order of operators in a PaddleOCR preprocessing pipeline?

The mandatory sequence starts with `DecodeImage` to normalize color channels, followed by optional training augmentations (`RecAug`, `RecConAug`), then resizing operators (`RecResizeImg` for recognition or `DetResizeForTest` for detection), followed by label encoders (`MultiLabelEncode`), and finally `KeepKeys` to filter the output dictionary. This ordering ensures that label encoders receive final image dimensions and that `valid_ratio` is computed before encoding.

### How do I prevent out-of-memory errors during detection preprocessing?

Reduce the `limit_side_len` parameter in `DetResizeForTest` (located in [`ppocr/data/imaug/operators.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/ppocr/data/imaug/operators.py)) to a moderate value such as 736, or switch from dynamic resizing to a fixed `image_shape`. This controls the maximum spatial dimensions of input tensors before they reach the GPU, preventing allocation failures on hardware with limited VRAM.

### Why is my CTC decoder producing incorrect sequence lengths?

This occurs when the `valid_ratio` field is missing from the data dictionary. Ensure that `RecResizeImg` (which calculates and injects `valid_ratio`) appears **before** `MultiLabelEncode` or `CTCLabelEncode` in the YAML transform list. The encoder uses this ratio to distinguish between valid content and padded regions in variable-length sequences.

### Should I use the same transforms for training and inference?

No. Training configs should include probabilistic augmentations like `RecConAug` and `RecAug` to improve generalization, while inference configs must exclude these stochastic operators entirely. Inference pipelines should use only deterministic operators: `DecodeImage`, resizing operators, and `KeepKeys`, ensuring that evaluation metrics reflect model performance without random variation.