# How to Perform Batch Processing with PaddleOCR: A Complete Guide to Batch Recognition

> Master batch processing with PaddleOCR using the TextRecognizer class and rec_batch_num parameter. Accelerate your text recognition by processing multiple images simultaneously for enhanced efficiency.

- Repository: [PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)
- Tags: how-to-guide
- Published: 2026-03-03

---

**PaddleOCR accelerates text recognition by processing multiple images simultaneously through the `TextRecognizer` class using the `--rec_batch_num` parameter, while the detection stage remains strictly single-image.**

Batch processing dramatically reduces inference latency when processing large volumes of text images. In the PaddlePaddle/PaddleOCR repository, batch inference is implemented specifically within the recognition pipeline, allowing efficient parallel processing of cropped text line images while maintaining compatibility with both Paddle Inference and ONNX runtimes.

## Understanding the Batch Processing Architecture

PaddleOCR separates its two-stage OCR pipeline into distinct batching behaviors. The **detection stage** (`TextDetector` in [`tools/infer/predict_det.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_det.py)) processes images individually because its predictor architecture expects a batch size of 1. Conversely, the **recognition stage** (`TextRecognizer` in [`tools/infer/predict_rec.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_rec.py)) fully supports batched inference, making it the primary target for throughput optimization.

The entry point for batch configuration resides in [`tools/infer/utility.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/utility.py), where the argument parser defines the `--rec_batch_num` flag. This parameter controls exactly how many text line images enter the recognition network simultaneously:

```python

# tools/infer/utility.py

parser.add_argument("--rec_batch_num", type=int, default=6, help="...")

```

## How Batch Processing Works in TextRecognizer

The `TextRecognizer.__call__` method orchestrates the entire batch workflow, from raw images to final text predictions. Understanding this internal pipeline is essential for optimizing performance.

### Step 1: Predictor Initialization

First, `utility.create_predictor` constructs the inference engine for `mode="rec"`. This factory function returns a predictor handle, an input tensor reference, and a list of output tensors configured for the specific recognition model (CTC-based, SRN, or others).

### Step 2: Dynamic Batch Preprocessing

Before inference, the input list undergoes intelligent batching:

1. **Aspect Ratio Sorting**: Images are sorted by width-to-height ratio to minimize padding overhead
2. **Batch Slicing**: The sorted list is divided into chunks of size `rec_batch_num`
3. **Normalization**: Each image passes through `TextRecognizer.resize_norm_img`, which resizes, normalizes pixel values, and pads images to a uniform tensor shape within the batch

### Step 3: Batch Inference Execution

For standard Paddle Inference (`use_onnx=False`), the entire batch tensor copies directly into the GPU memory:

```python
self.input_tensor.copy_from_cpu(norm_img_batch)
self.predictor.run()

```

When using ONNX runtime, the batch feeds through a dictionary mapping:

```python
input_dict[self.input_tensor.name] = norm_img_batch

```

### Step 4: Aligned Post-Processing

Raw network outputs feed into the post-processor (CTC decoder, SRN head, etc.) which returns `[text, score]` pairs. Crucially, results align with the **original image order**, not the sorted batch order, ensuring proper correspondence between input images and recognized text.

## Command-Line Batch Recognition

Execute batch processing directly from the terminal using [`predict_rec.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/predict_rec.py). This approach handles directory traversal and internal batching automatically:

```bash
python tools/infer/predict_rec.py \
    --image_dir ./doc/imgs \
    --rec_model_dir ./inference/ch_ppocr_mobile_v2.0_rec_infer \
    --rec_batch_num 8 \
    --use_gpu True

```

The script loads all images from `./doc/imgs`, groups them into batches of 8, and processes them through the recognition pipeline.

## Python API Implementation Methods

### High-Level TextRecognizer API

For programmatic control, instantiate `TextRecognizer` directly and pass image lists:

```python
import cv2
from tools.infer.predict_rec import TextRecognizer
from tools.infer import utility

# Initialize configuration

args = utility.init_args().parse_args([
    "--rec_model_dir", "./inference/ch_ppocr_mobile_v2.0_rec_infer",
    "--rec_batch_num", "8",
    "--use_gpu", "True"
])

# Build recognizer

recognizer = TextRecognizer(args)

# Load heterogeneous images

imgs = [cv2.imread(p) for p in ["img1.jpg", "img2.jpg", "img3.jpg"]]

# Execute batch inference with automatic internal slicing

results, elapsed = recognizer(imgs)

# Results maintain input order

for path, (txt, score) in zip(["img1.jpg", "img2.jpg", "img3.jpg"], results):
    print(f"{path}: {txt} (score={score:.3f})")

```

### Low-Level Predictor API

For advanced scenarios requiring manual tensor manipulation, interact with the predictor directly:

```python
from tools.infer.utility import create_predictor
import numpy as np
import cv2

# Create predictor once

predictor, input_tensor, output_tensors, cfg = create_predictor(
    args, "rec", utility.get_logger()
)

# Manually preprocess and stack batch

batch_imgs = np.stack([
    recognizer.resize_norm_img(cv2.imread(p), 1.0) 
    for p in ["img1.jpg", "img2.jpg"]
])

# Feed batch to GPU

input_tensor.copy_from_cpu(batch_imgs)
predictor.run()

# Extract raw predictions

outputs = [t.copy_to_cpu() for t in output_tensors]

```

## Key Configuration and Performance Considerations

- **`rec_batch_num`**: Default values vary by model (typically 6 for mobile, higher for server), but can increase to 32 or 64 depending on GPU memory capacity
- **Memory Scaling**: Each batch element consumes GPU memory proportional to image dimensions; larger batches require proportional VRAM increases
- **Aspect Ratio Efficiency**: The internal sorting by aspect ratio minimizes padding waste, but extreme size variations within a single batch may reduce computational efficiency
- **Model Compatibility**: All recognition heads in `ppocr/modeling` (CTC, SRN, NRTR) support batch dimensions of `[B, C, H, W]`

## Summary

- **Batch processing in PaddleOCR applies exclusively to the recognition stage** (`TextRecognizer`); detection remains single-image only
- Control batch size through the `--rec_batch_num` argument defined in [`tools/infer/utility.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/utility.py)
- The `TextRecognizer.__call__` method handles automatic sorting, padding, and batch slicing internally
- Inference executes via `input_tensor.copy_from_cpu()` followed by `predictor.run()` for Paddle Inference, or dictionary input for ONNX
- Results maintain original image order despite internal aspect-ratio sorting for padding efficiency

## Frequently Asked Questions

### Can I batch process images through the detection stage in PaddleOCR?

No. According to the source code in [`tools/infer/predict_det.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_det.py), the `TextDetector` class and its underlying predictor are architected strictly for single-image processing with a fixed batch size of 1. You must run detection sequentially on each image, then feed the resulting cropped text regions into the batched recognition pipeline.

### What is the optimal batch size for PaddleOCR recognition?

The optimal value depends on your GPU memory and image dimensions. The default `rec_batch_num` of 6 works for most consumer GPUs, but server-grade hardware with 16GB+ VRAM can handle 32 or 64 images simultaneously. Monitor GPU memory usage and increase the value until you approach memory limits, as larger batches generally improve throughput until the GPU becomes saturated.

### Does batch processing work with the ONNX runtime in PaddleOCR?

Yes, but the input mechanism differs slightly. When `use_onnx=True`, the batch tensor passes through a dictionary mapping (`input_dict[self.input_tensor.name] = norm_img_batch`) rather than the `copy_from_cpu` method used by the standard Paddle Inference engine. Both paths execute batch inference in a single forward pass.

### How does PaddleOCR handle images of different sizes in a single batch?

The `TextRecognizer.__call__` method first sorts images by aspect ratio to group similarly-shaped images together, then resizes and pads each image within the batch to a common tensor shape using `resize_norm_img`. This minimizes padding waste while ensuring the neural network receives a uniform `[B, C, H, W]` input tensor. The post-processing stage then re-aligns results with the original input order.