How to Perform Batch Processing with PaddleOCR: A Complete Guide to Batch Recognition
PaddleOCR accelerates text recognition by processing multiple images simultaneously through the TextRecognizer class using the --rec_batch_num parameter, while the detection stage remains strictly single-image.
Batch processing dramatically reduces inference latency when processing large volumes of text images. In the PaddlePaddle/PaddleOCR repository, batch inference is implemented specifically within the recognition pipeline, allowing efficient parallel processing of cropped text line images while maintaining compatibility with both Paddle Inference and ONNX runtimes.
Understanding the Batch Processing Architecture
PaddleOCR separates its two-stage OCR pipeline into distinct batching behaviors. The detection stage (TextDetector in tools/infer/predict_det.py) processes images individually because its predictor architecture expects a batch size of 1. Conversely, the recognition stage (TextRecognizer in tools/infer/predict_rec.py) fully supports batched inference, making it the primary target for throughput optimization.
The entry point for batch configuration resides in tools/infer/utility.py, where the argument parser defines the --rec_batch_num flag. This parameter controls exactly how many text line images enter the recognition network simultaneously:
# tools/infer/utility.py
parser.add_argument("--rec_batch_num", type=int, default=6, help="...")
How Batch Processing Works in TextRecognizer
The TextRecognizer.__call__ method orchestrates the entire batch workflow, from raw images to final text predictions. Understanding this internal pipeline is essential for optimizing performance.
Step 1: Predictor Initialization
First, utility.create_predictor constructs the inference engine for mode="rec". This factory function returns a predictor handle, an input tensor reference, and a list of output tensors configured for the specific recognition model (CTC-based, SRN, or others).
Step 2: Dynamic Batch Preprocessing
Before inference, the input list undergoes intelligent batching:
- Aspect Ratio Sorting: Images are sorted by width-to-height ratio to minimize padding overhead
- Batch Slicing: The sorted list is divided into chunks of size
rec_batch_num - Normalization: Each image passes through
TextRecognizer.resize_norm_img, which resizes, normalizes pixel values, and pads images to a uniform tensor shape within the batch
Step 3: Batch Inference Execution
For standard Paddle Inference (use_onnx=False), the entire batch tensor copies directly into the GPU memory:
self.input_tensor.copy_from_cpu(norm_img_batch)
self.predictor.run()
When using ONNX runtime, the batch feeds through a dictionary mapping:
input_dict[self.input_tensor.name] = norm_img_batch
Step 4: Aligned Post-Processing
Raw network outputs feed into the post-processor (CTC decoder, SRN head, etc.) which returns [text, score] pairs. Crucially, results align with the original image order, not the sorted batch order, ensuring proper correspondence between input images and recognized text.
Command-Line Batch Recognition
Execute batch processing directly from the terminal using predict_rec.py. This approach handles directory traversal and internal batching automatically:
python tools/infer/predict_rec.py \
--image_dir ./doc/imgs \
--rec_model_dir ./inference/ch_ppocr_mobile_v2.0_rec_infer \
--rec_batch_num 8 \
--use_gpu True
The script loads all images from ./doc/imgs, groups them into batches of 8, and processes them through the recognition pipeline.
Python API Implementation Methods
High-Level TextRecognizer API
For programmatic control, instantiate TextRecognizer directly and pass image lists:
import cv2
from tools.infer.predict_rec import TextRecognizer
from tools.infer import utility
# Initialize configuration
args = utility.init_args().parse_args([
"--rec_model_dir", "./inference/ch_ppocr_mobile_v2.0_rec_infer",
"--rec_batch_num", "8",
"--use_gpu", "True"
])
# Build recognizer
recognizer = TextRecognizer(args)
# Load heterogeneous images
imgs = [cv2.imread(p) for p in ["img1.jpg", "img2.jpg", "img3.jpg"]]
# Execute batch inference with automatic internal slicing
results, elapsed = recognizer(imgs)
# Results maintain input order
for path, (txt, score) in zip(["img1.jpg", "img2.jpg", "img3.jpg"], results):
print(f"{path}: {txt} (score={score:.3f})")
Low-Level Predictor API
For advanced scenarios requiring manual tensor manipulation, interact with the predictor directly:
from tools.infer.utility import create_predictor
import numpy as np
import cv2
# Create predictor once
predictor, input_tensor, output_tensors, cfg = create_predictor(
args, "rec", utility.get_logger()
)
# Manually preprocess and stack batch
batch_imgs = np.stack([
recognizer.resize_norm_img(cv2.imread(p), 1.0)
for p in ["img1.jpg", "img2.jpg"]
])
# Feed batch to GPU
input_tensor.copy_from_cpu(batch_imgs)
predictor.run()
# Extract raw predictions
outputs = [t.copy_to_cpu() for t in output_tensors]
Key Configuration and Performance Considerations
rec_batch_num: Default values vary by model (typically 6 for mobile, higher for server), but can increase to 32 or 64 depending on GPU memory capacity- Memory Scaling: Each batch element consumes GPU memory proportional to image dimensions; larger batches require proportional VRAM increases
- Aspect Ratio Efficiency: The internal sorting by aspect ratio minimizes padding waste, but extreme size variations within a single batch may reduce computational efficiency
- Model Compatibility: All recognition heads in
ppocr/modeling(CTC, SRN, NRTR) support batch dimensions of[B, C, H, W]
Summary
- Batch processing in PaddleOCR applies exclusively to the recognition stage (
TextRecognizer); detection remains single-image only - Control batch size through the
--rec_batch_numargument defined intools/infer/utility.py - The
TextRecognizer.__call__method handles automatic sorting, padding, and batch slicing internally - Inference executes via
input_tensor.copy_from_cpu()followed bypredictor.run()for Paddle Inference, or dictionary input for ONNX - Results maintain original image order despite internal aspect-ratio sorting for padding efficiency
Frequently Asked Questions
Can I batch process images through the detection stage in PaddleOCR?
No. According to the source code in tools/infer/predict_det.py, the TextDetector class and its underlying predictor are architected strictly for single-image processing with a fixed batch size of 1. You must run detection sequentially on each image, then feed the resulting cropped text regions into the batched recognition pipeline.
What is the optimal batch size for PaddleOCR recognition?
The optimal value depends on your GPU memory and image dimensions. The default rec_batch_num of 6 works for most consumer GPUs, but server-grade hardware with 16GB+ VRAM can handle 32 or 64 images simultaneously. Monitor GPU memory usage and increase the value until you approach memory limits, as larger batches generally improve throughput until the GPU becomes saturated.
Does batch processing work with the ONNX runtime in PaddleOCR?
Yes, but the input mechanism differs slightly. When use_onnx=True, the batch tensor passes through a dictionary mapping (input_dict[self.input_tensor.name] = norm_img_batch) rather than the copy_from_cpu method used by the standard Paddle Inference engine. Both paths execute batch inference in a single forward pass.
How does PaddleOCR handle images of different sizes in a single batch?
The TextRecognizer.__call__ method first sorts images by aspect ratio to group similarly-shaped images together, then resizes and pads each image within the batch to a common tensor shape using resize_norm_img. This minimizes padding waste while ensuring the neural network receives a uniform [B, C, H, W] input tensor. The post-processing stage then re-aligns results with the original input order.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →