How to Optimize PaddleOCR for Inference Speed: 7 Proven Acceleration Techniques
You can optimize PaddleOCR inference speed by combining lightweight model variants, hardware-optimized backends (TensorRT/MKL-DNN), the High-Performance Inference (HPI) plugin, ONNX Runtime conversion, INT8 quantization, and batch processing to achieve maximum throughput with minimal accuracy loss.
PaddleOCR from the PaddlePaddle ecosystem offers multiple orthogonal optimization pathways for production deployment. To optimize PaddleOCR for inference speed effectively, you must select the right combination of model architecture, inference backend, and runtime configuration based on your target hardware. The following techniques can be stacked cumulatively to minimize latency per image.
Select Lightweight Model Architectures
The foundation of fast inference starts with selecting models specifically designed for speed rather than accuracy. PaddleOCR provides PP-OCRv5, PP-LCNet, and SVTR-v2 variants that trade marginal accuracy for significant computational efficiency.
Choose configurations labeled "slim" or "mobile" in the repository's config directory. For example, configs/rec/PP-OCRv5/PP-OCRv5_rec.yml specifies a lightweight backbone optimized for edge deployment. These models reduce floating-point operations (FLOPs) by 40-60% compared to standard ResNet-based backbones while maintaining competitive OCR accuracy.
Enable Hardware-Optimized Inference Backends
PaddleOCR supports highly optimized kernels through MKL-DNN for Intel CPUs and TensorRT for NVIDIA GPUs. These backends replace generic implementations with vendor-optimized primitives for convolution and matrix multiplication operations.
In tools/infer/utility.py (lines 133-158), the argument parser defines the critical flags:
--enable_mkldnn Trueactivates Intel's Math Kernel Library for Deep Neural Networks on compatible CPUs--use_tensorrt Trueenables NVIDIA's TensorRT engine for GPU inference--precision fp16or--precision int8triggers mixed-precision or quantized execution when combined with TensorRT
The create_predictor function in tools/infer/predict_rec.py (lines 183-196) processes these arguments, calling config.enable_tensorrt_engine() when GPU acceleration is requested or config.enable_mkldnn() for CPU optimization.
Deploy the High-Performance Inference (HPI) Plugin
The High-Performance Inference (HPI) plugin automatically selects the optimal backend from Paddle Inference, ONNX Runtime, OpenVINO, or TensorRT based on your hardware and caches the compiled engine for subsequent runs. This eliminates manual backend tuning.
As documented in docs/version3.x/deployment/high_performance_inference.en.md, enable HPI via the CLI flag --enable_hpi True or the Python API parameter enable_hpi=True. When activated, the plugin inspects the model and hardware capabilities, then routes inference through the fastest available execution path. The first run compiles and caches the engine; subsequent runs achieve maximum throughput without warmup penalties.
Convert Models to ONNX Runtime
Converting PaddleOCR models to ONNX format often yields faster inference on diverse hardware, particularly when combined with TensorRT or OpenVINO execution providers. The conversion logic resides in tools/infer/predict_rec.py (lines 183-188), which loads an ONNX Runtime session when args.use_onnx evaluates to true.
Enable this optimization by passing --use_onnx True in the CLI or setting use_onnx=True in the PaddleOCR class constructor. ONNX Runtime abstracts hardware-specific optimizations and frequently outperforms native Paddle Inference on x86 CPUs and edge devices due to superior kernel implementations.
Apply Quantization and Model Pruning
Quantization-aware training (QAT) reduces weight and activation precision to INT8, shrinking model size by 75% and enabling integer-only arithmetic on supported hardware. The QAT implementation lives in deploy/slim/quantization/quant.py (line 65), which fine-tunes models to maintain accuracy under reduced precision.
After training, export the quantized model using deploy/slim/quantization/export_model.py (line 35). During inference, combine the exported model with --precision int8 to force the backend (TensorRT or MKL-DNN) to use INT8 kernels.
For additional speedup, apply model pruning via deploy/slim/prune/export_prune_model.py to remove redundant filters and channels, directly decreasing FLOPs and memory bandwidth requirements.
Maximize Throughput with Batching and Parallelism
Reduce per-image overhead by processing multiple images simultaneously. Increase the rec_batch_num parameter in your configuration file or pass --batch_num N via CLI to enable batch inference. The predictor in tools/infer/predict_rec.py processes these batches in a single forward pass, amortizing the fixed cost of kernel launches across multiple images.
For multi-GPU or multi-core CPU systems, implement parallel inference as shown in docs/version3.x/pipeline_usage/instructions/parallel_inference.en.md (line 35). Spawn separate processes targeting specific devices (e.g., gpu:0, gpu:1) to saturate all available compute units. Use --benchmark True (defined in tools/infer/utility.py line 153) to measure per-step latencies and verify throughput gains.
Code Examples
CLI: Maximum Speed on GPU with TensorRT and HPI
paddleocr ocr \
--rec_model_dir ./inference/ch_PP-OCRv5_rec_infer \
--use_tensorrt True \
--precision int8 \
--enable_hpi True \
--rec_batch_num 8 \
--benchmark True \
--image_dir ./test_images
This configuration enables TensorRT with INT8 precision processing, activates the HPI plugin for automatic backend selection, and processes eight images per batch to minimize overhead.
Python API: Programmatic Optimization
from paddleocr import PaddleOCR
ocr = PaddleOCR(
rec_model_dir="./inference/ch_PP-OCRv5_rec_infer",
use_tensorrt=True,
precision="int8",
enable_hpi=True,
rec_batch_num=8,
)
result = ocr.ocr("./test_images")
The PaddleOCR constructor forwards these parameters to the internal predictor creation logic, triggering the same optimizations as the CLI.
Quantization-Aware Training Pipeline
# Step 1: Train with quantization
python deploy/slim/quantization/quant.py \
-c configs/det/ch_PP-OCRv3/ch_PP-OCRv3_det_cml.yml \
-o Global.pretrained_model=./ch_PP-OCRv3_det_distill_train/best_accuracy \
Global.save_model_dir=./output/quant_model_distill/
# Step 2: Export quantized inference model
python deploy/slim/quantization/export_model.py \
-c configs/det/ch_PP-OCRv3/ch_PP-OCRv3_det_cml.yml \
-o Global.checkpoints=output/quant_model_distill/best_accuracy \
Global.save_inference_dir=./output/quant_inference_model
Multi-Process Parallel Inference
import multiprocessing as mp
from paddleocr import PaddleOCR
def worker(rank, device):
ocr = PaddleOCR(device=device, rec_batch_num=8)
ocr.ocr(f"./data/part_{rank}")
if __name__ == "__main__":
devices = ["gpu:0", "gpu:1"]
processes = []
for i, dev in enumerate(devices):
p = mp.Process(target=worker, args=(i, dev))
p.start()
processes.append(p)
for p in processes:
p.join()
Summary
- Choose lightweight architectures like PP-OCRv5 via configs such as
configs/rec/PP-OCRv5/PP-OCRv5_rec.ymlto reduce baseline computational requirements. - Enable hardware backends by setting
--enable_mkldnn Truefor Intel CPUs or--use_tensorrt Truefor NVIDIA GPUs intools/infer/utility.pyarguments. - Activate HPI with
--enable_hpi Trueto automatically select between Paddle Inference, ONNX Runtime, OpenVINO, and TensorRT without manual tuning. - Convert to ONNX using
--use_onnx Truewhen targeting heterogeneous hardware or when ONNX Runtime provides superior kernel implementations. - Quantize to INT8 using
deploy/slim/quantization/quant.pyandexport_model.py, then inference with--precision int8for 2-4x throughput gains on compatible hardware. - Prune models via
deploy/slim/prune/export_prune_model.pyto eliminate redundant parameters and reduce FLOPs. - Batch and parallelize inference by adjusting
rec_batch_numand spawning multiple processes per GPU to saturate compute resources.
Frequently Asked Questions
How do I know if TensorRT is actually running during inference?
Enable the benchmark flag by adding --benchmark True to your command. According to the implementation in tools/infer/utility.py (line 153), this prints detailed per-step timings including preprocessing, inference, and postprocessing durations. If TensorRT is active, you will see engine building messages on the first run and significantly faster inference times (typically 3-5x speedup over native GPU inference) on subsequent runs.
Can I combine quantization with the HPI plugin?
Yes. The HPI plugin in PaddleOCR automatically detects quantized models and routes them to the appropriate INT8-capable backend (TensorRT or OpenVINO) when available. Pass your quantized model directory to --rec_model_dir and ensure --precision int8 is set. As documented in docs/version3.x/deployment/high_performance_inference.en.md, the plugin caches the optimized engine after the first warm-up run, combining the benefits of quantization with automatic backend selection.
What is the difference between MKL-DNN and TensorRT optimization?
MKL-DNN (now oneDNN) optimizes CPU inference by utilizing Intel-specific AVX-512 and AMX instructions for deep learning primitives, accessible via --enable_mkldnn True in tools/infer/utility.py. TensorRT optimizes GPU inference through layer fusion, kernel auto-tuning, and mixed-precision execution on NVIDIA hardware, enabled via --use_tensorrt True. MKL-DNN targets x86_64 CPUs exclusively, while TensorRT requires NVIDIA GPUs. Both can be combined with INT8 quantization for additional speedup on their respective hardware platforms.
How does batch size affect inference speed in PaddleOCR?
Increasing rec_batch_num (or --batch_num in CLI) reduces per-image overhead by processing multiple images in a single forward pass through the neural network. In tools/infer/predict_rec.py, the predictor processes these batches as a single tensor operation, amortizing the fixed cost of kernel launches and memory transfers. However, diminishing returns occur when batch sizes exceed GPU memory capacity or when the bottleneck shifts to I/O rather than computation. Benchmark with --benchmark True to identify the optimal batch size for your hardware (typically 4-8 for consumer GPUs, 16-32 for datacenter cards).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →