# How to Optimize PaddleOCR for Inference Speed: 7 Proven Acceleration Techniques

> Boost PaddleOCR inference speed with 7 proven techniques. Learn to optimize models, use hardware backends, INT8 quantization, and batch processing for faster results.

- Repository: [PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)
- Tags: performance
- Published: 2026-03-03

---

**You can optimize PaddleOCR inference speed by combining lightweight model variants, hardware-optimized backends (TensorRT/MKL-DNN), the High-Performance Inference (HPI) plugin, ONNX Runtime conversion, INT8 quantization, and batch processing to achieve maximum throughput with minimal accuracy loss.**

PaddleOCR from the PaddlePaddle ecosystem offers multiple orthogonal optimization pathways for production deployment. To optimize PaddleOCR for inference speed effectively, you must select the right combination of model architecture, inference backend, and runtime configuration based on your target hardware. The following techniques can be stacked cumulatively to minimize latency per image.

## Select Lightweight Model Architectures

The foundation of fast inference starts with selecting models specifically designed for speed rather than accuracy. PaddleOCR provides **PP-OCRv5**, **PP-LCNet**, and **SVTR-v2** variants that trade marginal accuracy for significant computational efficiency.

Choose configurations labeled "slim" or "mobile" in the repository's config directory. For example, [`configs/rec/PP-OCRv5/PP-OCRv5_rec.yml`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/configs/rec/PP-OCRv5/PP-OCRv5_rec.yml) specifies a lightweight backbone optimized for edge deployment. These models reduce floating-point operations (FLOPs) by 40-60% compared to standard ResNet-based backbones while maintaining competitive OCR accuracy.

## Enable Hardware-Optimized Inference Backends

PaddleOCR supports highly optimized kernels through **MKL-DNN** for Intel CPUs and **TensorRT** for NVIDIA GPUs. These backends replace generic implementations with vendor-optimized primitives for convolution and matrix multiplication operations.

In [`tools/infer/utility.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/utility.py) (lines 133-158), the argument parser defines the critical flags:
- `--enable_mkldnn True` activates Intel's Math Kernel Library for Deep Neural Networks on compatible CPUs
- `--use_tensorrt True` enables NVIDIA's TensorRT engine for GPU inference
- `--precision fp16` or `--precision int8` triggers mixed-precision or quantized execution when combined with TensorRT

The `create_predictor` function in [`tools/infer/predict_rec.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_rec.py) (lines 183-196) processes these arguments, calling `config.enable_tensorrt_engine()` when GPU acceleration is requested or `config.enable_mkldnn()` for CPU optimization.

## Deploy the High-Performance Inference (HPI) Plugin

The **High-Performance Inference (HPI)** plugin automatically selects the optimal backend from Paddle Inference, ONNX Runtime, OpenVINO, or TensorRT based on your hardware and caches the compiled engine for subsequent runs. This eliminates manual backend tuning.

As documented in [`docs/version3.x/deployment/high_performance_inference.en.md`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/docs/version3.x/deployment/high_performance_inference.en.md), enable HPI via the CLI flag `--enable_hpi True` or the Python API parameter `enable_hpi=True`. When activated, the plugin inspects the model and hardware capabilities, then routes inference through the fastest available execution path. The first run compiles and caches the engine; subsequent runs achieve maximum throughput without warmup penalties.

## Convert Models to ONNX Runtime

Converting PaddleOCR models to **ONNX format** often yields faster inference on diverse hardware, particularly when combined with TensorRT or OpenVINO execution providers. The conversion logic resides in [`tools/infer/predict_rec.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_rec.py) (lines 183-188), which loads an ONNX Runtime session when `args.use_onnx` evaluates to true.

Enable this optimization by passing `--use_onnx True` in the CLI or setting `use_onnx=True` in the `PaddleOCR` class constructor. ONNX Runtime abstracts hardware-specific optimizations and frequently outperforms native Paddle Inference on x86 CPUs and edge devices due to superior kernel implementations.

## Apply Quantization and Model Pruning

**Quantization-aware training (QAT)** reduces weight and activation precision to INT8, shrinking model size by 75% and enabling integer-only arithmetic on supported hardware. The QAT implementation lives in [`deploy/slim/quantization/quant.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/deploy/slim/quantization/quant.py) (line 65), which fine-tunes models to maintain accuracy under reduced precision.

After training, export the quantized model using [`deploy/slim/quantization/export_model.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/deploy/slim/quantization/export_model.py) (line 35). During inference, combine the exported model with `--precision int8` to force the backend (TensorRT or MKL-DNN) to use INT8 kernels.

For additional speedup, apply **model pruning** via [`deploy/slim/prune/export_prune_model.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/deploy/slim/prune/export_prune_model.py) to remove redundant filters and channels, directly decreasing FLOPs and memory bandwidth requirements.

## Maximize Throughput with Batching and Parallelism

Reduce per-image overhead by processing multiple images simultaneously. Increase the `rec_batch_num` parameter in your configuration file or pass `--batch_num N` via CLI to enable batch inference. The predictor in [`tools/infer/predict_rec.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_rec.py) processes these batches in a single forward pass, amortizing the fixed cost of kernel launches across multiple images.

For multi-GPU or multi-core CPU systems, implement parallel inference as shown in [`docs/version3.x/pipeline_usage/instructions/parallel_inference.en.md`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/docs/version3.x/pipeline_usage/instructions/parallel_inference.en.md) (line 35). Spawn separate processes targeting specific devices (e.g., `gpu:0`, `gpu:1`) to saturate all available compute units. Use `--benchmark True` (defined in [`tools/infer/utility.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/utility.py) line 153) to measure per-step latencies and verify throughput gains.

## Code Examples

### CLI: Maximum Speed on GPU with TensorRT and HPI

```bash
paddleocr ocr \
  --rec_model_dir ./inference/ch_PP-OCRv5_rec_infer \
  --use_tensorrt True \
  --precision int8 \
  --enable_hpi True \
  --rec_batch_num 8 \
  --benchmark True \
  --image_dir ./test_images

```

This configuration enables TensorRT with INT8 precision processing, activates the HPI plugin for automatic backend selection, and processes eight images per batch to minimize overhead.

### Python API: Programmatic Optimization

```python
from paddleocr import PaddleOCR

ocr = PaddleOCR(
    rec_model_dir="./inference/ch_PP-OCRv5_rec_infer",
    use_tensorrt=True,
    precision="int8",
    enable_hpi=True,
    rec_batch_num=8,
)
result = ocr.ocr("./test_images")

```

The `PaddleOCR` constructor forwards these parameters to the internal predictor creation logic, triggering the same optimizations as the CLI.

### Quantization-Aware Training Pipeline

```bash

# Step 1: Train with quantization

python deploy/slim/quantization/quant.py \
  -c configs/det/ch_PP-OCRv3/ch_PP-OCRv3_det_cml.yml \
  -o Global.pretrained_model=./ch_PP-OCRv3_det_distill_train/best_accuracy \
     Global.save_model_dir=./output/quant_model_distill/

# Step 2: Export quantized inference model

python deploy/slim/quantization/export_model.py \
  -c configs/det/ch_PP-OCRv3/ch_PP-OCRv3_det_cml.yml \
  -o Global.checkpoints=output/quant_model_distill/best_accuracy \
     Global.save_inference_dir=./output/quant_inference_model

```

### Multi-Process Parallel Inference

```python
import multiprocessing as mp
from paddleocr import PaddleOCR

def worker(rank, device):
    ocr = PaddleOCR(device=device, rec_batch_num=8)
    ocr.ocr(f"./data/part_{rank}")

if __name__ == "__main__":
    devices = ["gpu:0", "gpu:1"]
    processes = []
    for i, dev in enumerate(devices):
        p = mp.Process(target=worker, args=(i, dev))
        p.start()
        processes.append(p)
    for p in processes:
        p.join()

```

## Summary

- **Choose lightweight architectures** like PP-OCRv5 via configs such as [`configs/rec/PP-OCRv5/PP-OCRv5_rec.yml`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/configs/rec/PP-OCRv5/PP-OCRv5_rec.yml) to reduce baseline computational requirements.
- **Enable hardware backends** by setting `--enable_mkldnn True` for Intel CPUs or `--use_tensorrt True` for NVIDIA GPUs in [`tools/infer/utility.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/utility.py) arguments.
- **Activate HPI** with `--enable_hpi True` to automatically select between Paddle Inference, ONNX Runtime, OpenVINO, and TensorRT without manual tuning.
- **Convert to ONNX** using `--use_onnx True` when targeting heterogeneous hardware or when ONNX Runtime provides superior kernel implementations.
- **Quantize to INT8** using [`deploy/slim/quantization/quant.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/deploy/slim/quantization/quant.py) and [`export_model.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/export_model.py), then inference with `--precision int8` for 2-4x throughput gains on compatible hardware.
- **Prune models** via [`deploy/slim/prune/export_prune_model.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/deploy/slim/prune/export_prune_model.py) to eliminate redundant parameters and reduce FLOPs.
- **Batch and parallelize** inference by adjusting `rec_batch_num` and spawning multiple processes per GPU to saturate compute resources.

## Frequently Asked Questions

### How do I know if TensorRT is actually running during inference?

Enable the benchmark flag by adding `--benchmark True` to your command. According to the implementation in [`tools/infer/utility.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/utility.py) (line 153), this prints detailed per-step timings including preprocessing, inference, and postprocessing durations. If TensorRT is active, you will see engine building messages on the first run and significantly faster inference times (typically 3-5x speedup over native GPU inference) on subsequent runs.

### Can I combine quantization with the HPI plugin?

Yes. The HPI plugin in PaddleOCR automatically detects quantized models and routes them to the appropriate INT8-capable backend (TensorRT or OpenVINO) when available. Pass your quantized model directory to `--rec_model_dir` and ensure `--precision int8` is set. As documented in [`docs/version3.x/deployment/high_performance_inference.en.md`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/docs/version3.x/deployment/high_performance_inference.en.md), the plugin caches the optimized engine after the first warm-up run, combining the benefits of quantization with automatic backend selection.

### What is the difference between MKL-DNN and TensorRT optimization?

**MKL-DNN** (now oneDNN) optimizes CPU inference by utilizing Intel-specific AVX-512 and AMX instructions for deep learning primitives, accessible via `--enable_mkldnn True` in [`tools/infer/utility.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/utility.py). **TensorRT** optimizes GPU inference through layer fusion, kernel auto-tuning, and mixed-precision execution on NVIDIA hardware, enabled via `--use_tensorrt True`. MKL-DNN targets x86_64 CPUs exclusively, while TensorRT requires NVIDIA GPUs. Both can be combined with INT8 quantization for additional speedup on their respective hardware platforms.

### How does batch size affect inference speed in PaddleOCR?

Increasing `rec_batch_num` (or `--batch_num` in CLI) reduces per-image overhead by processing multiple images in a single forward pass through the neural network. In [`tools/infer/predict_rec.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/tools/infer/predict_rec.py), the predictor processes these batches as a single tensor operation, amortizing the fixed cost of kernel launches and memory transfers. However, diminishing returns occur when batch sizes exceed GPU memory capacity or when the bottleneck shifts to I/O rather than computation. Benchmark with `--benchmark True` to identify the optimal batch size for your hardware (typically 4-8 for consumer GPUs, 16-32 for datacenter cards).