# How to Optimize YOLOv5 Inference Speed on Edge Devices: 8 Proven Techniques

> Boost YOLOv5 inference speed on edge devices using half-precision, TensorRT, OpenVINO, and reduced resolution. Achieve real-time performance on resource-constrained hardware.

- Repository: [Ultralytics/yolov5](https://github.com/ultralytics/yolov5)
- Tags: performance
- Published: 2026-03-06

---

**Enable half-precision inference with `--half`, export to TensorRT or OpenVINO using [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py), and reduce input resolution via `--imgsz` to achieve real-time YOLOv5 performance on resource-constrained hardware.**

The `ultralytics/yolov5` repository provides a modular inference pipeline specifically designed for deployment flexibility on edge hardware. By targeting the appropriate backend runtime and precision format, you can significantly reduce latency and memory footprint without retraining the model. This guide maps each optimization to its exact implementation in the source code, providing actionable commands and API examples for NVIDIA Jetson, Intel Movidius, Raspberry Pi, and Apple Neural Engine platforms.

## Enable FP16 Half-Precision Inference

The fastest way to reduce memory bandwidth and computation time is to switch from 32-bit floating point (FP32) to 16-bit floating point (FP16). In [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py), the `--half` flag toggles half-precision mode by passing `fp16=True` to the `DetectMultiBackend` class【detect.py†L65-L66】.

When enabled, `DetectMultiBackend` (located in [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py)) automatically casts input tensors to `float16` if the device supports it, halving the data transfer requirements and enabling faster tensor operations on modern GPUs and NPUs.

```bash
python detect.py --weights yolov5s.pt --half --device 0

```

For edge devices with TensorRT support, FP16 inference is essential for maximizing throughput on NVIDIA Jetson platforms.

## Export to Hardware-Optimized Backends

The [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) script converts PyTorch models into hardware-specific formats that eliminate Python interpreter overhead and leverage platform-specific accelerators. The `DetectMultiBackend` abstraction in [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py) automatically selects the optimal runtime based on file extensions present in the working directory.

### ONNX and OpenCV DNN for CPU Devices

For ARM-based CPUs like the Raspberry Pi or generic x86 edge devices, export to ONNX and enable the OpenCV DNN backend. In [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py), the `--dnn` flag passes `dnn=True` to `DetectMultiBackend`【detect.py†L65-L66】, which loads the model via OpenCV's DNN module instead of PyTorch.

```bash

# Export once

python export.py --weights yolov5s.pt --include onnx

# Run inference on edge CPU with OpenCV DNN acceleration

python detect.py --weights yolov5s.onnx --dnn --half --device cpu --source images/

```

This backend avoids the PyTorch dependency entirely, reducing binary size and memory footprint for embedded Linux systems.

### TensorRT Engine for NVIDIA Jetson

For NVIDIA Jetson Nano, TX2, or Xavier devices, generate a TensorRT engine file to access kernel-level optimizations and INT8/FP16 inference capabilities. When you export with `--include engine`, [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) creates a serialized `.engine` file【export.py†L30-L34】. `DetectMultiBackend` automatically loads this file preferentially over the PyTorch model when present.

```bash

# Export on workstation or Jetson with sufficient memory

python export.py --weights yolov5s.pt --include engine --half

# Inference using the TensorRT backend

python detect.py --weights yolov5s.engine --half --device 0

```

TensorRT performs layer fusion, kernel auto-tuning, and precision calibration specific to the target GPU architecture, delivering the lowest possible latency on NVIDIA edge hardware.

### OpenVINO for Intel VPUs

Deploy YOLOv5 on Intel Movidius VPU, Neural Compute Stick 2 (NCS2), or integrated Intel GPUs by exporting to OpenVINO Intermediate Representation (IR) format.

```bash
python export.py --weights yolov5s.pt --include openvino
python detect.py --weights yolov5s_openvino_model/ --source 0

```

`DetectMultiBackend` recognizes the `.xml`/`.bin` file pair and loads the OpenVINO runtime, enabling VPU-accelerated inference without CUDA dependencies.

### CoreML for Apple Neural Engine

For iOS devices and Apple Silicon Macs running as edge servers, export to CoreML format to leverage the Neural Engine.

```bash
python export.py --weights yolov5s.pt --include coreml

```

The resulting `.mlmodel` file runs through Apple's CoreML runtime, providing dedicated NPU acceleration while maintaining battery efficiency on mobile devices.

## Optimize Input Resolution and Batch Configuration

The computational cost of YOLOv5 scales quadratically with input resolution. In [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py), the `--imgsz` argument controls the inference dimensions【detect.py†L74-L75】. Reducing the default 640x640 to 320 or 416 decreases FLOPs proportionally, often enabling real-time inference on modest CPUs with minimal accuracy degradation.

```bash
python detect.py --weights yolov5s.pt --imgsz 416 --half

```

For edge deployment, maintain a batch size of 1 to avoid memory copies and buffering overhead. The [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py) script explicitly sets `bs = 1` for single-image inference pipelines【detect.py†L70-L71】, which is the optimal configuration for streaming video inputs on memory-constrained devices.

## Execute Warm-Up and Minimize Post-Processing

Before measuring performance or processing production data, trigger the warm-up routine to force kernel compilation and memory allocation. The `model.warmup()` method is called in [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py) before the inference loop begins【detect.py†L81-L83】, ensuring steady-state latency by eliminating first-run compilation overhead on GPUs and TensorRT engines.

Additionally, reduce CPU overhead by disabling unnecessary post-processing when only raw detections are required:

- Use `--nosave` to skip disk I/O for output images
- Use `--save-txt` only if bounding box coordinates must be persisted
- Consider modifying the code to skip `non_max_suppression` if your application can tolerate redundant detections or handles filtering downstream

## Implement INT8 Quantization for Extreme Compression

While not built into the main repository, the ONNX export path enables external quantization workflows. Convert your exported ONNX model to 8-bit integer format using `onnxruntime` quantization tools for deployment on CPU-only edge devices where FP16 support is unavailable.

```python
import onnxruntime as ort
from onnxruntime.quantization import quantize_dynamic, QuantType

# Convert FP32 ONNX to INT8

quantize_dynamic(
    model_input='yolov5s.onnx',
    model_output='yolov5s_int8.onnx',
    weight_type=QuantType.QInt8
)

# Inference session

sess = ort.InferenceSession('yolov5s_int8.onnx')
input_name = sess.get_inputs()[0].name
outputs = sess.run(None, {input_name: img_array})

```

INT8 quantization reduces model size by 75% and significantly increases inference speed on x86 and ARM CPUs at the cost of minor precision loss.

## Load Models via Python API for Embedded Integration

For custom edge applications requiring direct Python integration, use `DetectMultiBackend` with exported TorchScript or ONNX models to bypass the command-line interface overhead.

```python
from pathlib import Path
from models.common import DetectMultiBackend
from utils.torch_utils import select_device
import torch
import cv2

# Initialize backend with TorchScript for ultra-low latency

device = select_device('cpu')
model = DetectMultiBackend(
    weights=Path('yolov5s.torchscript'),
    device=device,
    dnn=False,
    data=None,
    fp16=True
)

# Preprocess image

img = cv2.imread('bus.jpg')
img = cv2.resize(img, (640, 640))
img_tensor = torch.from_numpy(img).permute(2, 0, 1).unsqueeze(0).float() / 255.0
if model.fp16:
    img_tensor = img_tensor.half()

# Inference

pred = model(img_tensor, augment=False, visualize=False)

```

This pattern is essential for robotics applications where the inference engine must run as a library within a larger control system.

## Summary

- **Enable `--half`** in [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py) to activate FP16 inference through `DetectMultiBackend`, reducing memory bandwidth by 50% on compatible hardware.
- **Export to specialized backends** using [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py)—choose TensorRT for NVIDIA Jetson, OpenVINO for Intel VPUs, CoreML for Apple devices, or ONNX+OpenCV DNN for generic ARM CPUs.
- **Reduce resolution** with `--imgsz` to 320-416 for edge devices, and maintain batch size of 1 to minimize memory copies.
- **Call `model.warmup()`** before production inference to ensure kernels are compiled and latency is consistent.
- **Quantize to INT8** via ONNX Runtime for maximum compression on CPU-only edge hardware where GPU acceleration is unavailable.

## Frequently Asked Questions

### What is the fastest backend for YOLOv5 on NVIDIA Jetson devices?

TensorRT provides the lowest latency and highest throughput on NVIDIA Jetson platforms. Export your model using `python export.py --weights yolov5s.pt --include engine --half`, then run inference with `python detect.py --weights yolov5s.engine --half --device 0`. The TensorRT engine performs layer fusion and kernel optimization specific to the Jetson GPU architecture.

### How do I run YOLOv5 on a Raspberry Pi without GPU acceleration?

Export the model to ONNX format using [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py), then use the OpenCV DNN backend. Run `python detect.py --weights yolov5s.onnx --dnn --half --device cpu --imgsz 416`. The OpenCV DNN backend eliminates the PyTorch dependency and runs efficiently on ARM Cortex CPUs, while reducing the input resolution to 416 or 320 maintains real-time performance.

### Does FP16 inference work on all edge devices?

No, FP16 requires hardware support. It works natively on NVIDIA GPUs (Jetson), Apple Neural Engine (via CoreML), and some modern ARM processors with FP16 SIMD instructions. For older CPUs without FP16 support, use INT8 quantization via ONNX Runtime or stick with FP32 to avoid emulation overhead.

### Where in the codebase does the warm-up functionality reside?

The warm-up routine is implemented in [`utils/general.py`](https://github.com/ultralytics/yolov5/blob/main/utils/general.py) and invoked in [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py) at lines 81-83【detect.py†L81-L83】. The code calls `model.warmup(imgsz=(1 if pt else bs, 3, *imgsz))` to execute a dummy forward pass, forcing CUDA kernel compilation and memory allocation before the actual inference loop begins.