How to Optimize YOLOv5 Inference Speed on Edge Devices: 8 Proven Techniques
Enable half-precision inference with --half, export to TensorRT or OpenVINO using export.py, and reduce input resolution via --imgsz to achieve real-time YOLOv5 performance on resource-constrained hardware.
The ultralytics/yolov5 repository provides a modular inference pipeline specifically designed for deployment flexibility on edge hardware. By targeting the appropriate backend runtime and precision format, you can significantly reduce latency and memory footprint without retraining the model. This guide maps each optimization to its exact implementation in the source code, providing actionable commands and API examples for NVIDIA Jetson, Intel Movidius, Raspberry Pi, and Apple Neural Engine platforms.
Enable FP16 Half-Precision Inference
The fastest way to reduce memory bandwidth and computation time is to switch from 32-bit floating point (FP32) to 16-bit floating point (FP16). In detect.py, the --half flag toggles half-precision mode by passing fp16=True to the DetectMultiBackend class【detect.py†L65-L66】.
When enabled, DetectMultiBackend (located in models/yolo.py) automatically casts input tensors to float16 if the device supports it, halving the data transfer requirements and enabling faster tensor operations on modern GPUs and NPUs.
python detect.py --weights yolov5s.pt --half --device 0
For edge devices with TensorRT support, FP16 inference is essential for maximizing throughput on NVIDIA Jetson platforms.
Export to Hardware-Optimized Backends
The export.py script converts PyTorch models into hardware-specific formats that eliminate Python interpreter overhead and leverage platform-specific accelerators. The DetectMultiBackend abstraction in models/yolo.py automatically selects the optimal runtime based on file extensions present in the working directory.
ONNX and OpenCV DNN for CPU Devices
For ARM-based CPUs like the Raspberry Pi or generic x86 edge devices, export to ONNX and enable the OpenCV DNN backend. In detect.py, the --dnn flag passes dnn=True to DetectMultiBackend【detect.py†L65-L66】, which loads the model via OpenCV's DNN module instead of PyTorch.
# Export once
python export.py --weights yolov5s.pt --include onnx
# Run inference on edge CPU with OpenCV DNN acceleration
python detect.py --weights yolov5s.onnx --dnn --half --device cpu --source images/
This backend avoids the PyTorch dependency entirely, reducing binary size and memory footprint for embedded Linux systems.
TensorRT Engine for NVIDIA Jetson
For NVIDIA Jetson Nano, TX2, or Xavier devices, generate a TensorRT engine file to access kernel-level optimizations and INT8/FP16 inference capabilities. When you export with --include engine, export.py creates a serialized .engine file【export.py†L30-L34】. DetectMultiBackend automatically loads this file preferentially over the PyTorch model when present.
# Export on workstation or Jetson with sufficient memory
python export.py --weights yolov5s.pt --include engine --half
# Inference using the TensorRT backend
python detect.py --weights yolov5s.engine --half --device 0
TensorRT performs layer fusion, kernel auto-tuning, and precision calibration specific to the target GPU architecture, delivering the lowest possible latency on NVIDIA edge hardware.
OpenVINO for Intel VPUs
Deploy YOLOv5 on Intel Movidius VPU, Neural Compute Stick 2 (NCS2), or integrated Intel GPUs by exporting to OpenVINO Intermediate Representation (IR) format.
python export.py --weights yolov5s.pt --include openvino
python detect.py --weights yolov5s_openvino_model/ --source 0
DetectMultiBackend recognizes the .xml/.bin file pair and loads the OpenVINO runtime, enabling VPU-accelerated inference without CUDA dependencies.
CoreML for Apple Neural Engine
For iOS devices and Apple Silicon Macs running as edge servers, export to CoreML format to leverage the Neural Engine.
python export.py --weights yolov5s.pt --include coreml
The resulting .mlmodel file runs through Apple's CoreML runtime, providing dedicated NPU acceleration while maintaining battery efficiency on mobile devices.
Optimize Input Resolution and Batch Configuration
The computational cost of YOLOv5 scales quadratically with input resolution. In detect.py, the --imgsz argument controls the inference dimensions【detect.py†L74-L75】. Reducing the default 640x640 to 320 or 416 decreases FLOPs proportionally, often enabling real-time inference on modest CPUs with minimal accuracy degradation.
python detect.py --weights yolov5s.pt --imgsz 416 --half
For edge deployment, maintain a batch size of 1 to avoid memory copies and buffering overhead. The detect.py script explicitly sets bs = 1 for single-image inference pipelines【detect.py†L70-L71】, which is the optimal configuration for streaming video inputs on memory-constrained devices.
Execute Warm-Up and Minimize Post-Processing
Before measuring performance or processing production data, trigger the warm-up routine to force kernel compilation and memory allocation. The model.warmup() method is called in detect.py before the inference loop begins【detect.py†L81-L83】, ensuring steady-state latency by eliminating first-run compilation overhead on GPUs and TensorRT engines.
Additionally, reduce CPU overhead by disabling unnecessary post-processing when only raw detections are required:
- Use
--nosaveto skip disk I/O for output images - Use
--save-txtonly if bounding box coordinates must be persisted - Consider modifying the code to skip
non_max_suppressionif your application can tolerate redundant detections or handles filtering downstream
Implement INT8 Quantization for Extreme Compression
While not built into the main repository, the ONNX export path enables external quantization workflows. Convert your exported ONNX model to 8-bit integer format using onnxruntime quantization tools for deployment on CPU-only edge devices where FP16 support is unavailable.
import onnxruntime as ort
from onnxruntime.quantization import quantize_dynamic, QuantType
# Convert FP32 ONNX to INT8
quantize_dynamic(
model_input='yolov5s.onnx',
model_output='yolov5s_int8.onnx',
weight_type=QuantType.QInt8
)
# Inference session
sess = ort.InferenceSession('yolov5s_int8.onnx')
input_name = sess.get_inputs()[0].name
outputs = sess.run(None, {input_name: img_array})
INT8 quantization reduces model size by 75% and significantly increases inference speed on x86 and ARM CPUs at the cost of minor precision loss.
Load Models via Python API for Embedded Integration
For custom edge applications requiring direct Python integration, use DetectMultiBackend with exported TorchScript or ONNX models to bypass the command-line interface overhead.
from pathlib import Path
from models.common import DetectMultiBackend
from utils.torch_utils import select_device
import torch
import cv2
# Initialize backend with TorchScript for ultra-low latency
device = select_device('cpu')
model = DetectMultiBackend(
weights=Path('yolov5s.torchscript'),
device=device,
dnn=False,
data=None,
fp16=True
)
# Preprocess image
img = cv2.imread('bus.jpg')
img = cv2.resize(img, (640, 640))
img_tensor = torch.from_numpy(img).permute(2, 0, 1).unsqueeze(0).float() / 255.0
if model.fp16:
img_tensor = img_tensor.half()
# Inference
pred = model(img_tensor, augment=False, visualize=False)
This pattern is essential for robotics applications where the inference engine must run as a library within a larger control system.
Summary
- Enable
--halfindetect.pyto activate FP16 inference throughDetectMultiBackend, reducing memory bandwidth by 50% on compatible hardware. - Export to specialized backends using
export.py—choose TensorRT for NVIDIA Jetson, OpenVINO for Intel VPUs, CoreML for Apple devices, or ONNX+OpenCV DNN for generic ARM CPUs. - Reduce resolution with
--imgszto 320-416 for edge devices, and maintain batch size of 1 to minimize memory copies. - Call
model.warmup()before production inference to ensure kernels are compiled and latency is consistent. - Quantize to INT8 via ONNX Runtime for maximum compression on CPU-only edge hardware where GPU acceleration is unavailable.
Frequently Asked Questions
What is the fastest backend for YOLOv5 on NVIDIA Jetson devices?
TensorRT provides the lowest latency and highest throughput on NVIDIA Jetson platforms. Export your model using python export.py --weights yolov5s.pt --include engine --half, then run inference with python detect.py --weights yolov5s.engine --half --device 0. The TensorRT engine performs layer fusion and kernel optimization specific to the Jetson GPU architecture.
How do I run YOLOv5 on a Raspberry Pi without GPU acceleration?
Export the model to ONNX format using export.py, then use the OpenCV DNN backend. Run python detect.py --weights yolov5s.onnx --dnn --half --device cpu --imgsz 416. The OpenCV DNN backend eliminates the PyTorch dependency and runs efficiently on ARM Cortex CPUs, while reducing the input resolution to 416 or 320 maintains real-time performance.
Does FP16 inference work on all edge devices?
No, FP16 requires hardware support. It works natively on NVIDIA GPUs (Jetson), Apple Neural Engine (via CoreML), and some modern ARM processors with FP16 SIMD instructions. For older CPUs without FP16 support, use INT8 quantization via ONNX Runtime or stick with FP32 to avoid emulation overhead.
Where in the codebase does the warm-up functionality reside?
The warm-up routine is implemented in utils/general.py and invoked in detect.py at lines 81-83【detect.py†L81-L83】. The code calls model.warmup(imgsz=(1 if pt else bs, 3, *imgsz)) to execute a dummy forward pass, forcing CUDA kernel compilation and memory allocation before the actual inference loop begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →