How to Export a YOLOv5 Model to ONNX Format for Inference
You can export a YOLOv5 model to ONNX format by running the export.py script with the --include onnx flag, which internally calls the export_onnx() function to convert PyTorch .pt checkpoints into optimized ONNX graphs for deployment across ONNX Runtime, OpenCV DNN, and TensorRT.
The ultralytics/yolov5 repository provides a dedicated export pipeline that streamlines the conversion of trained PyTorch models to production-ready formats. Understanding how to export a YOLOv5 model to ONNX format unlocks framework-agnostic inference capabilities, allowing you to deploy object detection models on edge devices and optimized serving platforms without requiring PyTorch dependencies.
Prerequisites for ONNX Export
Before initiating the conversion, install the required dependencies. The export.py script verifies package availability before attempting conversion and requires these libraries for graph optimization and validation.
pip install onnx onnx-simplifier onnxruntime
The Export Pipeline Architecture
The export process follows a three-stage architecture defined in export.py. Understanding these stages helps troubleshoot conversion errors and optimize output for specific deployment targets.
Model Loading with attempt_load()
In models/experimental.py, the attempt_load() function parses the PyTorch checkpoint and reconstructs the model graph on the target device. This function handles version compatibility and loads weights into the YOLOv5 architecture before export begins.
Dummy Input Preparation
The export pipeline creates a dummy tensor matching the training image size—typically (1, 3, 640, 640)—to trace the model's computation graph during the conversion process. This tensor flows through the network to define the ONNX graph structure.
The export_onnx() Function
The core conversion logic resides in export_onnx() (lines 79–106 of export.py). This function wraps torch.onnx.export and manages ONNX-specific configurations through the following signature:
def export_onnx(
model, # torch.nn.Module – loaded YOLOv5 model
im, # torch.Tensor – dummy input (B, C, H, W)
file, # pathlib.Path – destination "model.onnx"
opset, # int – ONNX opset version (e.g., 12)
dynamic, # bool – add dynamic batch/height/width axes
simplify, # bool – run onnx-simplifier after export
prefix=colorstr("ONNX:") )
Key implementation details include:
- Dynamic axes configuration for batch, height, and width dimensions when
dynamic=True, enabling variable input sizes - Metadata injection storing
strideand classnamesinside the ONNX protobuf for downstream post-processing - Model validation using
onnx.checker.check_modelimmediately after file generation to ensure graph integrity - Graph simplification via
onnx-slim(oronnx-simplifier) whensimplify=True, which prunes redundant nodes to reduce file size and improve inference speed
Export Methods
You can initiate the export via command line for quick conversions or programmatically via the Python API for integration into custom training pipelines.
Command-Line Interface
The fastest way to export a YOLOv5 model to ONNX format uses the built-in CLI. Execute the following from the repository root to generate an optimized model with dynamic axes:
python export.py --weights yolov5s.pt --include onnx --dynamic --simplify
This command loads yolov5s.pt using attempt_load(), exports to yolov5s.onnx with opset 12, enables dynamic batching and input dimensions, and applies graph simplification to optimize the model structure.
Python API Integration
For automated deployment workflows or custom training pipelines, import the export utilities directly:
from pathlib import Path
import torch
from utils.torch_utils import select_device
from models.experimental import attempt_load
from export import export_onnx
# Load checkpoint
weights = 'yolov5s.pt'
device = select_device('cpu') # or '0' for CUDA
model = attempt_load(weights, device=device)
# Create dummy input (batch, channels, height, width)
dummy = torch.zeros(1, 3, 640, 640).to(device)
# Export with dynamic axes and simplification
export_onnx(
model=model,
im=dummy,
file=Path('yolov5s.onnx'),
opset=12,
dynamic=True,
simplify=True
)
Running Inference with Exported ONNX Models
Once converted, ONNX models support inference through YOLOv5's native detection tools or custom runtime implementations.
Using detect.py
The detect.py script automatically detects the .onnx extension and routes inference through the appropriate backend (lines 27–33). This allows immediate validation without modifying your existing detection pipeline:
python detect.py --weights yolov5s.onnx --source data/images/bus.jpg
Custom ONNX Runtime Implementation
For production applications requiring direct control over pre-processing and batching, use ONNX Runtime:
import onnxruntime as ort
import numpy as np
import cv2
# Initialize session
session = ort.InferenceSession('yolov5s.onnx')
input_name = session.get_inputs()[0].name
# Pre-process image (BGR → RGB, resize, normalize)
img = cv2.imread('data/images/bus.jpg')
img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
img = cv2.resize(img, (640, 640))
img = img.astype(np.float32) / 255.0
img = np.transpose(img, (2, 0, 1))[np.newaxis, ...] # (1, 3, 640, 640)
# Run inference
outputs = session.run(None, {input_name: img})
detections = outputs[0] # Bounding boxes, confidences, and class IDs
Summary
- Use
export.pyas the primary entry point for converting PyTorch checkpoints to ONNX format in the ultralytics/yolov5 repository. - Call
export_onnx()programmatically when integrating the conversion into custom workflows; this function handles dynamic axes, metadata embedding, and optional graph simplification. - Leverage
detect.pyfor immediate validation of exported models—it automatically selects ONNX Runtime or OpenCV DNN based on the file extension. - Install dependencies (
onnx,onnx-simplifier,onnxruntime) before attempting export to ensure the simplification and validation steps succeed.
Frequently Asked Questions
What is the difference between dynamic and static ONNX export in YOLOv5?
Dynamic export (enabled with --dynamic) configures the ONNX model to accept variable batch sizes and input dimensions by setting dynamic axes for batch, height, and width in the export_onnx() function. Static export produces a model fixed to the input shape specified during conversion (default 640×640). Use dynamic export when your inference pipeline processes varying image sizes or batch counts; use static export for maximum performance optimization in fixed-input scenarios.
Does the ONNX export include NMS (Non-Maximum Suppression)?
Standard YOLOv5 ONNX exports do not include built-in NMS operations. The exported model outputs raw detection tensors (bounding boxes, objectness scores, and class probabilities) that require post-processing. You must implement NMS separately in your inference code, or use the EfficientNMS plugin if exporting to TensorRT through ONNX. The detect.py script handles this post-processing automatically when running inference.
Which ONNX opset version should I use for YOLOv5?
YOLOv5's export_onnx() function defaults to opset 12, which provides broad compatibility across inference engines including ONNX Runtime, OpenCV DNN, and older TensorRT versions. While newer opsets (14–17) are supported by passing --opset N to the CLI, opset 12 remains the recommended default for maximum deployment compatibility unless you require specific operators introduced in later versions.
Can I export quantized or INT8 ONNX models using the YOLOv5 export script?
The standard export.py script focuses on FP32 and FP16 ONNX conversion. For INT8 quantization, YOLOv5 requires separate calibration and export workflows typically handled through TensorRT's explicit quantization or ONNX Runtime's quantization tools after exporting the standard FP32 model. The export_onnx() function does not directly support INT8 quantization flags; use TensorRT export (--include engine) for optimized INT8 inference on NVIDIA hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →