How to Export a YOLOv5 Model to ONNX Format for Inference

You can export a YOLOv5 model to ONNX format by running the export.py script with the --include onnx flag, which internally calls the export_onnx() function to convert PyTorch .pt checkpoints into optimized ONNX graphs for deployment across ONNX Runtime, OpenCV DNN, and TensorRT.

The ultralytics/yolov5 repository provides a dedicated export pipeline that streamlines the conversion of trained PyTorch models to production-ready formats. Understanding how to export a YOLOv5 model to ONNX format unlocks framework-agnostic inference capabilities, allowing you to deploy object detection models on edge devices and optimized serving platforms without requiring PyTorch dependencies.

Prerequisites for ONNX Export

Before initiating the conversion, install the required dependencies. The export.py script verifies package availability before attempting conversion and requires these libraries for graph optimization and validation.

pip install onnx onnx-simplifier onnxruntime

The Export Pipeline Architecture

The export process follows a three-stage architecture defined in export.py. Understanding these stages helps troubleshoot conversion errors and optimize output for specific deployment targets.

Model Loading with attempt_load()

In models/experimental.py, the attempt_load() function parses the PyTorch checkpoint and reconstructs the model graph on the target device. This function handles version compatibility and loads weights into the YOLOv5 architecture before export begins.

Dummy Input Preparation

The export pipeline creates a dummy tensor matching the training image size—typically (1, 3, 640, 640)—to trace the model's computation graph during the conversion process. This tensor flows through the network to define the ONNX graph structure.

The export_onnx() Function

The core conversion logic resides in export_onnx() (lines 79–106 of export.py). This function wraps torch.onnx.export and manages ONNX-specific configurations through the following signature:

def export_onnx(
    model,               # torch.nn.Module – loaded YOLOv5 model

    im,                  # torch.Tensor – dummy input (B, C, H, W)

    file,                # pathlib.Path – destination "model.onnx"

    opset,               # int – ONNX opset version (e.g., 12)

    dynamic,             # bool – add dynamic batch/height/width axes

    simplify,            # bool – run onnx-simplifier after export

    prefix=colorstr("ONNX:") )

Key implementation details include:

  • Dynamic axes configuration for batch, height, and width dimensions when dynamic=True, enabling variable input sizes
  • Metadata injection storing stride and class names inside the ONNX protobuf for downstream post-processing
  • Model validation using onnx.checker.check_model immediately after file generation to ensure graph integrity
  • Graph simplification via onnx-slim (or onnx-simplifier) when simplify=True, which prunes redundant nodes to reduce file size and improve inference speed

Export Methods

You can initiate the export via command line for quick conversions or programmatically via the Python API for integration into custom training pipelines.

Command-Line Interface

The fastest way to export a YOLOv5 model to ONNX format uses the built-in CLI. Execute the following from the repository root to generate an optimized model with dynamic axes:

python export.py --weights yolov5s.pt --include onnx --dynamic --simplify

This command loads yolov5s.pt using attempt_load(), exports to yolov5s.onnx with opset 12, enables dynamic batching and input dimensions, and applies graph simplification to optimize the model structure.

Python API Integration

For automated deployment workflows or custom training pipelines, import the export utilities directly:

from pathlib import Path
import torch
from utils.torch_utils import select_device
from models.experimental import attempt_load
from export import export_onnx

# Load checkpoint

weights = 'yolov5s.pt'
device = select_device('cpu')  # or '0' for CUDA

model = attempt_load(weights, device=device)

# Create dummy input (batch, channels, height, width)

dummy = torch.zeros(1, 3, 640, 640).to(device)

# Export with dynamic axes and simplification

export_onnx(
    model=model,
    im=dummy,
    file=Path('yolov5s.onnx'),
    opset=12,
    dynamic=True,
    simplify=True
)

Running Inference with Exported ONNX Models

Once converted, ONNX models support inference through YOLOv5's native detection tools or custom runtime implementations.

Using detect.py

The detect.py script automatically detects the .onnx extension and routes inference through the appropriate backend (lines 27–33). This allows immediate validation without modifying your existing detection pipeline:

python detect.py --weights yolov5s.onnx --source data/images/bus.jpg

Custom ONNX Runtime Implementation

For production applications requiring direct control over pre-processing and batching, use ONNX Runtime:

import onnxruntime as ort
import numpy as np
import cv2

# Initialize session

session = ort.InferenceSession('yolov5s.onnx')
input_name = session.get_inputs()[0].name

# Pre-process image (BGR → RGB, resize, normalize)

img = cv2.imread('data/images/bus.jpg')
img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
img = cv2.resize(img, (640, 640))
img = img.astype(np.float32) / 255.0
img = np.transpose(img, (2, 0, 1))[np.newaxis, ...]  # (1, 3, 640, 640)

# Run inference

outputs = session.run(None, {input_name: img})
detections = outputs[0]  # Bounding boxes, confidences, and class IDs

Summary

  • Use export.py as the primary entry point for converting PyTorch checkpoints to ONNX format in the ultralytics/yolov5 repository.
  • Call export_onnx() programmatically when integrating the conversion into custom workflows; this function handles dynamic axes, metadata embedding, and optional graph simplification.
  • Leverage detect.py for immediate validation of exported models—it automatically selects ONNX Runtime or OpenCV DNN based on the file extension.
  • Install dependencies (onnx, onnx-simplifier, onnxruntime) before attempting export to ensure the simplification and validation steps succeed.

Frequently Asked Questions

What is the difference between dynamic and static ONNX export in YOLOv5?

Dynamic export (enabled with --dynamic) configures the ONNX model to accept variable batch sizes and input dimensions by setting dynamic axes for batch, height, and width in the export_onnx() function. Static export produces a model fixed to the input shape specified during conversion (default 640×640). Use dynamic export when your inference pipeline processes varying image sizes or batch counts; use static export for maximum performance optimization in fixed-input scenarios.

Does the ONNX export include NMS (Non-Maximum Suppression)?

Standard YOLOv5 ONNX exports do not include built-in NMS operations. The exported model outputs raw detection tensors (bounding boxes, objectness scores, and class probabilities) that require post-processing. You must implement NMS separately in your inference code, or use the EfficientNMS plugin if exporting to TensorRT through ONNX. The detect.py script handles this post-processing automatically when running inference.

Which ONNX opset version should I use for YOLOv5?

YOLOv5's export_onnx() function defaults to opset 12, which provides broad compatibility across inference engines including ONNX Runtime, OpenCV DNN, and older TensorRT versions. While newer opsets (14–17) are supported by passing --opset N to the CLI, opset 12 remains the recommended default for maximum deployment compatibility unless you require specific operators introduced in later versions.

Can I export quantized or INT8 ONNX models using the YOLOv5 export script?

The standard export.py script focuses on FP32 and FP16 ONNX conversion. For INT8 quantization, YOLOv5 requires separate calibration and export workflows typically handled through TensorRT's explicit quantization or ONNX Runtime's quantization tools after exporting the standard FP32 model. The export_onnx() function does not directly support INT8 quantization flags; use TensorRT export (--include engine) for optimized INT8 inference on NVIDIA hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →