# How to Export a YOLOv5 Model to ONNX Format for Inference

> Easily export your YOLOv5 model to ONNX format using the export.py script. Optimize your PyTorch checkpoints for fast inference with ONNX Runtime, OpenCV DNN, and TensorRT.

- Repository: [Ultralytics/yolov5](https://github.com/ultralytics/yolov5)
- Tags: how-to-guide
- Published: 2026-03-06

---

**You can export a YOLOv5 model to ONNX format by running the [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) script with the `--include onnx` flag, which internally calls the `export_onnx()` function to convert PyTorch `.pt` checkpoints into optimized ONNX graphs for deployment across ONNX Runtime, OpenCV DNN, and TensorRT.**

The ultralytics/yolov5 repository provides a dedicated export pipeline that streamlines the conversion of trained PyTorch models to production-ready formats. Understanding how to export a YOLOv5 model to ONNX format unlocks framework-agnostic inference capabilities, allowing you to deploy object detection models on edge devices and optimized serving platforms without requiring PyTorch dependencies.

## Prerequisites for ONNX Export

Before initiating the conversion, install the required dependencies. The [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) script verifies package availability before attempting conversion and requires these libraries for graph optimization and validation.

```bash
pip install onnx onnx-simplifier onnxruntime

```

## The Export Pipeline Architecture

The export process follows a three-stage architecture defined in [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py). Understanding these stages helps troubleshoot conversion errors and optimize output for specific deployment targets.

### Model Loading with attempt_load()

In [`models/experimental.py`](https://github.com/ultralytics/yolov5/blob/main/models/experimental.py), the `attempt_load()` function parses the PyTorch checkpoint and reconstructs the model graph on the target device. This function handles version compatibility and loads weights into the YOLOv5 architecture before export begins.

### Dummy Input Preparation

The export pipeline creates a dummy tensor matching the training image size—typically `(1, 3, 640, 640)`—to trace the model's computation graph during the conversion process. This tensor flows through the network to define the ONNX graph structure.

### The export_onnx() Function

The core conversion logic resides in `export_onnx()` (lines 79–106 of [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py)). This function wraps `torch.onnx.export` and manages ONNX-specific configurations through the following signature:

```python
def export_onnx(
    model,               # torch.nn.Module – loaded YOLOv5 model

    im,                  # torch.Tensor – dummy input (B, C, H, W)

    file,                # pathlib.Path – destination "model.onnx"

    opset,               # int – ONNX opset version (e.g., 12)

    dynamic,             # bool – add dynamic batch/height/width axes

    simplify,            # bool – run onnx-simplifier after export

    prefix=colorstr("ONNX:") )

```

Key implementation details include:

- **Dynamic axes** configuration for batch, height, and width dimensions when `dynamic=True`, enabling variable input sizes
- **Metadata injection** storing `stride` and class `names` inside the ONNX protobuf for downstream post-processing
- **Model validation** using `onnx.checker.check_model` immediately after file generation to ensure graph integrity
- **Graph simplification** via `onnx-slim` (or `onnx-simplifier`) when `simplify=True`, which prunes redundant nodes to reduce file size and improve inference speed

## Export Methods

You can initiate the export via command line for quick conversions or programmatically via the Python API for integration into custom training pipelines.

### Command-Line Interface

The fastest way to export a YOLOv5 model to ONNX format uses the built-in CLI. Execute the following from the repository root to generate an optimized model with dynamic axes:

```bash
python export.py --weights yolov5s.pt --include onnx --dynamic --simplify

```

This command loads `yolov5s.pt` using `attempt_load()`, exports to `yolov5s.onnx` with opset 12, enables dynamic batching and input dimensions, and applies graph simplification to optimize the model structure.

### Python API Integration

For automated deployment workflows or custom training pipelines, import the export utilities directly:

```python
from pathlib import Path
import torch
from utils.torch_utils import select_device
from models.experimental import attempt_load
from export import export_onnx

# Load checkpoint

weights = 'yolov5s.pt'
device = select_device('cpu')  # or '0' for CUDA

model = attempt_load(weights, device=device)

# Create dummy input (batch, channels, height, width)

dummy = torch.zeros(1, 3, 640, 640).to(device)

# Export with dynamic axes and simplification

export_onnx(
    model=model,
    im=dummy,
    file=Path('yolov5s.onnx'),
    opset=12,
    dynamic=True,
    simplify=True
)

```

## Running Inference with Exported ONNX Models

Once converted, ONNX models support inference through YOLOv5's native detection tools or custom runtime implementations.

### Using detect.py

The [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py) script automatically detects the `.onnx` extension and routes inference through the appropriate backend (lines 27–33). This allows immediate validation without modifying your existing detection pipeline:

```bash
python detect.py --weights yolov5s.onnx --source data/images/bus.jpg

```

### Custom ONNX Runtime Implementation

For production applications requiring direct control over pre-processing and batching, use ONNX Runtime:

```python
import onnxruntime as ort
import numpy as np
import cv2

# Initialize session

session = ort.InferenceSession('yolov5s.onnx')
input_name = session.get_inputs()[0].name

# Pre-process image (BGR → RGB, resize, normalize)

img = cv2.imread('data/images/bus.jpg')
img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
img = cv2.resize(img, (640, 640))
img = img.astype(np.float32) / 255.0
img = np.transpose(img, (2, 0, 1))[np.newaxis, ...]  # (1, 3, 640, 640)

# Run inference

outputs = session.run(None, {input_name: img})
detections = outputs[0]  # Bounding boxes, confidences, and class IDs

```

## Summary

- **Use [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py)** as the primary entry point for converting PyTorch checkpoints to ONNX format in the ultralytics/yolov5 repository.
- **Call `export_onnx()`** programmatically when integrating the conversion into custom workflows; this function handles dynamic axes, metadata embedding, and optional graph simplification.
- **Leverage [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py)** for immediate validation of exported models—it automatically selects ONNX Runtime or OpenCV DNN based on the file extension.
- **Install dependencies** (`onnx`, `onnx-simplifier`, `onnxruntime`) before attempting export to ensure the simplification and validation steps succeed.

## Frequently Asked Questions

### What is the difference between dynamic and static ONNX export in YOLOv5?

**Dynamic export** (enabled with `--dynamic`) configures the ONNX model to accept variable batch sizes and input dimensions by setting dynamic axes for batch, height, and width in the `export_onnx()` function. **Static export** produces a model fixed to the input shape specified during conversion (default 640×640). Use dynamic export when your inference pipeline processes varying image sizes or batch counts; use static export for maximum performance optimization in fixed-input scenarios.

### Does the ONNX export include NMS (Non-Maximum Suppression)?

**Standard YOLOv5 ONNX exports do not include built-in NMS operations.** The exported model outputs raw detection tensors (bounding boxes, objectness scores, and class probabilities) that require post-processing. You must implement NMS separately in your inference code, or use the EfficientNMS plugin if exporting to TensorRT through ONNX. The [`detect.py`](https://github.com/ultralytics/yolov5/blob/main/detect.py) script handles this post-processing automatically when running inference.

### Which ONNX opset version should I use for YOLOv5?

**YOLOv5's `export_onnx()` function defaults to opset 12**, which provides broad compatibility across inference engines including ONNX Runtime, OpenCV DNN, and older TensorRT versions. While newer opsets (14–17) are supported by passing `--opset N` to the CLI, opset 12 remains the recommended default for maximum deployment compatibility unless you require specific operators introduced in later versions.

### Can I export quantized or INT8 ONNX models using the YOLOv5 export script?

**The standard [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) script focuses on FP32 and FP16 ONNX conversion.** For INT8 quantization, YOLOv5 requires separate calibration and export workflows typically handled through TensorRT's explicit quantization or ONNX Runtime's quantization tools after exporting the standard FP32 model. The `export_onnx()` function does not directly support INT8 quantization flags; use TensorRT export (`--include engine`) for optimized INT8 inference on NVIDIA hardware.