# How to Use TensorRT for YOLOv5 Inference Optimization

> Optimize YOLOv5 inference speed with TensorRT. Convert your PyTorch model using export.py for significant GPU acceleration with the same API.

- Repository: [Ultralytics/yolov5](https://github.com/ultralytics/yolov5)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Convert your YOLOv5 PyTorch model to a TensorRT engine using [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) and load it via `DetectMultiBackend` to achieve significant GPU acceleration while maintaining the same inference API.**

YOLOv5 supports NVIDIA TensorRT to accelerate object detection on NVIDIA GPUs. According to the ultralytics/yolov5 source code, the framework provides built-in utilities in [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) to export trained models to TensorRT format and seamlessly load them for optimized inference using a unified backend abstraction.

## Exporting YOLOv5 Models to TensorRT Format

The TensorRT export pipeline converts a PyTorch `.pt` file into an optimized `.engine` file through an ONNX intermediate. In [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py), the `export_engine` function (lines 84-108) handles this conversion using the TensorRT Python API.

The export process first generates an ONNX representation of the model, then instantiates a `trt.Builder` to construct the engine. The implementation checks TensorRT version compatibility using `check_version(trt.__version__, "7.0.0", hard=True)`, with additional validation for TensorRT 8+ requiring version `>= 8.0.0`.

### Configuring FP16 Precision and Dynamic Shapes

To enable **FP16 inference**, set `half=True` during export. The code verifies GPU support via `builder.platform_has_fast_fp16` before setting the `trt.BuilderFlag.FP16` flag. For **dynamic input shapes**, pass `dynamic=True` to create an optimization profile using `builder.create_optimization_profile()`, allowing variable batch sizes and image dimensions at runtime.

### Workspace Size and Build Cache

The `workspace` parameter (in GB) controls the maximum memory pool available to the builder. For faster iterative builds, supply a `cache` file path to enable TensorRT's timing cache, which reuses optimized layer implementations across exports.

## Running Inference with TensorRT Engines

Once exported, load the `.engine` file using `DetectMultiBackend` from [`models/common.py`](https://github.com/ultralytics/yolov5/blob/main/models/common.py) (lines 531-545). This class detects the `.engine` suffix and initializes a TensorRT runtime, deserializes the engine, and creates execution bindings that map input tensors to GPU memory.

The implementation abstracts TensorRT specifics behind the standard `.forward()` API. When you call `model(img)`, `DetectMultiBackend` internally invokes `engine.create_execution_context()` and `execute_v2()` to run inference, returning tensors compatible with the same post-processing pipeline used for PyTorch models.

## Command-Line Workflow

Export and run inference via the CLI tools provided in the repository:

```bash

# Export yolov5s.pt to TensorRT with FP16 and dynamic shapes

python export.py --weights yolov5s.pt \
                 --include engine \
                 --half \
                 --dynamic \
                 --workspace 8

```

After export, run detection using the optimized engine:

```bash
python detect.py --weights yolov5s.engine \
                 --source data/images \
                 --half \
                 --device 0

```

## Python API Implementation

For custom applications, use the programmatic API to load and run the TensorRT engine:

```python
import torch
from models.common import DetectMultiBackend
from utils.general import non_max_suppression

# Load TensorRT engine

model = DetectMultiBackend(
    weights='yolov5s.engine', 
    device='0', 
    dnn=False, 
    data='data/coco128.yaml'
)

# Warmup to allocate GPU buffers

model.warmup(imgsz=(1, 3, 640, 640))

# Prepare input tensor (NCHW format)

img = torch.randn(1, 3, 640, 640).cuda()

# TensorRT inference

pred = model(img, augment=False, visualize=False)
pred = non_max_suppression(pred, conf_thres=0.25, iou_thres=0.45)[0]

print(f'Detections: {pred}')

```

## Summary

- **Export pipeline**: Use [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) with `export_engine` (lines 84-108) to convert PyTorch models to TensorRT format via ONNX.
- **Version requirements**: TensorRT >= 7.0.0 is mandatory, with additional checks for TensorRT 8+.
- **Optimization features**: Enable FP16 precision, dynamic shapes, and workspace limits during export for maximum performance.
- **Inference abstraction**: `DetectMultiBackend` in [`models/common.py`](https://github.com/ultralytics/yolov5/blob/main/models/common.py) (lines 531-545) handles TensorRT runtime initialization and provides a unified API compatible with existing post-processing code.

## Frequently Asked Questions

### What TensorRT version is required for YOLOv5?

YOLOv5 requires TensorRT version 7.0.0 or higher, enforced via `check_version(trt.__version__, "7.0.0", hard=True)` in [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py). For TensorRT 8 and above, the code performs additional validation to ensure version compatibility before proceeding with engine construction.

### How do I enable FP16 precision in TensorRT for YOLOv5?

Pass `--half` to [`export.py`](https://github.com/ultralytics/yolov5/blob/main/export.py) or set `half=True` when calling `export_engine()`. The implementation automatically checks `builder.platform_has_fast_fp16` to verify your GPU supports fast FP16 operations, then sets the `trt.BuilderFlag.FP16` flag during engine compilation.

### Does YOLOv5 TensorRT support dynamic batch sizes?

Yes. Specify `--dynamic` during export to create an optimization profile via `builder.create_optimization_profile()`. This allows the engine to accept variable input shapes at runtime, though you must specify the minimum, optimal, and maximum shape bounds during the export process.

### Can I use a TensorRT engine on different GPUs or TensorRT versions?

No. TensorRT engines are hardware-specific and tied to the exact GPU architecture, TensorRT version, and library dependencies used during export. You must re-export the `.engine` file on each target deployment device to ensure compatibility.