How to Use TensorRT for YOLOv5 Inference Optimization

Convert your YOLOv5 PyTorch model to a TensorRT engine using export.py and load it via DetectMultiBackend to achieve significant GPU acceleration while maintaining the same inference API.

YOLOv5 supports NVIDIA TensorRT to accelerate object detection on NVIDIA GPUs. According to the ultralytics/yolov5 source code, the framework provides built-in utilities in export.py to export trained models to TensorRT format and seamlessly load them for optimized inference using a unified backend abstraction.

Exporting YOLOv5 Models to TensorRT Format

The TensorRT export pipeline converts a PyTorch .pt file into an optimized .engine file through an ONNX intermediate. In export.py, the export_engine function (lines 84-108) handles this conversion using the TensorRT Python API.

The export process first generates an ONNX representation of the model, then instantiates a trt.Builder to construct the engine. The implementation checks TensorRT version compatibility using check_version(trt.__version__, "7.0.0", hard=True), with additional validation for TensorRT 8+ requiring version >= 8.0.0.

Configuring FP16 Precision and Dynamic Shapes

To enable FP16 inference, set half=True during export. The code verifies GPU support via builder.platform_has_fast_fp16 before setting the trt.BuilderFlag.FP16 flag. For dynamic input shapes, pass dynamic=True to create an optimization profile using builder.create_optimization_profile(), allowing variable batch sizes and image dimensions at runtime.

Workspace Size and Build Cache

The workspace parameter (in GB) controls the maximum memory pool available to the builder. For faster iterative builds, supply a cache file path to enable TensorRT's timing cache, which reuses optimized layer implementations across exports.

Running Inference with TensorRT Engines

Once exported, load the .engine file using DetectMultiBackend from models/common.py (lines 531-545). This class detects the .engine suffix and initializes a TensorRT runtime, deserializes the engine, and creates execution bindings that map input tensors to GPU memory.

The implementation abstracts TensorRT specifics behind the standard .forward() API. When you call model(img), DetectMultiBackend internally invokes engine.create_execution_context() and execute_v2() to run inference, returning tensors compatible with the same post-processing pipeline used for PyTorch models.

Command-Line Workflow

Export and run inference via the CLI tools provided in the repository:


# Export yolov5s.pt to TensorRT with FP16 and dynamic shapes

python export.py --weights yolov5s.pt \
                 --include engine \
                 --half \
                 --dynamic \
                 --workspace 8

After export, run detection using the optimized engine:

python detect.py --weights yolov5s.engine \
                 --source data/images \
                 --half \
                 --device 0

Python API Implementation

For custom applications, use the programmatic API to load and run the TensorRT engine:

import torch
from models.common import DetectMultiBackend
from utils.general import non_max_suppression

# Load TensorRT engine

model = DetectMultiBackend(
    weights='yolov5s.engine', 
    device='0', 
    dnn=False, 
    data='data/coco128.yaml'
)

# Warmup to allocate GPU buffers

model.warmup(imgsz=(1, 3, 640, 640))

# Prepare input tensor (NCHW format)

img = torch.randn(1, 3, 640, 640).cuda()

# TensorRT inference

pred = model(img, augment=False, visualize=False)
pred = non_max_suppression(pred, conf_thres=0.25, iou_thres=0.45)[0]

print(f'Detections: {pred}')

Summary

  • Export pipeline: Use export.py with export_engine (lines 84-108) to convert PyTorch models to TensorRT format via ONNX.
  • Version requirements: TensorRT >= 7.0.0 is mandatory, with additional checks for TensorRT 8+.
  • Optimization features: Enable FP16 precision, dynamic shapes, and workspace limits during export for maximum performance.
  • Inference abstraction: DetectMultiBackend in models/common.py (lines 531-545) handles TensorRT runtime initialization and provides a unified API compatible with existing post-processing code.

Frequently Asked Questions

What TensorRT version is required for YOLOv5?

YOLOv5 requires TensorRT version 7.0.0 or higher, enforced via check_version(trt.__version__, "7.0.0", hard=True) in export.py. For TensorRT 8 and above, the code performs additional validation to ensure version compatibility before proceeding with engine construction.

How do I enable FP16 precision in TensorRT for YOLOv5?

Pass --half to export.py or set half=True when calling export_engine(). The implementation automatically checks builder.platform_has_fast_fp16 to verify your GPU supports fast FP16 operations, then sets the trt.BuilderFlag.FP16 flag during engine compilation.

Does YOLOv5 TensorRT support dynamic batch sizes?

Yes. Specify --dynamic during export to create an optimization profile via builder.create_optimization_profile(). This allows the engine to accept variable input shapes at runtime, though you must specify the minimum, optimal, and maximum shape bounds during the export process.

Can I use a TensorRT engine on different GPUs or TensorRT versions?

No. TensorRT engines are hardware-specific and tied to the exact GPU architecture, TensorRT version, and library dependencies used during export. You must re-export the .engine file on each target deployment device to ensure compatibility.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →