How to Export an RF-DETR Model to TensorRT Format

RF-DETR provides a built-in export pipeline that converts trained detection models into optimized TensorRT engines via an intermediate ONNX representation, accessible through both Python API and CLI interfaces.

Exporting RF-DETR to TensorRT enables high-performance inference on NVIDIA GPUs. According to the roboflow/rf-detr source code, the export functionality is implemented in the core RFDETR class and orchestrated through dedicated export modules, eliminating the need for manual conversion scripts.

Installing TensorRT Dependencies

TensorRT support is optional to keep the base installation lightweight. Install the required dependencies using the tensorrt extra, which pulls in the polygraphy wrapper library:

pip install "rfdetr[tensorrt]"

This extra ensures the TensorRT Python API is available through polygraphy, allowing in-process engine building without launching external trtexec subprocesses.

Exporting via Python API

The RFDETR.export() method in src/rfdetr/detr.py drives the entire conversion pipeline. This method handles ONNX generation and TensorRT engine serialization in a single call.

Loading the Model

Instantiate a pre-trained model or load a custom checkpoint before export:

import rfdetr

# Load a pre-trained small variant

model = rfdetr.detr.RFDETRSmall()

# Optional: configure input resolution (must be set before export)

model.resolution = (640, 640)
model.eval()  # Ensure deterministic export mode

Configuring Export Parameters

The export() method accepts specific parameters for TensorRT optimization:

  • format: Must be set to "tensorrt" to trigger the TensorRT builder
  • fp16: Boolean flag for FP16 precision (default True). Set to False for FP32 engines when targeting TensorRT builds without FP16 support
  • output_dir: Directory path for saving the ONNX intermediate and final .engine file
  • engine_name: Base filename for the serialized engine (e.g., "rfdetr_small" generates rfdetr_small-fp16.engine)
  • backbone_only: Export only the backbone architecture (useful for downstream fine-tuning pipelines)

Executing the Export

Run the export with your desired configuration:

engine_path = model.export(
    format="tensorrt",
    fp16=True,                # Set False for FP32 precision

    output_dir="./exported",  # Creates directory if missing

    engine_name="rfdetr_small"
)

print(f"TensorRT engine saved to: {engine_path}")

Under the hood, RFDETR.export() in src/rfdetr/detr.py performs three operations:

  1. ONNX Generation – Calls torch.onnx.export to create an ONNX graph, implemented in src/rfdetr/export/main.py
  2. Engine Building – Passes the ONNX file to src/rfdetr/export/_tensorrt.py, which uses polygraphy's build_engine API to create the serialized TensorRT engine
  3. Filename Resolution – Uses utilities in src/rfdetr/export/_naming.py to generate consistent filenames (e.g., <stem>-fp16.engine or <stem>-fp32.engine)

Exporting via Command Line

The same functionality is exposed through the CLI entry point in src/rfdetr/__main__.py. This interface constructs the model via the factory pattern and forwards arguments to the Python API:

python -m rfdetr export \
    --model rfdetr-small \
    --format tensorrt \
    --fp16 \
    --output-dir ./exported \
    --engine-name rfdetr_small

Supported model names include rfdetr-nano, rfdetr-small, rfdetr-medium, and others. Omit the --fp16 flag to generate an FP32 engine instead.

Validating the TensorRT Engine

After export, verify engine functionality using the benchmark utility in src/rfdetr/export/benchmark.py:

from rfdetr.export.benchmark import run_tensorrt_inference

results = run_tensorrt_inference(
    engine_path="exported/rfdetr_small-fp16.engine",
    image_path="sample.jpg"
)

This loads the engine using polygraphy.trt.Engine.from_file, executes forward passes, and reports latency statistics. The implementation mirrors the validation logic in tests/export/test_tensorrt_export.py, which ensures FP16 and FP32 export paths function correctly.

Troubleshooting Common Issues

FP16 Builder Limitations

Some minimal TensorRT installations lack FP16 support. The builder in src/rfdetr/export/_tensorrt.py (lines 126-134) automatically detects this condition, falls back to FP32, and emits a warning rather than failing.

ScatterND Compatibility

The transformer export path deliberately avoids ScatterND operations because TensorRT cannot consume them. As noted in src/rfdetr/models/transformer.py (lines 342-347), the implementation uses alternative graph constructions to ensure compatibility.

GPU Availability Requirements

TensorRT export requires an NVIDIA GPU with CUDA drivers installed. The code checks for CUDA availability before invoking the builder and raises a clear error if running on CPU-only systems.

Summary

Frequently Asked Questions

What is the difference between ONNX and TensorRT export in RF-DETR?

ONNX export produces a portable graph representation suitable for cross-platform deployment, while TensorRT export generates a highly optimized, NVIDIA-specific serialized engine. The TensorRT pipeline in src/rfdetr/export/main.py actually uses ONNX as an intermediate step, then immediately builds the TensorRT engine from that ONNX file using the builder in src/rfdetr/export/_tensorrt.py.

Can I export RF-DETR to TensorRT on a machine without a GPU?

No. The TensorRT builder requires a physical NVIDIA GPU to optimize and serialize the engine, even if you later intend to run inference on a different device. The export code explicitly checks for CUDA availability and will raise a runtime error if no GPU is detected.

How do I choose between FP16 and FP32 precision for TensorRT export?

Use FP16 (the default) for maximum inference speed and reduced memory bandwidth, which is ideal for most modern NVIDIA GPUs (Turing architecture and newer). Select FP32 only if you encounter numerical precision issues with FP16 or if deploying to older hardware with limited FP16 support. The exporter automatically falls back to FP32 if FP16 builder flags are unavailable.

Where does RF-DETR save the TensorRT engine file?

The engine is saved to the directory specified by the output_dir parameter, using the naming convention defined in src/rfdetr/export/_naming.py. The filename combines your specified engine_name with the precision suffix, producing files like rfdetr_small-fp16.engine or rfdetr_small-fp32.engine in the designated output directory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →