How to Export an RF-DETR Model to TensorRT Format
RF-DETR provides a built-in export pipeline that converts trained detection models into optimized TensorRT engines via an intermediate ONNX representation, accessible through both Python API and CLI interfaces.
Exporting RF-DETR to TensorRT enables high-performance inference on NVIDIA GPUs. According to the roboflow/rf-detr source code, the export functionality is implemented in the core RFDETR class and orchestrated through dedicated export modules, eliminating the need for manual conversion scripts.
Installing TensorRT Dependencies
TensorRT support is optional to keep the base installation lightweight. Install the required dependencies using the tensorrt extra, which pulls in the polygraphy wrapper library:
pip install "rfdetr[tensorrt]"
This extra ensures the TensorRT Python API is available through polygraphy, allowing in-process engine building without launching external trtexec subprocesses.
Exporting via Python API
The RFDETR.export() method in src/rfdetr/detr.py drives the entire conversion pipeline. This method handles ONNX generation and TensorRT engine serialization in a single call.
Loading the Model
Instantiate a pre-trained model or load a custom checkpoint before export:
import rfdetr
# Load a pre-trained small variant
model = rfdetr.detr.RFDETRSmall()
# Optional: configure input resolution (must be set before export)
model.resolution = (640, 640)
model.eval() # Ensure deterministic export mode
Configuring Export Parameters
The export() method accepts specific parameters for TensorRT optimization:
format: Must be set to"tensorrt"to trigger the TensorRT builderfp16: Boolean flag for FP16 precision (defaultTrue). Set toFalsefor FP32 engines when targeting TensorRT builds without FP16 supportoutput_dir: Directory path for saving the ONNX intermediate and final.enginefileengine_name: Base filename for the serialized engine (e.g.,"rfdetr_small"generatesrfdetr_small-fp16.engine)backbone_only: Export only the backbone architecture (useful for downstream fine-tuning pipelines)
Executing the Export
Run the export with your desired configuration:
engine_path = model.export(
format="tensorrt",
fp16=True, # Set False for FP32 precision
output_dir="./exported", # Creates directory if missing
engine_name="rfdetr_small"
)
print(f"TensorRT engine saved to: {engine_path}")
Under the hood, RFDETR.export() in src/rfdetr/detr.py performs three operations:
- ONNX Generation – Calls
torch.onnx.exportto create an ONNX graph, implemented insrc/rfdetr/export/main.py - Engine Building – Passes the ONNX file to
src/rfdetr/export/_tensorrt.py, which usespolygraphy'sbuild_engineAPI to create the serialized TensorRT engine - Filename Resolution – Uses utilities in
src/rfdetr/export/_naming.pyto generate consistent filenames (e.g.,<stem>-fp16.engineor<stem>-fp32.engine)
Exporting via Command Line
The same functionality is exposed through the CLI entry point in src/rfdetr/__main__.py. This interface constructs the model via the factory pattern and forwards arguments to the Python API:
python -m rfdetr export \
--model rfdetr-small \
--format tensorrt \
--fp16 \
--output-dir ./exported \
--engine-name rfdetr_small
Supported model names include rfdetr-nano, rfdetr-small, rfdetr-medium, and others. Omit the --fp16 flag to generate an FP32 engine instead.
Validating the TensorRT Engine
After export, verify engine functionality using the benchmark utility in src/rfdetr/export/benchmark.py:
from rfdetr.export.benchmark import run_tensorrt_inference
results = run_tensorrt_inference(
engine_path="exported/rfdetr_small-fp16.engine",
image_path="sample.jpg"
)
This loads the engine using polygraphy.trt.Engine.from_file, executes forward passes, and reports latency statistics. The implementation mirrors the validation logic in tests/export/test_tensorrt_export.py, which ensures FP16 and FP32 export paths function correctly.
Troubleshooting Common Issues
FP16 Builder Limitations
Some minimal TensorRT installations lack FP16 support. The builder in src/rfdetr/export/_tensorrt.py (lines 126-134) automatically detects this condition, falls back to FP32, and emits a warning rather than failing.
ScatterND Compatibility
The transformer export path deliberately avoids ScatterND operations because TensorRT cannot consume them. As noted in src/rfdetr/models/transformer.py (lines 342-347), the implementation uses alternative graph constructions to ensure compatibility.
GPU Availability Requirements
TensorRT export requires an NVIDIA GPU with CUDA drivers installed. The code checks for CUDA availability before invoking the builder and raises a clear error if running on CPU-only systems.
Summary
- Install TensorRT support via
pip install "rfdetr[tensorrt]" - Use
model.export(format="tensorrt", ...)fromsrc/rfdetr/detr.pyfor Python-based exports - Leverage
src/rfdetr/export/_tensorrt.pyfor low-level engine building viapolygraphy - Automate exports using the CLI interface in
src/rfdetr/__main__.py - Validate engines with the benchmark utilities in
src/rfdetr/export/benchmark.py - Handle FP16 limitations gracefully with automatic fallback to FP32 precision
Frequently Asked Questions
What is the difference between ONNX and TensorRT export in RF-DETR?
ONNX export produces a portable graph representation suitable for cross-platform deployment, while TensorRT export generates a highly optimized, NVIDIA-specific serialized engine. The TensorRT pipeline in src/rfdetr/export/main.py actually uses ONNX as an intermediate step, then immediately builds the TensorRT engine from that ONNX file using the builder in src/rfdetr/export/_tensorrt.py.
Can I export RF-DETR to TensorRT on a machine without a GPU?
No. The TensorRT builder requires a physical NVIDIA GPU to optimize and serialize the engine, even if you later intend to run inference on a different device. The export code explicitly checks for CUDA availability and will raise a runtime error if no GPU is detected.
How do I choose between FP16 and FP32 precision for TensorRT export?
Use FP16 (the default) for maximum inference speed and reduced memory bandwidth, which is ideal for most modern NVIDIA GPUs (Turing architecture and newer). Select FP32 only if you encounter numerical precision issues with FP16 or if deploying to older hardware with limited FP16 support. The exporter automatically falls back to FP32 if FP16 builder flags are unavailable.
Where does RF-DETR save the TensorRT engine file?
The engine is saved to the directory specified by the output_dir parameter, using the naming convention defined in src/rfdetr/export/_naming.py. The filename combines your specified engine_name with the precision suffix, producing files like rfdetr_small-fp16.engine or rfdetr_small-fp32.engine in the designated output directory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →