How to Export RF-DETR Models to TFLite Format: A Complete Guide
You can export RF-DETR models to TensorFlow Lite format by calling the export() method with format="tflite" on any trained model instance, optionally enabling INT8 or FP16 quantization for optimized edge deployment.
RF-DETR is an open-source real-time detection transformer from Roboflow. When deploying these models on mobile devices or embedded systems with limited compute, converting them to TFLite format significantly reduces model size and improves inference latency. This guide walks through the complete workflow to export RF-DETR model to TFLite format using the library's built-in export pipeline.
Understanding the TFLite Export Architecture
The export process follows a two-stage pipeline defined in the source code. First, the PyTorch model converts to ONNX format as an intermediate representation. Then, the pipeline transforms the ONNX graph into a TensorFlow SavedModel before generating the final .tflite flatbuffer file.
According to src/rfdetr/detr.py (lines 1620-1625), the RFDETR.export() method dispatches to format-specific converters. When you specify format="tflite", the pipeline executes the conversion logic implemented in src/rfdetr/export/main.py.
Step 1 - ONNX Intermediate Generation
As implemented in src/rfdetr/export/main.py (lines 51-67), the exporter first invokes torch.export or torch.onnx.export to generate a temporary ONNX file. This intermediate representation preserves the model's computation graph and weights in a framework-agnostic format that bridges the PyTorch and TensorFlow ecosystems.
Step 2 - ONNX to TFLite Conversion
The src/rfdetr/export/_tflite/converter.py module handles the core transformation logic. It utilizes the onnx2tf library (version >= 2.4.0) to convert the ONNX graph to TensorFlow format, then applies tf.lite.TFLiteConverter to produce the final optimized flatbuffer. The final file is written adjacent to the ONNX artifact, with the stem derived from the output_name parameter.
Handling RF-DETR Specific Operations
RF-DETR utilizes a GridSample operation that lacks native TFLite support. The converter addresses this limitation through sophisticated graph rewriting that maintains mathematical equivalence.
GridSample Patch Implementation
In src/rfdetr/export/_tflite/converter.py (lines 100-154), the patch_grid_sample helper traverses the ONNX graph and rewrites each GridSample node into a TFLite-safe gather-based sub-graph. This transformation ensures compatibility with the TFLite runtime while preserving the deformable attention mechanisms that make RF-DETR accurate.
Operator Replacement Strategy
The converter also replaces unsupported activation functions like Erf and GeLU with pseudo-operators that TFLite implements natively (see lines 900-1006 of converter.py). This prevents "FlexErf" runtime errors that commonly occur when deploying transformer architectures on edge devices.
How to Export RF-DETR to TFLite
Before exporting, ensure you have installed onnx2tf>=2.4.0 and either tensorflow or tflite-runtime. The pipeline logs a dependency warning if these are missing, as the TFLite export path is currently marked experimental.
Basic Export Example
from rfdetr import RFDETRSmall
# Load a pretrained model
model = RFDETRSmall(pretrain_weights="rf-detr-small.pth")
# Export to TFLite
tflite_path = model.export(
format="tflite", # Required: selects TFLite backend
output_name="my_rfdet", # Base filename for outputs
quantization="fp32", # Options: "fp32", "int8", "float16"
verbose=True # Prints conversion pipeline steps
)
print(f"TFLite model saved to: {tflite_path}")
Export Method Parameters
- format: Must be
"tflite"to trigger the TensorFlow Lite conversion path - output_name: Stem for the intermediate ONNX file (
my_rfdet.onnx) and final TFLite file (my_rfdet_fp32.tfliteormy_rfdet_int8.tflite) - quantization:
"fp32"(default),"int8"for dynamic-range/full INT8 quantization, or"float16"for half-precision - onnx_path: Optional path to an existing ONNX file; if
None, a temporary ONNX file is generated automatically - verbose: When
True, emits detailed logs from the conversion pipeline for debugging
Quantization Implementation
For INT8 quantization, the converter builds a representative dataset and performs calibration as implemented in src/rfdetr/export/_tflite/converter.py (lines 820-850). This produces either dynamic-range quantized models or full INT8 models depending on the specific TFLite converter configuration generated by onnx2tf.
Running Inference with TFLite Models
RF-DETR provides lightweight runtime wrappers in src/rfdetr/export/_tflite/inference.py that handle TFLite interpreter allocation and input preprocessing.
from rfdetr.export._tflite.inference import _create_interpreter, _run_tflite
import numpy as np
# Load the model once
interpreter = _create_interpreter("my_rfdet_int8.tflite")
# Prepare a single image in HWC uint8 format (640x640x3)
image = np.random.randint(0, 255, (640, 640, 3), dtype=np.uint8)
# Run inference
detections = _run_tflite(interpreter, image)
# Returns a supervision.Detections object compatible with the RF-DETR ecosystem
print(detections)
The _run_tflite function (lines 30-80 of inference.py) automatically handles channel-first to channel-last conversion and input normalization, returning a standard supervision.Detections object compatible with the rest of the RF-DETR evaluation pipeline.
Summary
- Two-stage pipeline: RF-DETR exports to ONNX first via
torch.onnx.export, then converts to TFLite usingonnx2tfandtf.lite.TFLiteConverteras orchestrated insrc/rfdetr/export/main.py - Automatic graph patching: The
patch_grid_samplehelper insrc/rfdetr/export/_tflite/converter.pyrewrites unsupported operations into TFLite-compatible equivalents - Quantization support: Optional INT8 and FP16 quantization reduces model size for resource-constrained edge devices
- Inference helpers: Use
_create_interpreterand_run_tflitefrominference.pyto handle runtime execution without manual tensor allocation - Source locations: Core export logic resides in
src/rfdetr/detr.pyandsrc/rfdetr/export/_tflite/converter.py
Frequently Asked Questions
What dependencies are required for TFLite export?
You must install onnx2tf>=2.4.0 and either tensorflow or tflite-runtime. The pipeline checks for these at runtime and emits a warning if they're missing, as these are optional dependencies in the RF-DETR package. The onnx2tf library handles the ONNX-to-TensorFlow graph translation, while TensorFlow provides the TFLite converter runtime.
Why is the TFLite export marked as experimental?
The experimental status reflects the complexity of translating transformer architectures to TFLite's constrained operator set. While the GridSample patching and custom operator replacements (like Erf to native TFLite ops) are automated, edge cases in model architectures may require manual graph adjustments. The Roboflow team continues to stabilize this pathway based on community feedback.
Can I use a custom ONNX file instead of generating one?
Yes. Pass the onnx_path parameter to model.export() with the path to your existing ONNX file. When provided, the pipeline skips the PyTorch-to-ONNX conversion step and proceeds directly to TFLite conversion using your specified model graph. This is useful if you've already optimized the ONNX model with external tools.
Does INT8 quantization significantly impact detection accuracy?
INT8 quantization typically introduces minor accuracy degradation, but RF-DETR's implementation uses representative data calibration to minimize loss. The dynamic-range quantization available in src/rfdetr/export/_tflite/converter.py (lines 820-850) automatically determines optimal scales during conversion. For critical applications, validate FP32 and INT8 outputs against your test set before production deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →