# How to Export RF-DETR Models to TFLite Format: A Complete Guide

> Easily export RF-DETR models to TFLite format using the export() method. Optimize edge deployment with INT8 or FP16 quantization. Get the complete guide now.

- Repository: [Roboflow/rf-detr](https://github.com/roboflow/rf-detr)
- Tags: how-to-guide
- Published: 2026-09-08

---

**You can export RF-DETR models to TensorFlow Lite format by calling the `export()` method with `format="tflite"` on any trained model instance, optionally enabling INT8 or FP16 quantization for optimized edge deployment.**

RF-DETR is an open-source real-time detection transformer from Roboflow. When deploying these models on mobile devices or embedded systems with limited compute, converting them to TFLite format significantly reduces model size and improves inference latency. This guide walks through the complete workflow to export RF-DETR model to TFLite format using the library's built-in export pipeline.

## Understanding the TFLite Export Architecture

The export process follows a two-stage pipeline defined in the source code. First, the PyTorch model converts to ONNX format as an intermediate representation. Then, the pipeline transforms the ONNX graph into a TensorFlow SavedModel before generating the final `.tflite` flatbuffer file.

According to [`src/rfdetr/detr.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/detr.py) (lines 1620-1625), the `RFDETR.export()` method dispatches to format-specific converters. When you specify `format="tflite"`, the pipeline executes the conversion logic implemented in [`src/rfdetr/export/main.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/main.py).

### Step 1 - ONNX Intermediate Generation

As implemented in [`src/rfdetr/export/main.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/main.py) (lines 51-67), the exporter first invokes `torch.export` or `torch.onnx.export` to generate a temporary ONNX file. This intermediate representation preserves the model's computation graph and weights in a framework-agnostic format that bridges the PyTorch and TensorFlow ecosystems.

### Step 2 - ONNX to TFLite Conversion

The [`src/rfdetr/export/_tflite/converter.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/_tflite/converter.py) module handles the core transformation logic. It utilizes the `onnx2tf` library (version >= 2.4.0) to convert the ONNX graph to TensorFlow format, then applies `tf.lite.TFLiteConverter` to produce the final optimized flatbuffer. The final file is written adjacent to the ONNX artifact, with the stem derived from the `output_name` parameter.

## Handling RF-DETR Specific Operations

RF-DETR utilizes a `GridSample` operation that lacks native TFLite support. The converter addresses this limitation through sophisticated graph rewriting that maintains mathematical equivalence.

### GridSample Patch Implementation

In [`src/rfdetr/export/_tflite/converter.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/_tflite/converter.py) (lines 100-154), the `patch_grid_sample` helper traverses the ONNX graph and rewrites each `GridSample` node into a **TFLite-safe gather-based sub-graph**. This transformation ensures compatibility with the TFLite runtime while preserving the deformable attention mechanisms that make RF-DETR accurate.

### Operator Replacement Strategy

The converter also replaces unsupported activation functions like `Erf` and `GeLU` with pseudo-operators that TFLite implements natively (see lines 900-1006 of [`converter.py`](https://github.com/roboflow/rf-detr/blob/main/converter.py)). This prevents "FlexErf" runtime errors that commonly occur when deploying transformer architectures on edge devices.

## How to Export RF-DETR to TFLite

Before exporting, ensure you have installed `onnx2tf>=2.4.0` and either `tensorflow` or `tflite-runtime`. The pipeline logs a dependency warning if these are missing, as the TFLite export path is currently marked experimental.

### Basic Export Example

```python
from rfdetr import RFDETRSmall

# Load a pretrained model

model = RFDETRSmall(pretrain_weights="rf-detr-small.pth")

# Export to TFLite

tflite_path = model.export(
    format="tflite",           # Required: selects TFLite backend

    output_name="my_rfdet",    # Base filename for outputs

    quantization="fp32",       # Options: "fp32", "int8", "float16"

    verbose=True               # Prints conversion pipeline steps

)

print(f"TFLite model saved to: {tflite_path}")

```

### Export Method Parameters

- **format**: Must be `"tflite"` to trigger the TensorFlow Lite conversion path
- **output_name**: Stem for the intermediate ONNX file (`my_rfdet.onnx`) and final TFLite file (`my_rfdet_fp32.tflite` or `my_rfdet_int8.tflite`)
- **quantization**: `"fp32"` (default), `"int8"` for dynamic-range/full INT8 quantization, or `"float16"` for half-precision
- **onnx_path**: Optional path to an existing ONNX file; if `None`, a temporary ONNX file is generated automatically
- **verbose**: When `True`, emits detailed logs from the conversion pipeline for debugging

### Quantization Implementation

For INT8 quantization, the converter builds a representative dataset and performs calibration as implemented in [`src/rfdetr/export/_tflite/converter.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/_tflite/converter.py) (lines 820-850). This produces either dynamic-range quantized models or full INT8 models depending on the specific TFLite converter configuration generated by `onnx2tf`.

## Running Inference with TFLite Models

RF-DETR provides lightweight runtime wrappers in [`src/rfdetr/export/_tflite/inference.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/_tflite/inference.py) that handle TFLite interpreter allocation and input preprocessing.

```python
from rfdetr.export._tflite.inference import _create_interpreter, _run_tflite
import numpy as np

# Load the model once

interpreter = _create_interpreter("my_rfdet_int8.tflite")

# Prepare a single image in HWC uint8 format (640x640x3)

image = np.random.randint(0, 255, (640, 640, 3), dtype=np.uint8)

# Run inference

detections = _run_tflite(interpreter, image)

# Returns a supervision.Detections object compatible with the RF-DETR ecosystem

print(detections)

```

The `_run_tflite` function (lines 30-80 of [`inference.py`](https://github.com/roboflow/rf-detr/blob/main/inference.py)) automatically handles **channel-first to channel-last conversion** and input normalization, returning a standard `supervision.Detections` object compatible with the rest of the RF-DETR evaluation pipeline.

## Summary

- **Two-stage pipeline**: RF-DETR exports to ONNX first via `torch.onnx.export`, then converts to TFLite using `onnx2tf` and `tf.lite.TFLiteConverter` as orchestrated in [`src/rfdetr/export/main.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/main.py)
- **Automatic graph patching**: The `patch_grid_sample` helper in [`src/rfdetr/export/_tflite/converter.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/_tflite/converter.py) rewrites unsupported operations into TFLite-compatible equivalents
- **Quantization support**: Optional INT8 and FP16 quantization reduces model size for resource-constrained edge devices
- **Inference helpers**: Use `_create_interpreter` and `_run_tflite` from [`inference.py`](https://github.com/roboflow/rf-detr/blob/main/inference.py) to handle runtime execution without manual tensor allocation
- **Source locations**: Core export logic resides in [`src/rfdetr/detr.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/detr.py) and [`src/rfdetr/export/_tflite/converter.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/_tflite/converter.py)

## Frequently Asked Questions

### What dependencies are required for TFLite export?

You must install `onnx2tf>=2.4.0` and either `tensorflow` or `tflite-runtime`. The pipeline checks for these at runtime and emits a warning if they're missing, as these are optional dependencies in the RF-DETR package. The `onnx2tf` library handles the ONNX-to-TensorFlow graph translation, while TensorFlow provides the TFLite converter runtime.

### Why is the TFLite export marked as experimental?

The experimental status reflects the complexity of translating transformer architectures to TFLite's constrained operator set. While the `GridSample` patching and custom operator replacements (like `Erf` to native TFLite ops) are automated, edge cases in model architectures may require manual graph adjustments. The Roboflow team continues to stabilize this pathway based on community feedback.

### Can I use a custom ONNX file instead of generating one?

Yes. Pass the `onnx_path` parameter to `model.export()` with the path to your existing ONNX file. When provided, the pipeline skips the PyTorch-to-ONNX conversion step and proceeds directly to TFLite conversion using your specified model graph. This is useful if you've already optimized the ONNX model with external tools.

### Does INT8 quantization significantly impact detection accuracy?

INT8 quantization typically introduces minor accuracy degradation, but RF-DETR's implementation uses representative data calibration to minimize loss. The dynamic-range quantization available in [`src/rfdetr/export/_tflite/converter.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/export/_tflite/converter.py) (lines 820-850) automatically determines optimal scales during conversion. For critical applications, validate FP32 and INT8 outputs against your test set before production deployment.