How to Export RF-DETR Models to OpenVINO Format: A Complete Developer Guide

RF-DETR provides a native export path that converts PyTorch models directly to OpenVINO IR format (.xml and .bin) using a single model.export(format="openvino") call without requiring intermediate ONNX conversion.

The RF-DETR repository by Roboflow includes a built-in export pipeline optimized for Intel hardware deployment. Converting your trained detection transformer to OpenVINO Intermediate Representation (IR) enables high-performance inference on CPUs, GPUs, and VPUs while eliminating PyTorch runtime dependencies. This guide walks through the export process using the actual source implementation in the roboflow/rf-detr codebase.

Prerequisites for OpenVINO Export

Before exporting, install the optional OpenVINO dependencies bundled with RF-DETR:

pip install "rfdetr[openvino]"

Ensure your model is loaded on the CPU device, as the exporter automatically handles device placement during conversion. The input tensor shape you specify during export will be baked into the IR graph, so verify your target resolution before running the conversion.

How the RF-DETR OpenVINO Export Pipeline Works

The export process follows a streamlined backend architecture that bypasses ONNX entirely. When you call model.export(format="openvino"), the system triggers a specific chain of operations defined in src/rfdetr/export/_backend.py and src/rfdetr/export/_openvino/exporter.py.

Backend Dispatch and Entry Points

The journey begins in RFDETR.export(), which detects format="openvino" and forwards the request to _export_openvino_format inside src/rfdetr/export/_backend.py (lines 29-44). This dispatcher performs a critical validation: it raises a NotImplementedError if you attempt to use dynamic_batch=True, because OpenVINO IR captures fixed input shapes (lines 62-66).

The dispatcher then calls _switch_to_export_mode(model), which transforms the model’s forward pass to return plain tensors instead of dictionaries. This normalization is essential because OpenVINO’s converter expects tensor tuples, not complex data structures.

Model Wrapping and Tensor Normalization

Inside src/rfdetr/export/_openvino/exporter.py, the export_openvino function (lines 68-82) wraps your model with ModelWrapper (lines 39-66). This wrapper guarantees that the forward pass yields a tuple of tensors, raising a clear NotImplementedError if the underlying model returns a dictionary.

The wrapper feeds into openvino.convert_model, which traces the computation graph using example input tensors and produces an OpenVINO ov.Model object (lines 61-64).

Precision Handling and Compression

The openvino_precision parameter controls weight compression in the saved IR files:

  • None or "float16" → Sets compress_to_fp16 = True, storing weights in half-precision to reduce file size (default behavior).
  • "float32" → Sets compress_to_fp16 = False, preserving full 32-bit precision for maximum accuracy.

Note that this only affects storage format; actual inference precision depends on the target device capabilities.

File Naming and Output Structure

The exporter writes files to <output_dir>/<stem>.xml and <stem>.bin using openvino.save_model (lines 42-45). The naming convention follows these rules:

  • Default: Uses the model variant (e.g., small.xml / small.bin).
  • Custom name: Set via output_name="my_detector".
  • Backbone-only: Appends -backbone to the stem when backbone_only=True (lines 36-41).

The function returns the absolute path to the .xml file, compatible with the OpenVINOInference runtime wrapper.

Exporting RF-DETR to OpenVINO: Code Examples

Basic Export to OpenVINO IR

Load a pre-trained model and export it to OpenVINO format with default FP16 compression:

from rfdetr import RFDETRSmall

# Load model from Roboflow Hub or local checkpoint

model = RFDETRSmall(pretrain_weights="roboflow/rf-detr/small")

# Export to OpenVINO

model.export(
    format="openvino",
    output_dir="openvino_out",
    verbose=True,
)

# Output: openvino_out/small.xml and openvino_out/small.bin

Exporting Backbone-Only Models with Custom Names

To export only the feature extractor with full precision weights:

model.export(
    format="openvino",
    output_dir="openvino_out",
    output_name="my_detector",
    backbone_only=True,
    openvino_precision="float32",
)

# Output: openvino_out/my_detector-backbone.xml and .bin

Loading and Running OpenVINO Inference

After exporting, load the IR files using the built-in inference wrapper defined in src/rfdetr/export/_openvino/inference.py:

from rfdetr.export._openvino.inference import OpenVINOInference

# Initialize runtime with exported IR

ov_inferencer = OpenVINOInference(
    model_path="openvino_out/small.xml",
    device="CPU",  # Supports "GPU", "MYRIAD", etc.

)

# Run inference (input must match export shape)

outputs = ov_inferencer.run(input_tensor)

# Returns: tuple of tensors (detections, labels, scores)

The OpenVINOInference class handles device compilation and tensor formatting, returning results compatible with RF-DETR’s post-processing pipeline.

Summary

  • Direct conversion: RF-DETR exports directly to OpenVINO IR without ONNX intermediates via model.export(format="openvino").
  • Fixed shapes: Dynamic batching is not supported; the IR captures static input dimensions.
  • Precision control: Use openvino_precision="float32" for full precision or "float16"/None for compressed weights.
  • Modular export: Support for backbone_only=True exports just the feature extractor for custom heads.
  • Runtime ready: Exported models load via OpenVINOInference for immediate deployment on Intel hardware.

Frequently Asked Questions

Does RF-DETR require ONNX conversion before OpenVINO?

No. According to the roboflow/rf-detr source code, the export pipeline in src/rfdetr/export/_openvino/exporter.py calls openvino.convert_model directly on the PyTorch model wrapped by ModelWrapper. This eliminates the ONNX conversion step typically required by other frameworks, reducing conversion time and potential compatibility issues.

Can I use dynamic batch sizes with RF-DETR OpenVINO exports?

No. The backend explicitly raises a NotImplementedError in src/rfdetr/export/_backend.py (lines 62-66) when dynamic_batch=True is specified. OpenVINO IR files encode fixed input tensor shapes during the tracing process. If you need variable batch sizes, export separate models for each batch size or use OpenVINO’s dynamic shape APIs after loading the static IR.

What precision formats does RF-DETR OpenVINO export support?

The exporter supports two storage precisions controlled by the openvino_precision argument: "float32" for uncompressed 32-bit weights and "float16" (or None) for FP16 compression. This setting only affects the .bin weight file; actual inference precision is determined by the target device’s capabilities and the OpenVINOInference runtime configuration.

How do I load and run inference on the exported OpenVINO IR files?

Use the OpenVINOInference class from src/rfdetr/export/_openvino/inference.py. Instantiate it with the path to your .xml file and target device (e.g., "CPU" or "GPU"), then call .run(input_tensor) with an input matching the fixed shape defined during export. The class returns a tuple of tensors containing detection coordinates, class labels, and confidence scores.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →