How to Use RFDETR for Object Detection: A Complete Guide with Code Examples

RFDETR provides a unified Python API for transformer-based object detection through size-specific model classes that accept images via file paths, URLs, or arrays and return structured supervision.Detections objects.

RFDETR is a state-of-the-art transformer-based detection model built on the DINOv2 vision transformer backbone. Developed by Roboflow and hosted in the roboflow/rf-detr repository, it consolidates complex vision transformer logic into simple, importable classes. This guide demonstrates how to use RFDETR for object detection using the exact implementation details found in the source code.

Core Architecture and Model Variants

The RFDETR ecosystem centers on a base implementation with multiple size-optimized variants that share an identical public API.

The Base RFDETR Class

All detection functionality inherits from the generic RFDETR class defined in [src/rfdetr/detr.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/detr.py#L458) at line 458. This base class encapsulates the complete torch.nn.Module containing the DINOv2 backbone, transformer encoder/decoder layers, and task-specific prediction heads. When you instantiate any variant, you are implicitly initializing this base architecture with specific configuration parameters.

Size-Specific Variants

Practical deployment flexibility comes from thin subclasses located in [src/rfdetr/variants.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/variants.py). The available detection variants include:

  • RFDETRNano – Fastest inference, lowest computational cost
  • RFDETRSmall – Balanced speed and accuracy
  • RFDETRMedium – Recommended for general purpose use
  • RFDETRLarge – Higher accuracy, increased latency
  • RFDETR-XL and 2XL – Maximum accuracy (Plus subscription only)

Each variant modifies only the backbone resolution and transformer layer count, leaving the high-level interface unchanged.

Package Entry Points

The top-level rfdetr package consolidates these exports in [src/rfdetr/__init__.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/__init__.py) between lines 44-61. This enables direct imports such as from rfdetr import RFDETRMedium without referencing internal file structures. The initialization module also implements lazy loading for training-specific symbols to minimize import overhead.

The Inference Pipeline

The RFDETR.predict() method provides the primary interface for object detection, implemented in [src/rfdetr/detr.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/detr.py#L514-L560) from lines 514 to 560.

Input Handling and Preprocessing

The predict() method accepts multiple input types: local file paths, remote URLs, PIL Image objects, or NumPy arrays. Internally, it constructs a ModelContext (exposed through rfdetr.inference) that handles deterministic resizing to the model-specific resolution defined in each variant's configuration. Preprocessing utilities in rfdetr.utilities.image manage format conversion and normalization automatically.

Output Format and Post-Processing

The method executes a forward pass through the underlying torch.nn.Module, applies post-processing including Non-Maximum Suppression (NMS) and score thresholding, and returns a supervision.Detections object. This container includes bounding box coordinates, confidence scores, class IDs, and metadata such as the source image reference.

Practical Code Examples

The following implementations demonstrate how to use RFDETR for object detection in real-world scenarios.

Basic Detection with RFDETRMedium

This example uses the medium variant with COCO-pretrained weights to detect objects in a remote image:

import supervision as sv
from rfdetr import RFDETRMedium
from rfdetr.assets.coco_classes import COCO_CLASSES

# Instantiate model – downloads weights automatically on first use

model = RFDETRMedium()

# Run inference with 50% confidence threshold

detections = model.predict(
    "https://media.roboflow.com/dog.jpg",
    threshold=0.5,
)

# Generate human-readable labels from COCO class IDs

labels = [f"{COCO_CLASSES[cid]}" for cid in detections.class_id]

# Visualize results using supervision annotators

annotated = sv.BoxAnnotator().annotate(detections.metadata["source_image"], detections)
annotated = sv.LabelAnnotator().annotate(annotated, detections, labels)

# Save output

annotated.save("dog_pred.jpg")

Switching Model Sizes

To optimize for speed or accuracy, change only the import statement and class instantiation:

from rfdetr import RFDETRNano  # Fastest variant

model = RFDETRNano()
detections = model.predict("image.jpg", threshold=0.5)

Alternative: Using the Inference Package

For environments requiring model resolution by string alias, use the inference package entry point:

from inference import get_model
import requests
from PIL import Image
import supervision as sv

# Resolve model by alias (matches README table)

model = get_model("rfdetr-medium")

# Load image

image = Image.open(requests.get("https://media.roboflow.com/dog.jpg", stream=True).raw)

# Run inference and convert to Detections

pred = model.infer(image, confidence=0.5)[0]
detections = sv.Detections.from_inference(pred)

# Visualize

annotated = sv.BoxAnnotator().annotate(image, detections)

Key Source Files

Understanding the repository structure helps with advanced customization:

Summary

  • RFDETR provides transformer-based detection through a unified API built on the DINOv2 backbone, with classes defined in src/rfdetr/detr.py.
  • Model variants (Nano, Small, Medium, Large, XL) in src/rfdetr/variants.py offer different speed-accuracy trade-offs while maintaining identical method signatures.
  • The predict() method accepts diverse input formats (paths, URLs, PIL, NumPy) and returns structured supervision.Detections objects containing bounding boxes, scores, and class IDs.
  • Automatic weight downloading occurs on first instantiation, and the package supports both direct Python imports and alternative resolution via the inference package.

Frequently Asked Questions

What is the difference between RFDETR model variants?

Each variant configures the same base architecture with different backbone resolutions and transformer layer counts. RFDETRNano prioritizes inference speed with reduced layers, while RFDETRMedium and RFDETRLarge increase model capacity for higher accuracy at the cost of computational requirements. All variants maintain the same predict() interface defined in src/rfdetr/detr.py.

How does RFDETR handle image preprocessing?

The predict() method automatically manages preprocessing through the internal ModelContext class. It performs deterministic resizing to the variant-specific input resolution, normalization, and tensor conversion before passing data to the transformer backbone. These steps occur transparently within the rfdetr.inference module.

Can I use RFDETR with custom trained weights?

Yes. While the examples demonstrate COCO-pretrained weights downloaded automatically on first use, the RFDETR base class architecture supports loading custom checkpoints. You would instantiate your chosen variant class and load state dictionaries into the underlying torch.nn.Module before running inference.

What output format does RFDETR.predict() return?

The method returns a supervision.Detections object containing xyxy bounding boxes, confidence scores, class IDs, and metadata including the original source image. This format integrates directly with the Supervision library's visualization tools, as demonstrated in the annotation examples using sv.BoxAnnotator().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →