How to Use RFDETR for Object Detection: A Complete Guide with Code Examples
RFDETR provides a unified Python API for transformer-based object detection through size-specific model classes that accept images via file paths, URLs, or arrays and return structured supervision.Detections objects.
RFDETR is a state-of-the-art transformer-based detection model built on the DINOv2 vision transformer backbone. Developed by Roboflow and hosted in the roboflow/rf-detr repository, it consolidates complex vision transformer logic into simple, importable classes. This guide demonstrates how to use RFDETR for object detection using the exact implementation details found in the source code.
Core Architecture and Model Variants
The RFDETR ecosystem centers on a base implementation with multiple size-optimized variants that share an identical public API.
The Base RFDETR Class
All detection functionality inherits from the generic RFDETR class defined in [src/rfdetr/detr.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/detr.py#L458) at line 458. This base class encapsulates the complete torch.nn.Module containing the DINOv2 backbone, transformer encoder/decoder layers, and task-specific prediction heads. When you instantiate any variant, you are implicitly initializing this base architecture with specific configuration parameters.
Size-Specific Variants
Practical deployment flexibility comes from thin subclasses located in [src/rfdetr/variants.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/variants.py). The available detection variants include:
- RFDETRNano – Fastest inference, lowest computational cost
- RFDETRSmall – Balanced speed and accuracy
- RFDETRMedium – Recommended for general purpose use
- RFDETRLarge – Higher accuracy, increased latency
- RFDETR-XL and 2XL – Maximum accuracy (Plus subscription only)
Each variant modifies only the backbone resolution and transformer layer count, leaving the high-level interface unchanged.
Package Entry Points
The top-level rfdetr package consolidates these exports in [src/rfdetr/__init__.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/__init__.py) between lines 44-61. This enables direct imports such as from rfdetr import RFDETRMedium without referencing internal file structures. The initialization module also implements lazy loading for training-specific symbols to minimize import overhead.
The Inference Pipeline
The RFDETR.predict() method provides the primary interface for object detection, implemented in [src/rfdetr/detr.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/detr.py#L514-L560) from lines 514 to 560.
Input Handling and Preprocessing
The predict() method accepts multiple input types: local file paths, remote URLs, PIL Image objects, or NumPy arrays. Internally, it constructs a ModelContext (exposed through rfdetr.inference) that handles deterministic resizing to the model-specific resolution defined in each variant's configuration. Preprocessing utilities in rfdetr.utilities.image manage format conversion and normalization automatically.
Output Format and Post-Processing
The method executes a forward pass through the underlying torch.nn.Module, applies post-processing including Non-Maximum Suppression (NMS) and score thresholding, and returns a supervision.Detections object. This container includes bounding box coordinates, confidence scores, class IDs, and metadata such as the source image reference.
Practical Code Examples
The following implementations demonstrate how to use RFDETR for object detection in real-world scenarios.
Basic Detection with RFDETRMedium
This example uses the medium variant with COCO-pretrained weights to detect objects in a remote image:
import supervision as sv
from rfdetr import RFDETRMedium
from rfdetr.assets.coco_classes import COCO_CLASSES
# Instantiate model – downloads weights automatically on first use
model = RFDETRMedium()
# Run inference with 50% confidence threshold
detections = model.predict(
"https://media.roboflow.com/dog.jpg",
threshold=0.5,
)
# Generate human-readable labels from COCO class IDs
labels = [f"{COCO_CLASSES[cid]}" for cid in detections.class_id]
# Visualize results using supervision annotators
annotated = sv.BoxAnnotator().annotate(detections.metadata["source_image"], detections)
annotated = sv.LabelAnnotator().annotate(annotated, detections, labels)
# Save output
annotated.save("dog_pred.jpg")
Switching Model Sizes
To optimize for speed or accuracy, change only the import statement and class instantiation:
from rfdetr import RFDETRNano # Fastest variant
model = RFDETRNano()
detections = model.predict("image.jpg", threshold=0.5)
Alternative: Using the Inference Package
For environments requiring model resolution by string alias, use the inference package entry point:
from inference import get_model
import requests
from PIL import Image
import supervision as sv
# Resolve model by alias (matches README table)
model = get_model("rfdetr-medium")
# Load image
image = Image.open(requests.get("https://media.roboflow.com/dog.jpg", stream=True).raw)
# Run inference and convert to Detections
pred = model.infer(image, confidence=0.5)[0]
detections = sv.Detections.from_inference(pred)
# Visualize
annotated = sv.BoxAnnotator().annotate(image, detections)
Key Source Files
Understanding the repository structure helps with advanced customization:
- Core detection logic: [
src/rfdetr/detr.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/detr.py) – Contains theRFDETRbase class andpredict()implementation - Model definitions: [
src/rfdetr/variants.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/variants.py) – Houses size-specific subclasses - Public API: [
src/rfdetr/__init__.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/__init__.py) – Re-exports classes and manages lazy loading - Class names: [
src/rfdetr/assets/coco_classes.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/assets/coco_classes.py) – Standard COCO category labels for pretrained models - Inference utilities: [
src/rfdetr/inference/__init__.py](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/inference/__init__.py) – InternalModelContexthelpers
Summary
- RFDETR provides transformer-based detection through a unified API built on the DINOv2 backbone, with classes defined in
src/rfdetr/detr.py. - Model variants (Nano, Small, Medium, Large, XL) in
src/rfdetr/variants.pyoffer different speed-accuracy trade-offs while maintaining identical method signatures. - The
predict()method accepts diverse input formats (paths, URLs, PIL, NumPy) and returns structuredsupervision.Detectionsobjects containing bounding boxes, scores, and class IDs. - Automatic weight downloading occurs on first instantiation, and the package supports both direct Python imports and alternative resolution via the
inferencepackage.
Frequently Asked Questions
What is the difference between RFDETR model variants?
Each variant configures the same base architecture with different backbone resolutions and transformer layer counts. RFDETRNano prioritizes inference speed with reduced layers, while RFDETRMedium and RFDETRLarge increase model capacity for higher accuracy at the cost of computational requirements. All variants maintain the same predict() interface defined in src/rfdetr/detr.py.
How does RFDETR handle image preprocessing?
The predict() method automatically manages preprocessing through the internal ModelContext class. It performs deterministic resizing to the variant-specific input resolution, normalization, and tensor conversion before passing data to the transformer backbone. These steps occur transparently within the rfdetr.inference module.
Can I use RFDETR with custom trained weights?
Yes. While the examples demonstrate COCO-pretrained weights downloaded automatically on first use, the RFDETR base class architecture supports loading custom checkpoints. You would instantiate your chosen variant class and load state dictionaries into the underlying torch.nn.Module before running inference.
What output format does RFDETR.predict() return?
The method returns a supervision.Detections object containing xyxy bounding boxes, confidence scores, class IDs, and metadata including the original source image. This format integrates directly with the Supervision library's visualization tools, as demonstrated in the annotation examples using sv.BoxAnnotator().
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →