RF-DETR Model Architecture Parameters: Complete Configuration Guide
The RF-DETR architecture exposes all hyperparameters through a Pydantic-based ModelConfig class defined in src/rfdetr/config.py, allowing granular control over the vision transformer backbone, decoder depth, attention mechanisms, and task-specific heads.
RF-DETR is a real-time transformer-based object detector developed by Roboflow that unifies detection, segmentation, and keypoint estimation behind a single configuration interface. Every architectural decision—from the DINOv2 windowed backbone selection to the number of deformable attention points—is exposed as a typed parameter in the configuration system. Understanding these configurable parameters enables researchers and engineers to scale the model from edge devices (Nano) to high-accuracy cloud deployments (XLarge) or adapt it for custom computer vision tasks.
Core Architectural Parameters
The base ModelConfig class in src/rfdetr/config.py (lines 33-120) defines the superset of all tunable architectural fields. Concrete variants inherit from this class and override values to create the Nano, Small, Medium, Large, and Segmentation profiles.
Backbone and Encoder Settings
The vision transformer backbone behavior is controlled through encoder-specific parameters:
- encoder: String identifier for the vision transformer backbone (e.g.,
dinov2_windowed_small,dinov2_windowed_base). This field is mandatory and defined by each concrete subclass (lines 33-41). - out_feature_indexes: List of layer indices whose feature maps are forwarded to the decoder (e.g.,
[3, 6, 9, 12]for the Small variant). These indices determine which hierarchical features feed the detection head (lines 40-44). - patch_size: ViT patch dimension in pixels, typically
14for Base models or16for smaller variants (lines 47-48). - num_windows: Number of windowed-attention windows in the backbone, ranging from 2-4 depending on model scale (lines 49-50).
- backbone_lora: Boolean flag to enable Low-Rank Adaptation (LoRA) fine-tuning of the backbone weights (default
False). - freeze_encoder: Boolean to freeze encoder parameters during training (default
False). Both flags are defined in lines 117-120.
Decoder and Attention Configuration
The transformer decoder architecture is governed by attention and depth parameters:
- dec_layers: Number of transformer decoder layers, varying by variant (e.g.,
3for Nano,4for Large). Defined in lines 43-44. - hidden_dim: Width of the decoder hidden state, typically
256for Base and Small variants (lines 45-46). - sa_nheads: Number of heads in decoder self-attention, usually 8-12 (lines 51-52).
- ca_nheads: Number of heads in decoder cross-attention, typically 16-24 (lines 53-54).
- dec_n_points: Deformable attention sampling points per head per pyramid level in the decoder, ranging from 2-4 (lines 55-56).
- projector_scale: List of pyramid levels fed to decoder cross-attention (e.g.,
["P4"]for Base,["P3", "P4", "P5"]for multi-scale). Specified in lines 78-80.
Input, Output, and Query Parameters
These fields control the model's input resolution, output space, and query mechanism:
- resolution: Input image size in pixels (square), ranging from 384-768 depending on variant (lines 99-100).
- positional_encoding_size: Side length of the sinusoidal positional grid in patches, computed automatically as
resolution // patch_size(lines 57-58). - num_queries: Number of object queries during inference, defaulting to
300(lines 60-66). This also determines training query groups. - num_select: Number of queries retained after post-processing, typically mirroring
num_queries(lines 87-89). - num_classes: Number of output classes, defaulting to
90for COCO (lines 95-96). - bbox_reparam: Boolean enabling bounding box re-parameterization for improved training stability (default
True, lines 90-92).
Training Optimization Flags
Memory and compute optimization parameters include:
- gradient_checkpointing: Trade compute for memory by checkpointing activations (default
False, lines 71-73). - amp: Enable automatic mixed precision training (default
True, lines 93-95). - compile: Use
torch.compilefor optimized inference graphs (defaultFalse, lines 66-68). - layer_norm: Apply layer normalization in the decoder (default
True, lines 93-95). - group_detr: Number of duplicate query groups for GroupPose-style training, defaulting to
13(lines 63-66). - lite_refpoint_refine: Use lightweight reference-point refinement (default
True, lines 92-94).
Task-Specific Parameters
For segmentation and keypoint variants, additional parameters appear in specialized config subclasses:
- mask_downsample_ratio: Down-sampling factor for segmentation masks, default
4(lines 115-117). - use_grouppose_keypoints: Enable GroupPose-style keypoint estimation (lines 107-110).
- keypoint_cross_attn: Enable cross-attention mechanisms for keypoint heads.
- inter_instance_kp_attn: Enable inter-instance keypoint attention.
- grouppose_keypoint_dim_downscale: Dimension reduction factor for keypoint representations.
- num_keypoints_per_class: List defining keypoint counts per class (e.g.,
[17]for COCO pose), defaulting to empty list for detection models (lines 113-115). - dual_projector: Use dual feature projection for multi-task heads.
- pretrain_weights: Path or URL to checkpoint loading (e.g.,
rf-detr-small.pth), orNonefor training from scratch (lines 96-98).
Model Variants and Default Configurations
Each concrete subclass in src/rfdetr/config.py inherits from ModelConfig and supplies variant-specific defaults. The class hierarchy includes detection variants (Nano, Small, Medium, Large), segmentation variants (SegMedium, SegLarge, SegXLarge), and keypoint variants (KeypointPreview).
- Nano: Configures 384px resolution, 2 windows, 2 decoder layers, patch size 16, and 300 queries (lines 801-810).
- Small: Uses 384-512px resolution, 3 decoder layers, and specific feature indexes
[3, 6, 9, 12]. - Large (Detection): Implements 704px resolution, 4 decoder layers, and loads
rf-detr-large-2026.pthpretrained weights (lines 837-860). - Segmentation XLarge: Configures 624px resolution, 6 decoder layers, and enables the segmentation head (lines 444-556).
- Keypoint Preview: Enables group-pose keypoints with 100 queries, 576px resolution, and dual projector configuration (lines 775-795).
Practical Configuration Examples
Customizing a Small Model for High-Resolution Input
Override default parameters when instantiating the model to adjust resolution and query count:
from rfdetr import RFDETR
from rfdetr.config import RFDETRSmallConfig
custom_cfg = RFDETRSmallConfig(
resolution=640, # Increase from default 384/512
num_queries=500, # More object queries for dense scenes
amp=False, # Disable mixed precision for debugging
compile=True, # Enable torch.compile for faster inference
mask_downsample_ratio=2 # Finer mask resolution for segmentation
)
model = RFDETR(cfg=custom_cfg)
Configuring a Segmentation Model
Segmentation variants extend the base config with mask-specific parameters:
from rfdetr import RFDETR
from rfdetr.config import RFDETRSegMediumConfig
seg_cfg = RFDETRSegMediumConfig(
resolution=640,
num_queries=200,
mask_downsample_ratio=4, # Standard for medium segmentation
segmentation_head=True # Explicitly enabled (default for Seg configs)
)
seg_model = RFDETR(cfg=seg_cfg)
Enabling Keypoint Estimation
Keypoint models require additional flags for pose estimation heads:
from rfdetr import RFDETR
from rfdetr.config import RFDETRKeypointPreviewConfig
kp_cfg = RFDETRKeypointPreviewConfig(
resolution=576,
num_keypoints_per_class=[17], # COCO-format 17 keypoints
use_grouppose_keypoints=True, # Enable GroupPose architecture
dual_projector=True, # Use dual projection layers
dual_projector_kp_only=True, # Isolate keypoint projection
num_queries=100 # Reduced queries for pose tasks
)
kp_model = RFDETR(cfg=kp_cfg)
Key Implementation Files
The configuration system spans several critical files in the roboflow/rf-detr repository:
src/rfdetr/config.py: DefinesModelConfigand all concrete variant configurations (lines 33-120 for base class, lines 400+ for variants).src/rfdetr/detr.py: Instantiates theRFDETRmodel from aModelConfiginstance, wiring the backbone, projector, and transformer decoder.src/rfdetr/models/transformer.py: Implements the decoder transformer usingdec_layers,hidden_dim, and attention head parameters.src/rfdetr/models/backbone/: Contains DINOv2 windowed backbone implementations referenced by theencoderparameter.src/rfdetr/models/postprocess.py: Consumesnum_queriesandnum_selectduring inference to generate final predictions.
Summary
- All RF-DETR architectures derive from the
ModelConfigPydantic class insrc/rfdetr/config.py, providing type-safe parameter validation. - Backbone selection is controlled via
encoderandpatch_size, while decoder complexity is tuned throughdec_layers,hidden_dim, and attention head counts (sa_nheads,ca_nheads). - Input resolution (
resolution) and query count (num_queries) are the primary levers for trading speed against accuracy, with variant configs ranging from 384px (Nano) to 768px (Large). - Task specialization occurs through subclass parameters like
mask_downsample_ratiofor segmentation anduse_grouppose_keypointsfor pose estimation. - Optimization features including
torch.compile(compile), gradient checkpointing, and automatic mixed precision (amp) are exposed as boolean flags in the config.
Frequently Asked Questions
How do I change the input resolution for a pretrained RF-DETR model?
Modify the resolution parameter when constructing your config instance. For example, instantiate RFDETRSmallConfig(resolution=640) instead of using the default 384 or 512 pixels. Note that changing resolution affects the computed positional_encoding_size (calculated as resolution // patch_size on lines 57-58 of src/rfdetr/config.py), so ensure your custom resolution is divisible by the model's patch_size (typically 14 or 16).
What is the difference between num_queries and num_select?
num_queries defines how many object queries the transformer decoder generates and processes during inference (default 300), affecting both memory consumption and the maximum number of detections per image. num_select determines how many of these queries survive post-processing and contribute to the final prediction set. By default, num_select mirrors num_queries (lines 87-89), but you can reduce it to enforce stricter confidence thresholds or speed up NMS operations.
How do I enable torch.compile for faster RF-DETR inference?
Set compile=True in your configuration object before instantiating the model. According to lines 66-68 in src/rfdetr/config.py, this flag triggers torch.compile optimization on the assembled model. This is particularly effective for Large variants deployed in production environments, though it may increase initial model loading time.
Which parameters control the decoder depth and complexity?
The decoder architecture is primarily governed by four parameters: dec_layers (number of transformer layers, lines 43-44), hidden_dim (channel dimension, lines 45-46), sa_nheads (self-attention heads, lines 51-52), and ca_nheads (cross-attention heads, lines 53-54). For example, the Nano variant uses dec_layers=2 while the Large variant uses dec_layers=4, directly impacting parameter count and inference latency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →