RF-DETR Model Architecture Parameters: Complete Configuration Guide

The RF-DETR architecture exposes all hyperparameters through a Pydantic-based ModelConfig class defined in src/rfdetr/config.py, allowing granular control over the vision transformer backbone, decoder depth, attention mechanisms, and task-specific heads.

RF-DETR is a real-time transformer-based object detector developed by Roboflow that unifies detection, segmentation, and keypoint estimation behind a single configuration interface. Every architectural decision—from the DINOv2 windowed backbone selection to the number of deformable attention points—is exposed as a typed parameter in the configuration system. Understanding these configurable parameters enables researchers and engineers to scale the model from edge devices (Nano) to high-accuracy cloud deployments (XLarge) or adapt it for custom computer vision tasks.

Core Architectural Parameters

The base ModelConfig class in src/rfdetr/config.py (lines 33-120) defines the superset of all tunable architectural fields. Concrete variants inherit from this class and override values to create the Nano, Small, Medium, Large, and Segmentation profiles.

Backbone and Encoder Settings

The vision transformer backbone behavior is controlled through encoder-specific parameters:

  • encoder: String identifier for the vision transformer backbone (e.g., dinov2_windowed_small, dinov2_windowed_base). This field is mandatory and defined by each concrete subclass (lines 33-41).
  • out_feature_indexes: List of layer indices whose feature maps are forwarded to the decoder (e.g., [3, 6, 9, 12] for the Small variant). These indices determine which hierarchical features feed the detection head (lines 40-44).
  • patch_size: ViT patch dimension in pixels, typically 14 for Base models or 16 for smaller variants (lines 47-48).
  • num_windows: Number of windowed-attention windows in the backbone, ranging from 2-4 depending on model scale (lines 49-50).
  • backbone_lora: Boolean flag to enable Low-Rank Adaptation (LoRA) fine-tuning of the backbone weights (default False).
  • freeze_encoder: Boolean to freeze encoder parameters during training (default False). Both flags are defined in lines 117-120.

Decoder and Attention Configuration

The transformer decoder architecture is governed by attention and depth parameters:

  • dec_layers: Number of transformer decoder layers, varying by variant (e.g., 3 for Nano, 4 for Large). Defined in lines 43-44.
  • hidden_dim: Width of the decoder hidden state, typically 256 for Base and Small variants (lines 45-46).
  • sa_nheads: Number of heads in decoder self-attention, usually 8-12 (lines 51-52).
  • ca_nheads: Number of heads in decoder cross-attention, typically 16-24 (lines 53-54).
  • dec_n_points: Deformable attention sampling points per head per pyramid level in the decoder, ranging from 2-4 (lines 55-56).
  • projector_scale: List of pyramid levels fed to decoder cross-attention (e.g., ["P4"] for Base, ["P3", "P4", "P5"] for multi-scale). Specified in lines 78-80.

Input, Output, and Query Parameters

These fields control the model's input resolution, output space, and query mechanism:

  • resolution: Input image size in pixels (square), ranging from 384-768 depending on variant (lines 99-100).
  • positional_encoding_size: Side length of the sinusoidal positional grid in patches, computed automatically as resolution // patch_size (lines 57-58).
  • num_queries: Number of object queries during inference, defaulting to 300 (lines 60-66). This also determines training query groups.
  • num_select: Number of queries retained after post-processing, typically mirroring num_queries (lines 87-89).
  • num_classes: Number of output classes, defaulting to 90 for COCO (lines 95-96).
  • bbox_reparam: Boolean enabling bounding box re-parameterization for improved training stability (default True, lines 90-92).

Training Optimization Flags

Memory and compute optimization parameters include:

  • gradient_checkpointing: Trade compute for memory by checkpointing activations (default False, lines 71-73).
  • amp: Enable automatic mixed precision training (default True, lines 93-95).
  • compile: Use torch.compile for optimized inference graphs (default False, lines 66-68).
  • layer_norm: Apply layer normalization in the decoder (default True, lines 93-95).
  • group_detr: Number of duplicate query groups for GroupPose-style training, defaulting to 13 (lines 63-66).
  • lite_refpoint_refine: Use lightweight reference-point refinement (default True, lines 92-94).

Task-Specific Parameters

For segmentation and keypoint variants, additional parameters appear in specialized config subclasses:

  • mask_downsample_ratio: Down-sampling factor for segmentation masks, default 4 (lines 115-117).
  • use_grouppose_keypoints: Enable GroupPose-style keypoint estimation (lines 107-110).
  • keypoint_cross_attn: Enable cross-attention mechanisms for keypoint heads.
  • inter_instance_kp_attn: Enable inter-instance keypoint attention.
  • grouppose_keypoint_dim_downscale: Dimension reduction factor for keypoint representations.
  • num_keypoints_per_class: List defining keypoint counts per class (e.g., [17] for COCO pose), defaulting to empty list for detection models (lines 113-115).
  • dual_projector: Use dual feature projection for multi-task heads.
  • pretrain_weights: Path or URL to checkpoint loading (e.g., rf-detr-small.pth), or None for training from scratch (lines 96-98).

Model Variants and Default Configurations

Each concrete subclass in src/rfdetr/config.py inherits from ModelConfig and supplies variant-specific defaults. The class hierarchy includes detection variants (Nano, Small, Medium, Large), segmentation variants (SegMedium, SegLarge, SegXLarge), and keypoint variants (KeypointPreview).

  • Nano: Configures 384px resolution, 2 windows, 2 decoder layers, patch size 16, and 300 queries (lines 801-810).
  • Small: Uses 384-512px resolution, 3 decoder layers, and specific feature indexes [3, 6, 9, 12].
  • Large (Detection): Implements 704px resolution, 4 decoder layers, and loads rf-detr-large-2026.pth pretrained weights (lines 837-860).
  • Segmentation XLarge: Configures 624px resolution, 6 decoder layers, and enables the segmentation head (lines 444-556).
  • Keypoint Preview: Enables group-pose keypoints with 100 queries, 576px resolution, and dual projector configuration (lines 775-795).

Practical Configuration Examples

Customizing a Small Model for High-Resolution Input

Override default parameters when instantiating the model to adjust resolution and query count:

from rfdetr import RFDETR
from rfdetr.config import RFDETRSmallConfig

custom_cfg = RFDETRSmallConfig(
    resolution=640,          # Increase from default 384/512

    num_queries=500,         # More object queries for dense scenes

    amp=False,               # Disable mixed precision for debugging

    compile=True,            # Enable torch.compile for faster inference

    mask_downsample_ratio=2  # Finer mask resolution for segmentation

)

model = RFDETR(cfg=custom_cfg)

Configuring a Segmentation Model

Segmentation variants extend the base config with mask-specific parameters:

from rfdetr import RFDETR
from rfdetr.config import RFDETRSegMediumConfig

seg_cfg = RFDETRSegMediumConfig(
    resolution=640,
    num_queries=200,
    mask_downsample_ratio=4,  # Standard for medium segmentation

    segmentation_head=True    # Explicitly enabled (default for Seg configs)

)

seg_model = RFDETR(cfg=seg_cfg)

Enabling Keypoint Estimation

Keypoint models require additional flags for pose estimation heads:

from rfdetr import RFDETR
from rfdetr.config import RFDETRKeypointPreviewConfig

kp_cfg = RFDETRKeypointPreviewConfig(
    resolution=576,
    num_keypoints_per_class=[17],      # COCO-format 17 keypoints

    use_grouppose_keypoints=True,      # Enable GroupPose architecture

    dual_projector=True,               # Use dual projection layers

    dual_projector_kp_only=True,       # Isolate keypoint projection

    num_queries=100                    # Reduced queries for pose tasks

)

kp_model = RFDETR(cfg=kp_cfg)

Key Implementation Files

The configuration system spans several critical files in the roboflow/rf-detr repository:

  • src/rfdetr/config.py: Defines ModelConfig and all concrete variant configurations (lines 33-120 for base class, lines 400+ for variants).
  • src/rfdetr/detr.py: Instantiates the RFDETR model from a ModelConfig instance, wiring the backbone, projector, and transformer decoder.
  • src/rfdetr/models/transformer.py: Implements the decoder transformer using dec_layers, hidden_dim, and attention head parameters.
  • src/rfdetr/models/backbone/: Contains DINOv2 windowed backbone implementations referenced by the encoder parameter.
  • src/rfdetr/models/postprocess.py: Consumes num_queries and num_select during inference to generate final predictions.

Summary

  • All RF-DETR architectures derive from the ModelConfig Pydantic class in src/rfdetr/config.py, providing type-safe parameter validation.
  • Backbone selection is controlled via encoder and patch_size, while decoder complexity is tuned through dec_layers, hidden_dim, and attention head counts (sa_nheads, ca_nheads).
  • Input resolution (resolution) and query count (num_queries) are the primary levers for trading speed against accuracy, with variant configs ranging from 384px (Nano) to 768px (Large).
  • Task specialization occurs through subclass parameters like mask_downsample_ratio for segmentation and use_grouppose_keypoints for pose estimation.
  • Optimization features including torch.compile (compile), gradient checkpointing, and automatic mixed precision (amp) are exposed as boolean flags in the config.

Frequently Asked Questions

How do I change the input resolution for a pretrained RF-DETR model?

Modify the resolution parameter when constructing your config instance. For example, instantiate RFDETRSmallConfig(resolution=640) instead of using the default 384 or 512 pixels. Note that changing resolution affects the computed positional_encoding_size (calculated as resolution // patch_size on lines 57-58 of src/rfdetr/config.py), so ensure your custom resolution is divisible by the model's patch_size (typically 14 or 16).

What is the difference between num_queries and num_select?

num_queries defines how many object queries the transformer decoder generates and processes during inference (default 300), affecting both memory consumption and the maximum number of detections per image. num_select determines how many of these queries survive post-processing and contribute to the final prediction set. By default, num_select mirrors num_queries (lines 87-89), but you can reduce it to enforce stricter confidence thresholds or speed up NMS operations.

How do I enable torch.compile for faster RF-DETR inference?

Set compile=True in your configuration object before instantiating the model. According to lines 66-68 in src/rfdetr/config.py, this flag triggers torch.compile optimization on the assembled model. This is particularly effective for Large variants deployed in production environments, though it may increase initial model loading time.

Which parameters control the decoder depth and complexity?

The decoder architecture is primarily governed by four parameters: dec_layers (number of transformer layers, lines 43-44), hidden_dim (channel dimension, lines 45-46), sa_nheads (self-attention heads, lines 51-52), and ca_nheads (cross-attention heads, lines 53-54). For example, the Nano variant uses dec_layers=2 while the Large variant uses dec_layers=4, directly impacting parameter count and inference latency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →