# RF-DETR Model Architecture Parameters: Complete Configuration Guide

> Explore RF-DETR model architecture parameters with this comprehensive guide. Understand and control hyperparameters for vision transformer backbones, decoders, attention, and heads.

- Repository: [Roboflow/rf-detr](https://github.com/roboflow/rf-detr)
- Tags: api-reference
- Published: 2026-09-08

---

**The RF-DETR architecture exposes all hyperparameters through a Pydantic-based `ModelConfig` class defined in [`src/rfdetr/config.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/config.py), allowing granular control over the vision transformer backbone, decoder depth, attention mechanisms, and task-specific heads.**

RF-DETR is a real-time transformer-based object detector developed by Roboflow that unifies detection, segmentation, and keypoint estimation behind a single configuration interface. Every architectural decision—from the DINOv2 windowed backbone selection to the number of deformable attention points—is exposed as a typed parameter in the configuration system. Understanding these configurable parameters enables researchers and engineers to scale the model from edge devices (Nano) to high-accuracy cloud deployments (XLarge) or adapt it for custom computer vision tasks.

## Core Architectural Parameters

The base `ModelConfig` class in [`src/rfdetr/config.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/config.py) (lines 33-120) defines the superset of all tunable architectural fields. Concrete variants inherit from this class and override values to create the Nano, Small, Medium, Large, and Segmentation profiles.

### Backbone and Encoder Settings

The vision transformer backbone behavior is controlled through encoder-specific parameters:

- **encoder**: String identifier for the vision transformer backbone (e.g., `dinov2_windowed_small`, `dinov2_windowed_base`). This field is mandatory and defined by each concrete subclass (lines 33-41).
- **out_feature_indexes**: List of layer indices whose feature maps are forwarded to the decoder (e.g., `[3, 6, 9, 12]` for the Small variant). These indices determine which hierarchical features feed the detection head (lines 40-44).
- **patch_size**: ViT patch dimension in pixels, typically `14` for Base models or `16` for smaller variants (lines 47-48).
- **num_windows**: Number of windowed-attention windows in the backbone, ranging from 2-4 depending on model scale (lines 49-50).
- **backbone_lora**: Boolean flag to enable Low-Rank Adaptation (LoRA) fine-tuning of the backbone weights (default `False`).
- **freeze_encoder**: Boolean to freeze encoder parameters during training (default `False`). Both flags are defined in lines 117-120.

### Decoder and Attention Configuration

The transformer decoder architecture is governed by attention and depth parameters:

- **dec_layers**: Number of transformer decoder layers, varying by variant (e.g., `3` for Nano, `4` for Large). Defined in lines 43-44.
- **hidden_dim**: Width of the decoder hidden state, typically `256` for Base and Small variants (lines 45-46).
- **sa_nheads**: Number of heads in decoder self-attention, usually 8-12 (lines 51-52).
- **ca_nheads**: Number of heads in decoder cross-attention, typically 16-24 (lines 53-54).
- **dec_n_points**: Deformable attention sampling points per head per pyramid level in the decoder, ranging from 2-4 (lines 55-56).
- **projector_scale**: List of pyramid levels fed to decoder cross-attention (e.g., `["P4"]` for Base, `["P3", "P4", "P5"]` for multi-scale). Specified in lines 78-80.

### Input, Output, and Query Parameters

These fields control the model's input resolution, output space, and query mechanism:

- **resolution**: Input image size in pixels (square), ranging from 384-768 depending on variant (lines 99-100).
- **positional_encoding_size**: Side length of the sinusoidal positional grid in patches, computed automatically as `resolution // patch_size` (lines 57-58).
- **num_queries**: Number of object queries during inference, defaulting to `300` (lines 60-66). This also determines training query groups.
- **num_select**: Number of queries retained after post-processing, typically mirroring `num_queries` (lines 87-89).
- **num_classes**: Number of output classes, defaulting to `90` for COCO (lines 95-96).
- **bbox_reparam**: Boolean enabling bounding box re-parameterization for improved training stability (default `True`, lines 90-92).

### Training Optimization Flags

Memory and compute optimization parameters include:

- **gradient_checkpointing**: Trade compute for memory by checkpointing activations (default `False`, lines 71-73).
- **amp**: Enable automatic mixed precision training (default `True`, lines 93-95).
- **compile**: Use `torch.compile` for optimized inference graphs (default `False`, lines 66-68).
- **layer_norm**: Apply layer normalization in the decoder (default `True`, lines 93-95).
- **group_detr**: Number of duplicate query groups for GroupPose-style training, defaulting to `13` (lines 63-66).
- **lite_refpoint_refine**: Use lightweight reference-point refinement (default `True`, lines 92-94).

### Task-Specific Parameters

For segmentation and keypoint variants, additional parameters appear in specialized config subclasses:

- **mask_downsample_ratio**: Down-sampling factor for segmentation masks, default `4` (lines 115-117).
- **use_grouppose_keypoints**: Enable GroupPose-style keypoint estimation (lines 107-110).
- **keypoint_cross_attn**: Enable cross-attention mechanisms for keypoint heads.
- **inter_instance_kp_attn**: Enable inter-instance keypoint attention.
- **grouppose_keypoint_dim_downscale**: Dimension reduction factor for keypoint representations.
- **num_keypoints_per_class**: List defining keypoint counts per class (e.g., `[17]` for COCO pose), defaulting to empty list for detection models (lines 113-115).
- **dual_projector**: Use dual feature projection for multi-task heads.
- **pretrain_weights**: Path or URL to checkpoint loading (e.g., `rf-detr-small.pth`), or `None` for training from scratch (lines 96-98).

## Model Variants and Default Configurations

Each concrete subclass in [`src/rfdetr/config.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/config.py) inherits from `ModelConfig` and supplies variant-specific defaults. The class hierarchy includes detection variants (Nano, Small, Medium, Large), segmentation variants (SegMedium, SegLarge, SegXLarge), and keypoint variants (KeypointPreview).

- **Nano**: Configures 384px resolution, 2 windows, 2 decoder layers, patch size 16, and 300 queries (lines 801-810).
- **Small**: Uses 384-512px resolution, 3 decoder layers, and specific feature indexes `[3, 6, 9, 12]`.
- **Large (Detection)**: Implements 704px resolution, 4 decoder layers, and loads `rf-detr-large-2026.pth` pretrained weights (lines 837-860).
- **Segmentation XLarge**: Configures 624px resolution, 6 decoder layers, and enables the segmentation head (lines 444-556).
- **Keypoint Preview**: Enables group-pose keypoints with 100 queries, 576px resolution, and dual projector configuration (lines 775-795).

## Practical Configuration Examples

### Customizing a Small Model for High-Resolution Input

Override default parameters when instantiating the model to adjust resolution and query count:

```python
from rfdetr import RFDETR
from rfdetr.config import RFDETRSmallConfig

custom_cfg = RFDETRSmallConfig(
    resolution=640,          # Increase from default 384/512

    num_queries=500,         # More object queries for dense scenes

    amp=False,               # Disable mixed precision for debugging

    compile=True,            # Enable torch.compile for faster inference

    mask_downsample_ratio=2  # Finer mask resolution for segmentation

)

model = RFDETR(cfg=custom_cfg)

```

### Configuring a Segmentation Model

Segmentation variants extend the base config with mask-specific parameters:

```python
from rfdetr import RFDETR
from rfdetr.config import RFDETRSegMediumConfig

seg_cfg = RFDETRSegMediumConfig(
    resolution=640,
    num_queries=200,
    mask_downsample_ratio=4,  # Standard for medium segmentation

    segmentation_head=True    # Explicitly enabled (default for Seg configs)

)

seg_model = RFDETR(cfg=seg_cfg)

```

### Enabling Keypoint Estimation

Keypoint models require additional flags for pose estimation heads:

```python
from rfdetr import RFDETR
from rfdetr.config import RFDETRKeypointPreviewConfig

kp_cfg = RFDETRKeypointPreviewConfig(
    resolution=576,
    num_keypoints_per_class=[17],      # COCO-format 17 keypoints

    use_grouppose_keypoints=True,      # Enable GroupPose architecture

    dual_projector=True,               # Use dual projection layers

    dual_projector_kp_only=True,       # Isolate keypoint projection

    num_queries=100                    # Reduced queries for pose tasks

)

kp_model = RFDETR(cfg=kp_cfg)

```

## Key Implementation Files

The configuration system spans several critical files in the `roboflow/rf-detr` repository:

- **[`src/rfdetr/config.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/config.py)**: Defines `ModelConfig` and all concrete variant configurations (lines 33-120 for base class, lines 400+ for variants).
- **[`src/rfdetr/detr.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/detr.py)**: Instantiates the `RFDETR` model from a `ModelConfig` instance, wiring the backbone, projector, and transformer decoder.
- **[`src/rfdetr/models/transformer.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/models/transformer.py)**: Implements the decoder transformer using `dec_layers`, `hidden_dim`, and attention head parameters.
- **`src/rfdetr/models/backbone/`**: Contains DINOv2 windowed backbone implementations referenced by the `encoder` parameter.
- **[`src/rfdetr/models/postprocess.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/models/postprocess.py)**: Consumes `num_queries` and `num_select` during inference to generate final predictions.

## Summary

- **All RF-DETR architectures** derive from the `ModelConfig` Pydantic class in [`src/rfdetr/config.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/config.py), providing type-safe parameter validation.
- **Backbone selection** is controlled via `encoder` and `patch_size`, while decoder complexity is tuned through `dec_layers`, `hidden_dim`, and attention head counts (`sa_nheads`, `ca_nheads`).
- **Input resolution** (`resolution`) and **query count** (`num_queries`) are the primary levers for trading speed against accuracy, with variant configs ranging from 384px (Nano) to 768px (Large).
- **Task specialization** occurs through subclass parameters like `mask_downsample_ratio` for segmentation and `use_grouppose_keypoints` for pose estimation.
- **Optimization features** including `torch.compile` (`compile`), gradient checkpointing, and automatic mixed precision (`amp`) are exposed as boolean flags in the config.

## Frequently Asked Questions

### How do I change the input resolution for a pretrained RF-DETR model?

Modify the `resolution` parameter when constructing your config instance. For example, instantiate `RFDETRSmallConfig(resolution=640)` instead of using the default 384 or 512 pixels. Note that changing resolution affects the computed `positional_encoding_size` (calculated as `resolution // patch_size` on lines 57-58 of [`src/rfdetr/config.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/config.py)), so ensure your custom resolution is divisible by the model's `patch_size` (typically 14 or 16).

### What is the difference between `num_queries` and `num_select`?

`num_queries` defines how many object queries the transformer decoder generates and processes during inference (default 300), affecting both memory consumption and the maximum number of detections per image. `num_select` determines how many of these queries survive post-processing and contribute to the final prediction set. By default, `num_select` mirrors `num_queries` (lines 87-89), but you can reduce it to enforce stricter confidence thresholds or speed up NMS operations.

### How do I enable `torch.compile` for faster RF-DETR inference?

Set `compile=True` in your configuration object before instantiating the model. According to lines 66-68 in [`src/rfdetr/config.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/config.py), this flag triggers `torch.compile` optimization on the assembled model. This is particularly effective for Large variants deployed in production environments, though it may increase initial model loading time.

### Which parameters control the decoder depth and complexity?

The decoder architecture is primarily governed by four parameters: `dec_layers` (number of transformer layers, lines 43-44), `hidden_dim` (channel dimension, lines 45-46), `sa_nheads` (self-attention heads, lines 51-52), and `ca_nheads` (cross-attention heads, lines 53-54). For example, the Nano variant uses `dec_layers=2` while the Large variant uses `dec_layers=4`, directly impacting parameter count and inference latency.