RF-DETR Resolution Constraints: Valid Input Sizes for Training and Inference

RF-DETR requires square image inputs where the side length must be divisible by the model's specific block size (patch_size × num_windows), which equals 32 for detection models and 24 for most segmentation variants.

The roboflow/rf-detr repository implements a transformer-based detection architecture that enforces strict resolution constraints to ensure proper patch embedding alignment. Understanding these RF-DETR resolution constraints is essential for training custom models, running inference, and exporting to production formats like ONNX or TensorFlow Lite.

The Divisor Rule: patch_size × num_windows

At the core of RF-DETR's input validation is a mathematical constraint enforced across the entire pipeline. According to src/rfdetr/variants.py, any resolution must satisfy:


resolution % (patch_size × num_windows) == 0

This formula ensures that the internal patch-embedding and positional-encoding layers can tile the image without leftover pixels. As documented in docs/learn/train/training-parameters.md, violating this rule causes immediate validation failures during model construction, preventing tensor shape mismatches in the transformer backbone.

Resolution Requirements by Model Type

Different RF-DETR variants utilize different architectural parameters that determine valid input dimensions.

Detection Models

All official detection checkpoints use patch_size=16 and num_windows=2, resulting in a block size of 32 pixels. Consequently, valid resolutions must be multiples of 32 (e.g., 320, 512, 640, 768). The FAQ documentation in docs/faq.md explicitly confirms this requirement for current detection checkpoints.

Segmentation Models

Most segmentation variants require input resolutions divisible by 24. However, the RFDETRSegNano model uses a reduced divisor of 12, enabling smaller input sizes optimized for edge deployment scenarios.

Validating Resolutions in Python

The model constructor validates resolutions immediately upon instantiation. If you provide an invalid resolution, the constructor raises a ValueError before loading any weights, as implemented in src/rfdetr/variants.py.

Valid custom resolution:

from rfdetr import RFDETRSmall

model = RFDETRSmall(resolution=640)   # 640 ÷ 32 = 20 → Valid

print(model.model_config.resolution)  # 640

Invalid resolution handling:

from rfdetr import RFDETRSmall

try:
    RFDETRSmall(resolution=630)       # 630 ÷ 32 = 19.6875 → Invalid

except ValueError as e:
    print(e)  # "resolution must be divisible by patch_size * num_windows (32)"

Constraints for Training, Inference, and Export

Training Configuration

When configuring training via TrainConfig, the resolution parameter follows the same divisor rule. The trainer validates this value before launching to prevent mid-training failures due to shape mismatches in the transformer layers.

Runtime Inference

During inference, the predict() method accepts a shape parameter that must respect the divisor rule. Both dimensions must be valid multiples, though RF-DETR expects square inputs for optimal detection performance.


# Valid inference shape

preds = model.predict(
    image_batch,
    shape=(768, 768)           # 768 ÷ 32 = 24 → Valid

)

Model Export

Export functions for ONNX and TFLite enforce these constraints to ensure compatibility with optimized runtime engines. As noted in docs/learn/export.md, the input shape must be divisible by the selected model's block size.

from rfdetr.export import export_onnx

export_onnx(
    model=model,
    output_path="rf_detr_small_640.onnx",
    shape=(640, 640),   # Both dimensions divisible by 32

)

Summary

  • RF-DETR requires square image inputs with side lengths divisible by patch_size × num_windows
  • Detection models require multiples of 32; most segmentation models require multiples of 24 (or 12 for Nano variants)
  • src/rfdetr/variants.py validates resolutions during construction, raising ValueError for invalid sizes
  • TrainConfig.resolution, inference calls, and export functions all enforce identical constraints
  • These rules ensure proper tensor alignment between patch embeddings and positional encodings across the transformer architecture

Frequently Asked Questions

What happens if I specify an invalid resolution for RF-DETR?

The model constructor immediately raises a ValueError with a descriptive message indicating the required divisor. This validation occurs in src/rfdetr/variants.py before weight initialization, preventing runtime errors during the forward pass.

Can I use non-square images with RF-DETR?

While RF-DETR is architecturally designed for square inputs, if you specify rectangular shapes during inference or export, both dimensions must individually satisfy the divisor rule. However, square inputs are strongly recommended to maintain detection accuracy and architectural consistency.

Why does RF-DETR have stricter resolution constraints than CNN-based detectors?

Transformer-based architectures rely on fixed-size patch embeddings that must tile the input image perfectly without fractional patches. The constraint ensures that positional encodings align correctly with the patch grid, which is critical for the attention mechanisms operating within the encoder layers.

How do I calculate the minimum valid resolution for my RF-DETR model?

Multiply the model's patch_size by num_windows to determine the block size. For official detection models, this equals 32 pixels, making 320×320 the theoretical minimum, though 640×640 is the default for most pretrained checkpoints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →