How to Use the Auto-Batch Feature for RF-DETR Training

Set batch_size='auto' in your training configuration to automatically calculate the maximum safe micro-batch size and required gradient accumulation steps for your GPU hardware.

The roboflow/rf-detr repository provides an intelligent auto-batch system that eliminates manual GPU memory tuning during transformer-based object detection training. This feature probes available VRAM using a synthetic worst-case training step to determine optimal batch parameters before training begins.

How the Auto-Batch System Works

The auto-batch mechanism resides in src/rfdetr/training/auto_batch.py and simulates a full training iteration—including optimizer state allocation—to find memory-safe configurations.

Core Probing Logic

The resolve_auto_batch_config function orchestrates the probe by building a synthetic batch and executing probe_max_micro_batch. This function performs an exponential-then-binary search to discover the largest batch size that fits in GPU memory while accounting for the shadow AdamW optimizer state. [auto_batch.py#L44-L55]

The _probe_step function creates synthetic inputs via _make_synthetic_batch and runs them through _build_shadow_optimizer, which mirrors the real model's parameter shapes without updating actual weights. This ensures the probe accounts for lazy allocator behavior while preserving loaded checkpoints. [_build_shadow_optimizer#L40-L56]

Gradient Accumulation Calculation

After determining the safe micro-batch size, recommend_grad_accum_steps calculates the required accumulation steps to reach your target effective batch size. The system returns an AutoBatchResult containing the safe micro-batch size, recommended gradient accumulation steps, effective batch size, and GPU device name. [recommend_grad_accum_steps#L24-L33]

Enabling Auto-Batch via Python API

Activate auto-batch by setting batch_size="auto" in your TrainConfig and specifying your desired effective batch size:

from rfdetr.config import ModelConfig, TrainConfig
from rfdetr.training.trainer import RFDETRTrainer

# Configure model

model_cfg = ModelConfig(
    model_name="rfdetr-small",
    resolution=640,
    num_classes=80,
    amp=True,
)

# Enable auto-batch

train_cfg = TrainConfig(
    batch_size="auto",                   # Activates GPU probing

    auto_batch_target_effective=64,     # Desired effective batch per device

    optimizer="adamw",
    devices=1,
)

# Initialize trainer (probe runs automatically)

trainer = RFDETRTrainer(model_cfg, train_cfg)
trainer.fit()

The probe executes during trainer initialization, automatically setting trainer.train_config.batch_size to the safe micro-batch value and trainer.train_config.grad_accum_steps to the calculated accumulation count.

Using Auto-Batch with the CLI

The training CLI in src/rfdetr/training/cli.py parses --batch-size auto and forwards the configuration to the trainer:

uv run --no-sync rfdetr train \
    --model-name rfdetr-small \
    --resolution 640 \
    --batch-size auto \
    --auto-batch-target-effective 64 \
    --optimizer adamw \
    --devices 1 \
    --amp

This triggers the same probing logic as the Python API, ensuring consistent behavior across interfaces.

Accessing Probe Results Programmatically

After initialization, inspect the AutoBatchResult object stored on the trainer instance:

result = trainer.auto_batch_result

print(f"Safe micro-batch: {result.safe_micro_batch}")
print(f"Gradient accumulation steps: {result.recommended_grad_accum_steps}")
print(f"Effective batch size: {result.effective_batch_size}")
print(f"GPU device: {result.device_name}")

This object is populated immediately after resolve_auto_batch_config completes the memory probe. [AutoBatchResult creation]

Advanced Configuration Options

Fine-tune the probing behavior using these TrainConfig parameters:

  • auto_batch_max_targets_per_image: Simulates worst-case memory usage by specifying the maximum number of objects per image (default varies by resolution). Increase this for datasets with dense annotations. [resolve_auto_batch_config#L84-L100]

  • auto_batch_ema_headroom: Applies additional memory buffer when using Exponential Moving Average (ema), specified as a float multiplier (e.g., 0.6 reserves 60% extra headroom).

  • safety_margin: Scales the final micro-batch size down by this factor (default 0.9) to prevent out-of-memory errors from CUDA fragmentation.

  • max_micro_batch: Sets the upper bound for the binary search (default 128). Increase this for GPUs with very large VRAM capacity.

Summary

  • Set batch_size="auto" in TrainConfig to enable automatic GPU memory probing
  • The system uses a shadow optimizer in src/rfdetr/training/auto_batch.py to simulate real training memory consumption without modifying model weights
  • probe_max_micro_batch performs exponential-then-binary search to find the maximum safe batch size
  • Specify auto_batch_target_effective to let the system calculate required gradient accumulation steps automatically
  • Access probe results via trainer.auto_batch_result after initialization
  • Customize probe behavior with auto_batch_max_targets_per_image, safety_margin, and max_micro_batch parameters

Frequently Asked Questions

What happens if the auto-batch probe fails to find a valid batch size?

If the probe cannot fit even a single sample into memory, it raises a RuntimeError indicating insufficient GPU memory. This typically occurs when using high resolutions (e.g., 1024px) on GPUs with limited VRAM. Reduce the --resolution or enable mixed precision (--amp) to decrease memory requirements.

Does auto-batch work with multi-GPU training?

Yes, the auto-batch feature probes memory on each GPU independently when using distributed training. Set devices to your desired GPU count or -1 for all available devices. Each device will calculate its own safe micro-batch size based on individual GPU memory capacity.

How does the shadow optimizer differ from the real optimizer?

The shadow optimizer in _build_shadow_optimizer creates parameter state tensors matching the shapes and dtypes of the real AdamW optimizer but never actually updates model parameters. This allows the probe to force CUDA memory allocation for optimizer states without corrupting loaded checkpoints or affecting random seeds.

Can I use auto-batch with different optimizers like SGD?

The current implementation in src/rfdetr/training/auto_batch.py specifically probes using AdamW state allocation patterns. While the calculated micro-batch size generally works for other optimizers, AdamW has higher per-parameter memory overhead (maintaining both first and second moment estimates), making it the safest baseline for memory estimation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →