# How to Use the Auto-Batch Feature for RF-DETR Training

> Unlock efficient RF-DETR training with auto-batch. Effortlessly set batch_size='auto' to optimize micro-batch size and gradient accumulation for your GPU, maximizing performance.

- Repository: [Roboflow/rf-detr](https://github.com/roboflow/rf-detr)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Set `batch_size='auto'` in your training configuration to automatically calculate the maximum safe micro-batch size and required gradient accumulation steps for your GPU hardware.**

The roboflow/rf-detr repository provides an intelligent auto-batch system that eliminates manual GPU memory tuning during transformer-based object detection training. This feature probes available VRAM using a synthetic worst-case training step to determine optimal batch parameters before training begins.

## How the Auto-Batch System Works

The auto-batch mechanism resides in [`src/rfdetr/training/auto_batch.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/training/auto_batch.py) and simulates a full training iteration—including optimizer state allocation—to find memory-safe configurations.

### Core Probing Logic

The **`resolve_auto_batch_config`** function orchestrates the probe by building a synthetic batch and executing **`probe_max_micro_batch`**. This function performs an exponential-then-binary search to discover the largest batch size that fits in GPU memory while accounting for the **shadow AdamW optimizer** state. [[auto_batch.py#L44-L55]](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/training/auto_batch.py#L44-L55)

The **`_probe_step`** function creates synthetic inputs via **`_make_synthetic_batch`** and runs them through **`_build_shadow_optimizer`**, which mirrors the real model's parameter shapes without updating actual weights. This ensures the probe accounts for lazy allocator behavior while preserving loaded checkpoints. [[_build_shadow_optimizer#L40-L56]](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/training/auto_batch.py#L40-L56)

### Gradient Accumulation Calculation

After determining the safe micro-batch size, **`recommend_grad_accum_steps`** calculates the required accumulation steps to reach your target effective batch size. The system returns an **`AutoBatchResult`** containing the safe micro-batch size, recommended gradient accumulation steps, effective batch size, and GPU device name. [[recommend_grad_accum_steps#L24-L33]](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/training/auto_batch.py#L24-L33)

## Enabling Auto-Batch via Python API

Activate auto-batch by setting `batch_size="auto"` in your `TrainConfig` and specifying your desired effective batch size:

```python
from rfdetr.config import ModelConfig, TrainConfig
from rfdetr.training.trainer import RFDETRTrainer

# Configure model

model_cfg = ModelConfig(
    model_name="rfdetr-small",
    resolution=640,
    num_classes=80,
    amp=True,
)

# Enable auto-batch

train_cfg = TrainConfig(
    batch_size="auto",                   # Activates GPU probing

    auto_batch_target_effective=64,     # Desired effective batch per device

    optimizer="adamw",
    devices=1,
)

# Initialize trainer (probe runs automatically)

trainer = RFDETRTrainer(model_cfg, train_cfg)
trainer.fit()

```

The probe executes during trainer initialization, automatically setting `trainer.train_config.batch_size` to the safe micro-batch value and `trainer.train_config.grad_accum_steps` to the calculated accumulation count.

## Using Auto-Batch with the CLI

The training CLI in [`src/rfdetr/training/cli.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/training/cli.py) parses `--batch-size auto` and forwards the configuration to the trainer:

```bash
uv run --no-sync rfdetr train \
    --model-name rfdetr-small \
    --resolution 640 \
    --batch-size auto \
    --auto-batch-target-effective 64 \
    --optimizer adamw \
    --devices 1 \
    --amp

```

This triggers the same probing logic as the Python API, ensuring consistent behavior across interfaces.

## Accessing Probe Results Programmatically

After initialization, inspect the `AutoBatchResult` object stored on the trainer instance:

```python
result = trainer.auto_batch_result

print(f"Safe micro-batch: {result.safe_micro_batch}")
print(f"Gradient accumulation steps: {result.recommended_grad_accum_steps}")
print(f"Effective batch size: {result.effective_batch_size}")
print(f"GPU device: {result.device_name}")

```

This object is populated immediately after `resolve_auto_batch_config` completes the memory probe. [[AutoBatchResult creation]](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/training/auto_batch.py#L275-L283)

## Advanced Configuration Options

Fine-tune the probing behavior using these `TrainConfig` parameters:

- **`auto_batch_max_targets_per_image`**: Simulates worst-case memory usage by specifying the maximum number of objects per image (default varies by resolution). Increase this for datasets with dense annotations. [[resolve_auto_batch_config#L84-L100]](https://github.com/roboflow/rf-detr/blob/develop/src/rfdetr/training/auto_batch.py#L84-L100)
  
- **`auto_batch_ema_headroom`**: Applies additional memory buffer when using Exponential Moving Average (ema), specified as a float multiplier (e.g., `0.6` reserves 60% extra headroom).

- **`safety_margin`**: Scales the final micro-batch size down by this factor (default `0.9`) to prevent out-of-memory errors from CUDA fragmentation.

- **`max_micro_batch`**: Sets the upper bound for the binary search (default `128`). Increase this for GPUs with very large VRAM capacity.

## Summary

- Set `batch_size="auto"` in `TrainConfig` to enable automatic GPU memory probing
- The system uses a **shadow optimizer** in [`src/rfdetr/training/auto_batch.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/training/auto_batch.py) to simulate real training memory consumption without modifying model weights
- **`probe_max_micro_batch`** performs exponential-then-binary search to find the maximum safe batch size
- Specify `auto_batch_target_effective` to let the system calculate required gradient accumulation steps automatically
- Access probe results via `trainer.auto_batch_result` after initialization
- Customize probe behavior with `auto_batch_max_targets_per_image`, `safety_margin`, and `max_micro_batch` parameters

## Frequently Asked Questions

### What happens if the auto-batch probe fails to find a valid batch size?

If the probe cannot fit even a single sample into memory, it raises a `RuntimeError` indicating insufficient GPU memory. This typically occurs when using high resolutions (e.g., 1024px) on GPUs with limited VRAM. Reduce the `--resolution` or enable mixed precision (`--amp`) to decrease memory requirements.

### Does auto-batch work with multi-GPU training?

Yes, the auto-batch feature probes memory on each GPU independently when using distributed training. Set `devices` to your desired GPU count or `-1` for all available devices. Each device will calculate its own safe micro-batch size based on individual GPU memory capacity.

### How does the shadow optimizer differ from the real optimizer?

The **shadow optimizer** in `_build_shadow_optimizer` creates parameter state tensors matching the shapes and dtypes of the real AdamW optimizer but never actually updates model parameters. This allows the probe to force CUDA memory allocation for optimizer states without corrupting loaded checkpoints or affecting random seeds.

### Can I use auto-batch with different optimizers like SGD?

The current implementation in [`src/rfdetr/training/auto_batch.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/training/auto_batch.py) specifically probes using AdamW state allocation patterns. While the calculated micro-batch size generally works for other optimizers, AdamW has higher per-parameter memory overhead (maintaining both first and second moment estimates), making it the safest baseline for memory estimation.