# How to Use Mixed Precision Training with RF-DETR: FP16 Acceleration Guide

> Accelerate RF-DETR training with FP16 mixed precision. Enable this feature via your YAML config and leverage PyTorch AMP for faster object detection.

- Repository: [Roboflow/rf-detr](https://github.com/roboflow/rf-detr)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Enable mixed precision training with RF-DETR by setting `amp: true` in your YAML configuration file, which automatically activates PyTorch’s automatic mixed precision (AMP) workflow with FP16 autocasting and gradient scaling.**

RF-DETR, the open-source real-time transformer object detector developed by Roboflow, supports mixed precision training to reduce GPU memory consumption and accelerate training throughput without manual tensor type management. The implementation follows PyTorch’s standard AMP pattern, automatically handling FP16 conversion for compatible operations while preserving FP32 precision for numerically sensitive layers. This guide explains the exact configuration flags and source code implementation found in the `roboflow/rf-detr` repository.

## Configuring Mixed Precision in YAML

RF-DETR uses YAML configuration files to control training hyperparameters, including the mixed precision setting. To enable FP16 training, add the `amp` boolean flag to your configuration file (e.g., [`configs/rfdetr_small.yaml`](https://github.com/roboflow/rf-detr/blob/main/configs/rfdetr_small.yaml)).

```yaml

# configs/rfdetr_small.yaml

amp: true          # Enables automatic mixed precision training

epochs: 50
batch_size: 32

```

When `amp` is set to `true`, the trainer instantiates a `torch.cuda.amp.GradScaler` and wraps the forward pass in an autocast context. If the flag is omitted or set to `false`, the training loop defaults to full FP32 precision.

## Implementation Details in the Training Loop

The core mixed precision logic resides in [`src/rfdetr/visualize/training.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/visualize/training.py). During initialization, the code checks the configuration object and creates a gradient scaler conditionally.

```python
from torch.cuda.amp import autocast, GradScaler
import contextlib

# Initialize scaler based on config

if cfg.amp:
    scaler = GradScaler()
else:
    scaler = None

```

During each training iteration, the forward pass executes within an `autocast` context manager that automatically converts eligible operations to FP16. The implementation uses `contextlib.nullcontext()` as a no-op alternative when AMP is disabled, ensuring identical code paths for both precision modes.

```python
for images, targets in dataloader:
    optimizer.zero_grad()
    
    # Autocast context for mixed precision

    with autocast(device_type="cuda", dtype=torch.float16) if cfg.amp else contextlib.nullcontext():
        outputs = model(images)
        loss = criterion(outputs, targets)

```

### Gradient Scaling and Optimizer Steps

Mixed precision training requires scaling loss values to prevent gradient underflow in FP16. The implementation uses the scaler to scale loss before backward computation, then updates the optimizer accordingly.

```python
    if cfg.amp:
        # Scale loss, compute gradients, update weights

        scaler.scale(loss).backward()
        scaler.step(optimizer)
        scaler.update()
    else:
        # Standard FP32 training path

        loss.backward()
        optimizer.step()

```

The `scaler.scale()` method multiplies the loss by a dynamic scale factor to preserve small gradient values. After the backward pass, `scaler.step(optimizer)` safely updates parameters, and `scaler.update()` adjusts the scale factor for the next iteration.

## Launching Training with AMP Enabled

Invoke the RF-DETR training module via the CLI, passing the configuration file containing `amp: true`.

```bash
python -m rfdetr.train \
    --config configs/rfdetr_small.yaml \
    --epochs 50 \
    --batch-size 32

```

The trainer automatically detects the AMP setting from the parsed configuration and initializes the appropriate precision context. Alternatively, the documentation in [`docs/reference/training.md`](https://github.com/roboflow/rf-detr/blob/main/docs/reference/training.md) indicates you can pass an `--amp` flag via the CLI to override settings programmatically.

## Verifying Mixed Precision Activation

To confirm that mixed precision training is active during development, inspect the configuration object directly in your training script or logs.

```python
print(f"AMP enabled: {cfg.amp}")   # Output: True when mixed-precision is active

```

Additionally, GPU monitoring tools such as `nvidia-smi` will show reduced memory utilization compared to FP32 training, and training logs should indicate faster iteration times on compatible hardware.

## Summary

- **Configuration**: Set `amp: true` in your YAML config file (e.g., [`configs/rfdetr_small.yaml`](https://github.com/roboflow/rf-detr/blob/main/configs/rfdetr_small.yaml)) to enable mixed precision training with RF-DETR.
- **Implementation**: The training loop in [`src/rfdetr/visualize/training.py`](https://github.com/roboflow/rf-detr/blob/main/src/rfdetr/visualize/training.py) uses `torch.autocast` with `dtype=torch.float16` for automatic type conversion and Tensor Core acceleration.
- **Stability**: `GradScaler` handles gradient scaling via `scaler.scale(loss).backward()` to prevent underflow in FP16 computations.
- **Usage**: Run the standard RF-DETR CLI training command; the system automatically applies AMP settings from the configuration.
- **Performance**: Expect 30-50% memory reduction and significant speedup on NVIDIA Volta, Turing, Ampere, or newer GPUs with Tensor Cores.

## Frequently Asked Questions

### Does mixed precision training reduce model accuracy in RF-DETR?

No, mixed precision training maintains equivalent final accuracy to full FP32 training. PyTorch’s AMP system automatically retains FP32 precision for operations that require higher numerical stability, such as certain normalization layers and loss computations, while accelerating matmul-heavy operations with FP16.

### What GPU hardware is required for mixed precision training?

Mixed precision training requires NVIDIA GPUs with CUDA compute capability 7.0 or higher (Volta, Turing, Ampere, or newer architectures). These GPUs contain Tensor Cores specifically designed to accelerate FP16 matrix operations, delivering the maximum training speedup.

### How much memory does mixed precision training save?

Mixed precision training typically reduces GPU memory usage by approximately 30-50% compared to full FP32 training, depending on model architecture and batch dimensions. This memory reduction allows for larger batch sizes or higher resolution inputs within the same hardware constraints.

### Can I enable mixed precision without modifying the YAML config file?

Yes, RF-DETR respects configuration overrides through the CLI. While the primary mechanism uses the `amp` flag in YAML files, the training system can parse the `--amp` command-line option to enable mixed precision dynamically, as documented in the training reference guide.