# How Exponential Moving Average (EMA) Improves YOLOv5 Model Performance

> Discover how Exponential Moving Average EMA boosts YOLOv5 performance. Achieve higher mAP and stable training with this weight smoothing technique. Learn more!

- Repository: [Ultralytics/yolov5](https://github.com/ultralytics/yolov5)
- Tags: performance
- Published: 2026-03-06

---

**Exponential Moving Average (EMA) in YOLOv5 maintains a shadow copy of model weights that smooths training updates, resulting in higher validation mAP and more stable convergence without increasing inference cost.**

The `ultralytics/yolov5` repository implements EMA through the `ModelEMA` class to reduce optimization variance and produce better-generalizing final models. This technique creates a slowly evolving average of training weights that dampens sudden parameter spikes, ultimately yielding superior detection accuracy compared to raw training weights alone.

## How EMA Works in YOLOv5

### Shadow Model Initialization

When training begins, the `ModelEMA` class in [`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py) constructs a deep-copied, FP32 shadow model from the current training weights and immediately sets it to evaluation mode. This initialization occurs at lines 44-55 of [`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py), ensuring the EMA model maintains a separate, high-precision copy of parameters throughout training.

### Adaptive Decay Scheduling

EMA applies a **decay factor** that follows a ramp-up schedule during early training epochs before stabilizing. According to the implementation at lines 50-57 of [`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py), the effective decay calculates as:

```python
decay * (1 - math.exp(-updates / tau))

```

Where `updates` tracks the number of EMA updates and `tau` controls the warmup speed. The decay stabilizes near the user-specified default of **0.9999** after the initial ramp-up period, allowing the shadow model to initially adapt quickly before settling into a slow-moving average.

### Weight Update Mechanism

After each optimizer step, the EMA weights blend the current model parameters with the historical average using the formula implemented at lines 60-70 of [`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py):

```python
ema_weight = decay * ema_weight + (1 - decay) * model_weight

```

This update acts as a low-pass filter over the optimization trajectory, preventing sudden weight fluctuations from destabilizing the learned representation. The `update()` method iterates through all trainable parameters, applying this blending operation to maintain the smoothed shadow copy.

## EMA Integration in the Training Loop

The primary training scripts—[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py), [`segment/train.py`](https://github.com/ultralytics/yolov5/blob/main/segment/train.py), and [`classify/train.py`](https://github.com/ultralytics/yolov5/blob/main/classify/train.py)—integrate EMA instantiation and updates into their standard workflows. As implemented in [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) at lines 63-66, the `ModelEMA` object initializes only on the primary process (`RANK in {-1, 0}`) to avoid redundant computation in distributed training scenarios.

During each training batch, the scripts call `ema.update(model)` immediately after `optimizer.step()`, ensuring the shadow weights evolve alongside the actively trained parameters. The EMA model also supports checkpoint resumption, allowing training to restore both the optimizer state and the shadow weights when continuing from a saved checkpoint.

## Performance Impact of EMA on YOLOv5

EMA improves YOLOv5 performance through three primary mechanisms:

- **Higher Validation mAP**: The smoothed EMA weights consistently achieve superior mean Average Precision on validation sets compared to raw training weights, as the averaged parameters represent a more optimal point in the loss landscape.
- **Stable Loss Curves**: By dampening optimization noise, EMA produces smoother training and validation loss trajectories, making it easier to detect true convergence versus temporary instability.
- **Zero Inference Cost**: The EMA weights become the final model parameters after training, delivering improved accuracy without requiring separate inference-time computations or model ensemble overhead.

The EMA model serves as the authoritative version for final evaluation and model export, meaning the accuracy improvements transfer directly to production inference.

## Implementing EMA in YOLOv5

### Enabling EMA in Custom Training Scripts

To manually implement EMA in a custom YOLOv5 training loop, instantiate `ModelEMA` after creating your model and optimizer:

```python
from utils.torch_utils import ModelEMA
import torch

# Assume `model` is your YOLOv5 model and `optimizer` is already defined

ema = ModelEMA(model)               # Creates shadow EMA model (FP32, eval mode)

for epoch in range(num_epochs):
    for imgs, targets in dataloader:
        optimizer.zero_grad()
        loss, *_ = model(imgs, targets)   # Normal training step

        loss.backward()
        optimizer.step()
        ema.update(model)                 # Update EMA weights each batch

```

This pattern mirrors the internal implementation found in [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) (line 63) and [`segment/train.py`](https://github.com/ultralytics/yolov5/blob/main/segment/train.py) (line 226), ensuring consistency with the official YOLOv5 training pipeline.

### Saving and Loading EMA Weights

To preserve EMA state between training sessions or deploy the optimized weights:

```python

# After training - save both raw and EMA weights

torch.save({
    'model': model.state_dict(),
    'ema': ema.ema.state_dict(),          # Store EMA shadow model separately

    'updates': ema.updates                # Save update counter for decay scheduling

}, 'best.pt')

# Loading for inference - replace model weights with EMA parameters

ckpt = torch.load('best.pt')
model.load_state_dict(ckpt['ema'])       # Load EMA weights as final model

model.eval()

```

### Using EMA with the Built-in CLI

When using the official training scripts, EMA activates automatically on the main process without additional configuration:

```bash
python train.py --data coco.yaml --cfg yolov5s.yaml --epochs 300

# EMA automatically instantiates via ModelEMA in utils/torch_utils.py

```

The decay factor and other EMA hyperparameters inherit from the default values defined in [`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py) unless explicitly modified in the training script.

## Summary

- **EMA creates a shadow model** in [`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py) that maintains FP32 precision and evaluation mode throughout YOLOv5 training.
- **Adaptive decay scheduling** ramps up from zero to 0.9999 using an exponential warmup formula to balance early flexibility with late-stage stability.
- **Weight updates occur per batch** via `ema.update(model)`, blending new parameters with historical averages to dampen optimization noise.
- **Primary process exclusivity** ensures EMA computation only occurs on `RANK` -1 or 0, preventing distributed training overhead.
- **Superior validation performance** makes EMA weights the default choice for final model evaluation and export in detection, segmentation, and classification tasks.

## Frequently Asked Questions

### What is the default EMA decay factor in YOLOv5?

The default decay factor is **0.9999**, as specified in the `ModelEMA` class initialization within [`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py). However, the implementation applies a warmup schedule using `decay * (1 - math.exp(-updates / tau))` during early training steps, allowing the effective decay to ramp up from zero rather than starting immediately at the full value.

### Does EMA increase training time or memory usage?

EMA adds minimal computational overhead during training because it only performs parameter blending after the optimizer step completes. Memory usage increases moderately due to storing a full FP32 copy of model parameters as the shadow weights, but this typically represents less than a 50% increase over baseline training memory and enables superior final model performance without inference-time cost.

### Can I disable EMA when training YOLOv5?

Yes, you can disable EMA by modifying the training script to skip the `ModelEMA` instantiation. In [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py), remove or comment out the lines where `ema = ModelEMA(model)` is created (around line 63) and subsequent `ema.update()` calls. However, this is not recommended for production models, as the validation mAP typically drops without the smoothing benefits EMA provides.

### How do I resume training while preserving EMA state?

When resuming from a checkpoint saved by the official [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) script, the EMA state automatically restores if the checkpoint contains an `'ema'` key. Ensure your saved checkpoint includes `{'ema': ema.ema.state_dict(), 'updates': ema.updates}` alongside the model and optimizer states. The training script checks for these keys at initialization and restores both the shadow weights and the update counter to maintain the decay schedule continuity.