# How to Set Up YOLOv5 for Multi-GPU Training with PyTorch DDP

> Learn to set up YOLOv5 for multi-GPU training with PyTorch DDP. Accelerate your object detection tasks by utilizing all your GPUs efficiently. Get started now.

- Repository: [Ultralytics/yolov5](https://github.com/ultralytics/yolov5)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Launch YOLOv5 using `torch.distributed.run` with `--nproc_per_node` matching your GPU count, ensure the global batch size is divisible by the number of GPUs, and the repository automatically initializes NCCL process groups, wraps models with `DistributedDataParallel`, and synchronizes gradients across devices.**

YOLOv5 utilizes PyTorch Distributed Data Parallel (DDP) to scale object detection training across multiple GPUs on a single machine or cluster node. According to the ultralytics/yolov5 source code, the implementation in [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) and [`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py) handles process spawning, device selection, and inter-GPU communication transparently when you use the standard launch command. This article provides the exact configuration, launch commands, and internal architecture details required for efficient multi-GPU training.

## Quick Start: Single-Node Multi-GPU Command

The standard method for YOLOv5 multi-GPU training uses the PyTorch elastic launcher to spawn one process per GPU. The command reference in **[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 9‑11** shows the recommended syntax:

```bash
python -m torch.distributed.run \
    --nproc_per_node 4 \
    --master_port 1 \
    train.py \
    --data data/coco128.yaml \
    --weights yolov5s.pt \
    --img 640 \
    --batch-size 64 \
    --device 0,1,2,3 \
    --sync-bn \
    --epochs 100

```

**Critical requirements for this command:**
- **`--nproc_per_node`** must match the number of GPUs you specify in `--device` (e.g., `4` for GPUs `0,1,2,3`)
- **`--batch-size`** represents the global batch size and must be divisible by the number of GPUs (`WORLD_SIZE`)
- **`--master_port`** can be any free port on your machine (default is 29500)

When executed, `torch.distributed.run` automatically sets the `RANK`, `LOCAL_RANK`, `WORLD_SIZE`, and `MASTER_PORT` environment variables required by the DDP backend.

## Internal DDP Architecture

Understanding how YOLOv5 handles distributed training helps debug issues and optimize performance. The workflow follows a strict initialization sequence:

### Process Initialization in main()

After parsing arguments, **[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 71‑85** check if DDP mode is active by detecting `LOCAL_RANK != -1`. If true, the script performs several validations:
- Confirms the global batch size is divisible by `WORLD_SIZE` (**[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 78‑80**)
- Verifies sufficient CUDA devices are available
- Initializes the process group using the NCCL backend (with Gloo fallback)

```python

# Conceptual flow from train.py

if LOCAL_RANK != -1:
    assert torch.cuda.device_count() > LOCAL_RANK, 'insufficient CUDA devices'
    assert batch_size % WORLD_SIZE == 0, '--batch-size must be multiple of CUDA device count'
    torch.distributed.init_process_group(backend='nccl' if is_nccl_available() else 'gloo')

```

### Device Selection with select_device()

The **`select_device()`** function in **[`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py) lines 71‑76** parses the `--device` argument and binds each process to its assigned GPU based on `LOCAL_RANK`. When DDP is active, the function ignores manual device lists and uses `LOCAL_RANK` to determine the specific CUDA device for the current process.

### Model Wrapping with smart_DDP()

Once the model is constructed, **`smart_DDP()`** in **[`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py) lines 94‑99** wraps the base `nn.Module` with `torch.nn.parallel.DistributedDataParallel`. This wrapper handles gradient synchronization across processes automatically during the backward pass, ensuring consistent weight updates without manual all-reduce operations.

### Data Loading and Batch Distribution

The global batch size specified via `--batch-size` is divided equally among ranks using integer division (`batch_size // WORLD_SIZE`). Each process calls `create_dataloader()` with `rank=LOCAL_RANK` to receive only its partition of the dataset. This sharding happens in [`utils/dataloaders.py`](https://github.com/ultralytics/yolov5/blob/main/utils/dataloaders.py), ensuring no data overlap between GPUs while maintaining the effective global batch size.

### Synchronized Batch Normalization

When you include the `--sync-bn` flag, YOLOv5 converts all `BatchNorm` layers to `SyncBatchNorm` (**[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 82‑84**). This operation uses `torch.nn.SyncBatchNorm.convert_sync_batchnorm()` to aggregate batch statistics across all GPUs during training, which improves convergence when the per-GPU batch size is small.

### Early Stopping and Checkpoint Broadcasting

To ensure clean shutdowns, **[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 105‑110** implement a broadcast mechanism. When the master process (rank 0) triggers early stopping, it broadcasts the `stop` flag to all ranks using `dist.broadcast_object_list`, ensuring all processes exit the training loop simultaneously.

## Alternative Launch Methods

While the command-line launcher is standard, YOLOv5 supports programmatic multi-GPU training.

### Using the Python API

You can invoke training via the `run()` function, which internally handles DDP setup when multiple GPUs are detected:

```python
from train import run

run(
    data='data/coco128.yaml',
    weights='yolov5s.pt',
    imgsz=640,
    batch_size=64,          # Global batch size

    epochs=100,
    device='0,1,2,3',       # Triggers DDP when len(device) > 1

    sync_bn=True
)

```

When `device` contains multiple IDs, the wrapper automatically configures the distributed environment before calling the training loop.

### Manual DDP Spawn (Advanced)

For custom initialization logic, manually spawn processes and reuse YOLOv5's training logic:

```python
import os
import torch
import torch.distributed as dist
from train import parse_opt, main
from utils.torch_utils import select_device

def worker(local_rank, world_size):
    os.environ['LOCAL_RANK'] = str(local_rank)
    os.environ['RANK'] = str(local_rank)
    os.environ['WORLD_SIZE'] = str(world_size)
    
    dist.init_process_group(
        backend='nccl',
        init_method='env://'
    )
    
    opt = parse_opt()
    opt.device = f'{local_rank}'
    opt.batch_size = 64
    opt.sync_bn = True
    
    main(opt)

if __name__ == '__main__':
    world_size = torch.cuda.device_count()
    torch.multiprocessing.spawn(worker, args=(world_size,), nprocs=world_size, join=True)

```

This approach mirrors the internal behavior of `torch.distributed.run` while allowing custom pre-processing or environment configuration.

## Key Source Files

The multi-GPU implementation spans several critical files in the repository:

- **[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py)** – Orchestrates DDP initialization (**lines 71‑85**), validates batch sizes (**lines 78‑80**), handles `--sync-bn` conversion (**lines 82‑84**), and manages early stopping broadcasts (**lines 105‑110**)
- **[`utils/torch_utils.py`](https://github.com/ultralytics/yolov5/blob/main/utils/torch_utils.py)** – Contains `select_device()` (**lines 71‑76**) for GPU binding and `smart_DDP()` (**lines 94‑99**) for model wrapping
- **[`utils/dataloaders.py`](https://github.com/ultralytics/yolov5/blob/main/utils/dataloaders.py)** – Implements per-rank data sharding via `rank=LOCAL_RANK` parameter in `create_dataloader()`
- **[`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py)** – Defines the base `Model` class that gets wrapped by DDP

## Summary

- **Launch mechanism**: Use `python -m torch.distributed.run --nproc_per_node N` to spawn N processes, setting `RANK` and `LOCAL_RANK` automatically
- **Batch size constraint**: Global `--batch-size` must be divisible by the number of GPUs to ensure equal distribution via `batch_size // WORLD_SIZE`
- **NCCL backend**: YOLOv5 automatically selects NCCL for CUDA GPUs, falling back to Gloo only when necessary
- **Model wrapping**: `smart_DDP()` automatically converts models to `DistributedDataParallel` after device selection
- **SyncBatchNorm**: Enable `--sync-bn` when using small per-GPU batch sizes to maintain accurate batch statistics across all devices
- **Clean shutdown**: The master rank broadcasts the stop flag to ensure all processes exit together

## Frequently Asked Questions

### What happens if my batch size is not divisible by the number of GPUs?

The training script will raise an assertion error during initialization. According to **[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 78‑80**, YOLOv5 explicitly checks `assert batch_size % WORLD_SIZE == 0` because DDP requires each rank to process an equal number of samples per iteration. You must adjust your batch size to be a multiple of your GPU count (e.g., 64 for 4 GPUs, not 63).

### Can I use automatic batch size tuning (`--batch-size -1`) with multiple GPUs?

No. As implemented in **[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 78‑80**, the auto-batch feature is disabled when `LOCAL_RANK != -1` because DDP requires a fixed, explicit batch size to properly split the workload across processes. You must specify a concrete integer value for `--batch-size`.

### Why should I use `--sync-bn` during multi-GPU training?

**`--sync-bn`** converts standard `BatchNorm` layers to `SyncBatchNorm` (**[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 82‑84**), which aggregates mean and variance statistics across all GPUs. This is essential when the per-GPU batch size (global batch size divided by GPU count) is small (typically < 16), as it prevents unstable batch statistics and improves model convergence.

### How does YOLOv5 handle checkpoint saving in DDP mode?

Only the master process (rank 0) writes checkpoints to disk, preventing file corruption from multiple processes writing simultaneously. However, **[`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) lines 105‑110** show that critical control flags like the early stopping signal are broadcast to all ranks using `dist.broadcast_object_list`, ensuring consistent state across all GPUs even though only rank 0 performs I/O operations.