How to Set Up YOLOv5 for Multi-GPU Training with PyTorch DDP

Launch YOLOv5 using torch.distributed.run with --nproc_per_node matching your GPU count, ensure the global batch size is divisible by the number of GPUs, and the repository automatically initializes NCCL process groups, wraps models with DistributedDataParallel, and synchronizes gradients across devices.

YOLOv5 utilizes PyTorch Distributed Data Parallel (DDP) to scale object detection training across multiple GPUs on a single machine or cluster node. According to the ultralytics/yolov5 source code, the implementation in train.py and utils/torch_utils.py handles process spawning, device selection, and inter-GPU communication transparently when you use the standard launch command. This article provides the exact configuration, launch commands, and internal architecture details required for efficient multi-GPU training.

Quick Start: Single-Node Multi-GPU Command

The standard method for YOLOv5 multi-GPU training uses the PyTorch elastic launcher to spawn one process per GPU. The command reference in train.py lines 9‑11 shows the recommended syntax:

python -m torch.distributed.run \
    --nproc_per_node 4 \
    --master_port 1 \
    train.py \
    --data data/coco128.yaml \
    --weights yolov5s.pt \
    --img 640 \
    --batch-size 64 \
    --device 0,1,2,3 \
    --sync-bn \
    --epochs 100

Critical requirements for this command:

  • --nproc_per_node must match the number of GPUs you specify in --device (e.g., 4 for GPUs 0,1,2,3)
  • --batch-size represents the global batch size and must be divisible by the number of GPUs (WORLD_SIZE)
  • --master_port can be any free port on your machine (default is 29500)

When executed, torch.distributed.run automatically sets the RANK, LOCAL_RANK, WORLD_SIZE, and MASTER_PORT environment variables required by the DDP backend.

Internal DDP Architecture

Understanding how YOLOv5 handles distributed training helps debug issues and optimize performance. The workflow follows a strict initialization sequence:

Process Initialization in main()

After parsing arguments, train.py lines 71‑85 check if DDP mode is active by detecting LOCAL_RANK != -1. If true, the script performs several validations:

  • Confirms the global batch size is divisible by WORLD_SIZE (train.py lines 78‑80)
  • Verifies sufficient CUDA devices are available
  • Initializes the process group using the NCCL backend (with Gloo fallback)

# Conceptual flow from train.py

if LOCAL_RANK != -1:
    assert torch.cuda.device_count() > LOCAL_RANK, 'insufficient CUDA devices'
    assert batch_size % WORLD_SIZE == 0, '--batch-size must be multiple of CUDA device count'
    torch.distributed.init_process_group(backend='nccl' if is_nccl_available() else 'gloo')

Device Selection with select_device()

The select_device() function in utils/torch_utils.py lines 71‑76 parses the --device argument and binds each process to its assigned GPU based on LOCAL_RANK. When DDP is active, the function ignores manual device lists and uses LOCAL_RANK to determine the specific CUDA device for the current process.

Model Wrapping with smart_DDP()

Once the model is constructed, smart_DDP() in utils/torch_utils.py lines 94‑99 wraps the base nn.Module with torch.nn.parallel.DistributedDataParallel. This wrapper handles gradient synchronization across processes automatically during the backward pass, ensuring consistent weight updates without manual all-reduce operations.

Data Loading and Batch Distribution

The global batch size specified via --batch-size is divided equally among ranks using integer division (batch_size // WORLD_SIZE). Each process calls create_dataloader() with rank=LOCAL_RANK to receive only its partition of the dataset. This sharding happens in utils/dataloaders.py, ensuring no data overlap between GPUs while maintaining the effective global batch size.

Synchronized Batch Normalization

When you include the --sync-bn flag, YOLOv5 converts all BatchNorm layers to SyncBatchNorm (train.py lines 82‑84). This operation uses torch.nn.SyncBatchNorm.convert_sync_batchnorm() to aggregate batch statistics across all GPUs during training, which improves convergence when the per-GPU batch size is small.

Early Stopping and Checkpoint Broadcasting

To ensure clean shutdowns, train.py lines 105‑110 implement a broadcast mechanism. When the master process (rank 0) triggers early stopping, it broadcasts the stop flag to all ranks using dist.broadcast_object_list, ensuring all processes exit the training loop simultaneously.

Alternative Launch Methods

While the command-line launcher is standard, YOLOv5 supports programmatic multi-GPU training.

Using the Python API

You can invoke training via the run() function, which internally handles DDP setup when multiple GPUs are detected:

from train import run

run(
    data='data/coco128.yaml',
    weights='yolov5s.pt',
    imgsz=640,
    batch_size=64,          # Global batch size

    epochs=100,
    device='0,1,2,3',       # Triggers DDP when len(device) > 1

    sync_bn=True
)

When device contains multiple IDs, the wrapper automatically configures the distributed environment before calling the training loop.

Manual DDP Spawn (Advanced)

For custom initialization logic, manually spawn processes and reuse YOLOv5's training logic:

import os
import torch
import torch.distributed as dist
from train import parse_opt, main
from utils.torch_utils import select_device

def worker(local_rank, world_size):
    os.environ['LOCAL_RANK'] = str(local_rank)
    os.environ['RANK'] = str(local_rank)
    os.environ['WORLD_SIZE'] = str(world_size)
    
    dist.init_process_group(
        backend='nccl',
        init_method='env://'
    )
    
    opt = parse_opt()
    opt.device = f'{local_rank}'
    opt.batch_size = 64
    opt.sync_bn = True
    
    main(opt)

if __name__ == '__main__':
    world_size = torch.cuda.device_count()
    torch.multiprocessing.spawn(worker, args=(world_size,), nprocs=world_size, join=True)

This approach mirrors the internal behavior of torch.distributed.run while allowing custom pre-processing or environment configuration.

Key Source Files

The multi-GPU implementation spans several critical files in the repository:

  • train.py – Orchestrates DDP initialization (lines 71‑85), validates batch sizes (lines 78‑80), handles --sync-bn conversion (lines 82‑84), and manages early stopping broadcasts (lines 105‑110)
  • utils/torch_utils.py – Contains select_device() (lines 71‑76) for GPU binding and smart_DDP() (lines 94‑99) for model wrapping
  • utils/dataloaders.py – Implements per-rank data sharding via rank=LOCAL_RANK parameter in create_dataloader()
  • models/yolo.py – Defines the base Model class that gets wrapped by DDP

Summary

  • Launch mechanism: Use python -m torch.distributed.run --nproc_per_node N to spawn N processes, setting RANK and LOCAL_RANK automatically
  • Batch size constraint: Global --batch-size must be divisible by the number of GPUs to ensure equal distribution via batch_size // WORLD_SIZE
  • NCCL backend: YOLOv5 automatically selects NCCL for CUDA GPUs, falling back to Gloo only when necessary
  • Model wrapping: smart_DDP() automatically converts models to DistributedDataParallel after device selection
  • SyncBatchNorm: Enable --sync-bn when using small per-GPU batch sizes to maintain accurate batch statistics across all devices
  • Clean shutdown: The master rank broadcasts the stop flag to ensure all processes exit together

Frequently Asked Questions

What happens if my batch size is not divisible by the number of GPUs?

The training script will raise an assertion error during initialization. According to train.py lines 78‑80, YOLOv5 explicitly checks assert batch_size % WORLD_SIZE == 0 because DDP requires each rank to process an equal number of samples per iteration. You must adjust your batch size to be a multiple of your GPU count (e.g., 64 for 4 GPUs, not 63).

Can I use automatic batch size tuning (--batch-size -1) with multiple GPUs?

No. As implemented in train.py lines 78‑80, the auto-batch feature is disabled when LOCAL_RANK != -1 because DDP requires a fixed, explicit batch size to properly split the workload across processes. You must specify a concrete integer value for --batch-size.

Why should I use --sync-bn during multi-GPU training?

--sync-bn converts standard BatchNorm layers to SyncBatchNorm (train.py lines 82‑84), which aggregates mean and variance statistics across all GPUs. This is essential when the per-GPU batch size (global batch size divided by GPU count) is small (typically < 16), as it prevents unstable batch statistics and improves model convergence.

How does YOLOv5 handle checkpoint saving in DDP mode?

Only the master process (rank 0) writes checkpoints to disk, preventing file corruption from multiple processes writing simultaneously. However, train.py lines 105‑110 show that critical control flags like the early stopping signal are broadcast to all ranks using dist.broadcast_object_list, ensuring consistent state across all GPUs even though only rank 0 performs I/O operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →