Setting up Multi-GPU Training with PyTorch DDP: A Complete Guide from LLMs-from-scratch

PyTorch DDP scales large language model training across multiple GPUs by replicating the model on each device and automatically synchronizing gradients after every backward pass, requiring only minimal changes to single-GPU code.

Training transformer models from scratch quickly exhausts the memory of a single GPU. The rasbt/LLMs-from-scratch repository demonstrates how to scale the training pipeline using Distributed Data Parallel (DDP), a PyTorch module that shards data across devices while keeping model replicas synchronized. Unlike the legacy DataParallel, DDP runs separate processes for each GPU, eliminating the GIL bottleneck and maximizing throughput.

Why DDP is Essential for LLM Training

Single-GPU training becomes impractical as model parameters grow. In the LLMs-from-scratch project, the baseline script at ch05/10_llm-training-speed/01_opt_single_gpu.py trains a GPT-style model on one device. Moving to multiple GPUs requires handling both data partitioning and gradient synchronization. DDP solves this by launching independent processes—one per GPU—that each hold a full model copy. After every backward pass, DDP performs an all-reduce operation to average gradients across the process group, ensuring identical model updates without manual intervention. This design minimizes inter-GPU communication overhead by transferring only gradient tensors rather than full parameter sets.

Core Implementation Files in LLMs-from-scratch

The repository provides two reference implementations for distributed training:

These files demonstrate that switching to multi-GPU training requires only three additional components: process-group setup, a DistributedSampler, and the DistributedDataParallel wrapper.

Step-by-Step DDP Setup

The complete workflow involves initializing the distributed environment, partitioning the dataset, wrapping the model, and executing the training loop with proper cleanup.

Initialize the Process Group

Each DDP process requires a unique rank (identifier) and knowledge of the world_size (total GPU count). The initialization sets communication backend (typically "nccl" for NVIDIA GPUs) and master-node coordinates:

import os
import torch.distributed as dist

def setup(rank, world_size):
    os.environ["MASTER_ADDR"] = "127.0.0.1"
    os.environ["MASTER_PORT"] = "29500"
    dist.init_process_group("nccl", rank=rank, world_size=world_size)

def cleanup():
    dist.destroy_process_group()

The setup() function must be called at the start of every training process, establishing the all-reduce communication channel that DDP uses to synchronize gradients.

Prepare Distributed Data Sampling

Standard data loaders send identical batches to every process. To shard the dataset across GPUs, use DistributedSampler from torch.utils.data.distributed:

from torch.utils.data import DataLoader, DistributedSampler

sampler = DistributedSampler(
    dataset, 
    num_replicas=world_size, 
    rank=rank, 
    shuffle=True
)
loader = DataLoader(dataset, batch_size=8, sampler=sampler)

Critical for proper shuffling: call sampler.set_epoch(epoch) at the beginning of each training epoch to ensure different data ordering across runs.

Wrap the Model with DDP

After moving the model to the target device, wrap it with DistributedDataParallel. This module handles gradient averaging automatically during loss.backward():

from torch.nn.parallel import DistributedDataParallel as DDP
from gpt import GPT  # From ch04/01_main-chapter-code/gpt.py

model = GPT().to(rank)
ddp_model = DDP(model, device_ids=[rank])
optimizer = torch.optim.AdamW(ddp_model.parameters(), lr=6e-4)

The device_ids parameter pins the model to the specific GPU corresponding to the process rank.

The Training Loop

The training loop remains nearly identical to single-GPU code. DDP intercepts the backward pass to synchronize gradients, so no manual averaging is required:

for epoch in range(5):
    sampler.set_epoch(epoch)  # Essential for proper shuffling

    for batch in loader:
        optimizer.zero_grad()
        loss = ddp_model(batch).loss
        loss.backward()
        optimizer.step()
    if rank == 0:
        print(f"Epoch {epoch} completed")
cleanup()

Launching Multi-GPU Jobs

The repository provides two methods to spawn distributed processes: the modern torchrun utility and manual Python launching.

Launch with torchrun

The recommended approach uses torchrun (included with PyTorch ≥ 2.0) to handle process spawning and environment variable configuration automatically:

torchrun --standalone --nproc_per_node=4 appendix-A/01_main-chapter-code/DDP-script-torchrun.py

The --standalone flag configures a local process group without requiring a separate master node. The --nproc_per_node argument should match the number of available GPUs.

Manual Python Launching

For environments where torchrun is unavailable, the repository includes DDP-script.py, which uses torch.multiprocessing.spawn to create processes manually:

if __name__ == "__main__":
    world_size = torch.cuda.device_count()
    torch.multiprocessing.spawn(
        train,
        args=(world_size, dataset),
        nprocs=world_size,
        join=True
    )

Launch this version with:

python appendix-A/01_main-chapter-code/DDP-script.py

Verifying Correct Distributed Behavior

To confirm DDP is functioning correctly, inspect the printed loss values from each rank. In proper operation, losses should remain nearly identical across all GPUs after the first few iterations, confirming successful gradient synchronization. GPU utilization should show near-identical activity patterns across devices, with no single GPU bottlenecking the others.

Summary

  • DDP requires minimal code changes: Add process-group initialization in setup(), wrap the model with DDP(), and substitute DistributedSampler for the standard sampler.
  • File locations matter: Reference implementations live in appendix-A/01_main-chapter-code/DDP-script.py and ch05/10_llm-training-speed/02_opt_multi_gpu_ddp.py.
  • Launch methods: Use torchrun --standalone --nproc_per_node=N for simplicity, or torch.multiprocessing.spawn for custom cluster configurations.
  • Performance gains: DDP provides near-linear scaling on a single node because only gradients (not parameters) are communicated during the all-reduce step.

Frequently Asked Questions

What is the difference between DDP and DataParallel?

DDP launches separate Python processes for each GPU, avoiding the Global Interpreter Lock (GIL) contention that plagues DataParallel. While DataParallel runs in a single process with threads and requires constant parameter scattering and gathering, DDP replicates the model once per process and synchronizes only gradients via all-reduce operations, resulting in significantly better scaling efficiency.

Do I need to change my model architecture to use DDP?

No, the model definition remains unchanged. The GPT class in ch04/01_main-chapter-code/gpt.py works identically for both single-GPU and multi-GPU training. You only modify the training script to wrap the instantiated model with DistributedDataParallel and handle the distributed sampler.

How do I scale DDP across multiple nodes?

Modify the MASTER_ADDR environment variable to point to the coordinator node's IP and ensure all nodes use the same MASTER_PORT. Set the world_size to the total GPU count across all nodes (GPUs per node × number of nodes), and assign a unique global rank to each process. The init_process_group call handles the rest of the multi-node communication topology automatically.

Why are my losses different across GPUs in DDP?

Unequal losses typically indicate improper sampler configuration. Ensure you call sampler.set_epoch(epoch) before each epoch so that DistributedSampler shuffles data differently per epoch. Additionally, verify that all processes use the identical random seed initialization for model weights; any asymmetry in initial parameters or data ordering causes divergent training trajectories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →