# Setting up Multi-GPU Training with PyTorch DDP: A Complete Guide from LLMs-from-scratch

> Learn to set up multi-GPU training with PyTorch DDP. This guide simplifies scaling your LLM training across multiple GPUs with minimal code changes.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-12

---

**PyTorch DDP scales large language model training across multiple GPUs by replicating the model on each device and automatically synchronizing gradients after every backward pass, requiring only minimal changes to single-GPU code.**

Training transformer models from scratch quickly exhausts the memory of a single GPU. The `rasbt/LLMs-from-scratch` repository demonstrates how to scale the training pipeline using **Distributed Data Parallel (DDP)**, a PyTorch module that shards data across devices while keeping model replicas synchronized. Unlike the legacy `DataParallel`, DDP runs separate processes for each GPU, eliminating the GIL bottleneck and maximizing throughput.

## Why DDP is Essential for LLM Training

Single-GPU training becomes impractical as model parameters grow. In the LLMs-from-scratch project, the baseline script at [`ch05/10_llm-training-speed/01_opt_single_gpu.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05/10_llm-training-speed/01_opt_single_gpu.py) trains a GPT-style model on one device. Moving to multiple GPUs requires handling both **data partitioning** and **gradient synchronization**. DDP solves this by launching independent processes—one per GPU—that each hold a full model copy. After every backward pass, DDP performs an all-reduce operation to average gradients across the process group, ensuring identical model updates without manual intervention. This design minimizes inter-GPU communication overhead by transferring only gradient tensors rather than full parameter sets.

## Core Implementation Files in LLMs-from-scratch

The repository provides two reference implementations for distributed training:

- **[`appendix-A/01_main-chapter-code/DDP-script.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/appendix-A/01_main-chapter-code/DDP-script.py)** – A minimal example showing manual process-group initialization and Python multiprocessing spawning.
- **[`ch05/10_llm-training-speed/02_opt_multi_gpu_ddp.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05/10_llm-training-speed/02_opt_multi_gpu_ddp.py)** – A production-ready script mirroring the single-GPU training logic but adapted for DDP.
- **[`ch04/01_main-chapter-code/gpt.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/01_main-chapter-code/gpt.py)** – Contains the `GPT` model definition used across all training configurations.

These files demonstrate that switching to multi-GPU training requires only three additional components: process-group setup, a `DistributedSampler`, and the `DistributedDataParallel` wrapper.

## Step-by-Step DDP Setup

The complete workflow involves initializing the distributed environment, partitioning the dataset, wrapping the model, and executing the training loop with proper cleanup.

### Initialize the Process Group

Each DDP process requires a unique **rank** (identifier) and knowledge of the **world_size** (total GPU count). The initialization sets communication backend (typically `"nccl"` for NVIDIA GPUs) and master-node coordinates:

```python
import os
import torch.distributed as dist

def setup(rank, world_size):
    os.environ["MASTER_ADDR"] = "127.0.0.1"
    os.environ["MASTER_PORT"] = "29500"
    dist.init_process_group("nccl", rank=rank, world_size=world_size)

def cleanup():
    dist.destroy_process_group()

```

The `setup()` function must be called at the start of every training process, establishing the all-reduce communication channel that DDP uses to synchronize gradients.

### Prepare Distributed Data Sampling

Standard data loaders send identical batches to every process. To shard the dataset across GPUs, use `DistributedSampler` from `torch.utils.data.distributed`:

```python
from torch.utils.data import DataLoader, DistributedSampler

sampler = DistributedSampler(
    dataset, 
    num_replicas=world_size, 
    rank=rank, 
    shuffle=True
)
loader = DataLoader(dataset, batch_size=8, sampler=sampler)

```

Critical for proper shuffling: call `sampler.set_epoch(epoch)` at the beginning of each training epoch to ensure different data ordering across runs.

### Wrap the Model with DDP

After moving the model to the target device, wrap it with `DistributedDataParallel`. This module handles gradient averaging automatically during `loss.backward()`:

```python
from torch.nn.parallel import DistributedDataParallel as DDP
from gpt import GPT  # From ch04/01_main-chapter-code/gpt.py

model = GPT().to(rank)
ddp_model = DDP(model, device_ids=[rank])
optimizer = torch.optim.AdamW(ddp_model.parameters(), lr=6e-4)

```

The `device_ids` parameter pins the model to the specific GPU corresponding to the process rank.

### The Training Loop

The training loop remains nearly identical to single-GPU code. DDP intercepts the backward pass to synchronize gradients, so no manual averaging is required:

```python
for epoch in range(5):
    sampler.set_epoch(epoch)  # Essential for proper shuffling

    for batch in loader:
        optimizer.zero_grad()
        loss = ddp_model(batch).loss
        loss.backward()
        optimizer.step()
    if rank == 0:
        print(f"Epoch {epoch} completed")
cleanup()

```

## Launching Multi-GPU Jobs

The repository provides two methods to spawn distributed processes: the modern `torchrun` utility and manual Python launching.

### Launch with torchrun

The recommended approach uses `torchrun` (included with PyTorch ≥ 2.0) to handle process spawning and environment variable configuration automatically:

```bash
torchrun --standalone --nproc_per_node=4 appendix-A/01_main-chapter-code/DDP-script-torchrun.py

```

The `--standalone` flag configures a local process group without requiring a separate master node. The `--nproc_per_node` argument should match the number of available GPUs.

### Manual Python Launching

For environments where `torchrun` is unavailable, the repository includes [`DDP-script.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/DDP-script.py), which uses `torch.multiprocessing.spawn` to create processes manually:

```python
if __name__ == "__main__":
    world_size = torch.cuda.device_count()
    torch.multiprocessing.spawn(
        train,
        args=(world_size, dataset),
        nprocs=world_size,
        join=True
    )

```

Launch this version with:

```bash
python appendix-A/01_main-chapter-code/DDP-script.py

```

## Verifying Correct Distributed Behavior

To confirm DDP is functioning correctly, inspect the printed loss values from each rank. In proper operation, losses should remain nearly identical across all GPUs after the first few iterations, confirming successful gradient synchronization. GPU utilization should show near-identical activity patterns across devices, with no single GPU bottlenecking the others.

## Summary

- **DDP requires minimal code changes**: Add process-group initialization in `setup()`, wrap the model with `DDP()`, and substitute `DistributedSampler` for the standard sampler.
- **File locations matter**: Reference implementations live in [`appendix-A/01_main-chapter-code/DDP-script.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/appendix-A/01_main-chapter-code/DDP-script.py) and [`ch05/10_llm-training-speed/02_opt_multi_gpu_ddp.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05/10_llm-training-speed/02_opt_multi_gpu_ddp.py).
- **Launch methods**: Use `torchrun --standalone --nproc_per_node=N` for simplicity, or `torch.multiprocessing.spawn` for custom cluster configurations.
- **Performance gains**: DDP provides near-linear scaling on a single node because only gradients (not parameters) are communicated during the all-reduce step.

## Frequently Asked Questions

### What is the difference between DDP and DataParallel?

**DDP launches separate Python processes for each GPU**, avoiding the Global Interpreter Lock (GIL) contention that plagues `DataParallel`. While `DataParallel` runs in a single process with threads and requires constant parameter scattering and gathering, DDP replicates the model once per process and synchronizes only gradients via all-reduce operations, resulting in significantly better scaling efficiency.

### Do I need to change my model architecture to use DDP?

**No, the model definition remains unchanged.** The `GPT` class in [`ch04/01_main-chapter-code/gpt.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/01_main-chapter-code/gpt.py) works identically for both single-GPU and multi-GPU training. You only modify the training script to wrap the instantiated model with `DistributedDataParallel` and handle the distributed sampler.

### How do I scale DDP across multiple nodes?

**Modify the `MASTER_ADDR` environment variable to point to the coordinator node's IP** and ensure all nodes use the same `MASTER_PORT`. Set the `world_size` to the total GPU count across all nodes (GPUs per node × number of nodes), and assign a unique global rank to each process. The `init_process_group` call handles the rest of the multi-node communication topology automatically.

### Why are my losses different across GPUs in DDP?

**Unequal losses typically indicate improper sampler configuration.** Ensure you call `sampler.set_epoch(epoch)` before each epoch so that `DistributedSampler` shuffles data differently per epoch. Additionally, verify that all processes use the identical random seed initialization for model weights; any asymmetry in initial parameters or data ordering causes divergent training trajectories.