# How Faceswap Handles Distributed Training Across Multiple GPUs

> Discover how Faceswap leverages multiple GPUs with distributed training using PyTorch DataParallel. Automatically optimizes batch sizes for faster deepfake model training.

- Repository: [deepfakes/faceswap](https://github.com/deepfakes/faceswap)
- Tags: internals
- Published: 2026-03-06

---

**Faceswap automatically detects multiple GPUs and switches to distributed training mode using PyTorch's `torch.nn.DataParallel`, automatically adjusting batch sizes and wrapping Keras models for parallel execution.**

The deepfakes/faceswap repository implements a seamless multi-GPU training pipeline that abstracts away the complexity of distributed computing. When you launch a training session, the framework detects your hardware configuration and dynamically selects between single-GPU and distributed training modes without requiring manual configuration changes.

## Automatic GPU Detection and Trainer Selection

Faceswap determines the training strategy at runtime based on GPU availability. In [`scripts/train.py`](https://github.com/deepfakes/faceswap/blob/main/scripts/train.py), the entry point checks `torch.cuda.device_count()` to count available GPUs before selecting the appropriate trainer class.

If the system detects more than one GPU, Faceswap instantiates the distributed trainer from [`plugins/train/trainer/distributed.py`](https://github.com/deepfakes/faceswap/blob/main/plugins/train/trainer/distributed.py). For single-GPU or CPU-only environments, it falls back to the `OriginalTrainer` class in [`plugins/train/trainer/original.py`](https://github.com/deepfakes/faceswap/blob/main/plugins/train/trainer/original.py). This automatic selection ensures that users do not need to modify code or configuration files to take advantage of multiple GPUs.

The detection occurs at line 99 in [`distributed.py`](https://github.com/deepfakes/faceswap/blob/main/distributed.py), where `self._gpu_count = torch.cuda.device_count()` captures the number of available CUDA devices.

## Batch Size Validation for Multi-GPU Setups

Before initializing the distributed model, Faceswap validates that the requested batch size can be efficiently distributed across all detected GPUs.

### Ensuring Divisibility Across GPUs

The `_validate_batch_size` method (lines 20-32 in [`distributed.py`](https://github.com/deepfakes/faceswap/blob/main/distributed.py)) enforces two critical constraints:

- The batch size must be **at least equal to** the number of GPUs
- The batch size must be **evenly divisible** by the GPU count

If your requested batch size violates either constraint, Faceswap emits a warning and automatically adjusts the value to the nearest valid configuration. For example, with 4 GPUs and a requested batch size of 30, the framework would adjust to 32 or 28 to ensure equal distribution across devices.

## Model Wrapping for Distributed Execution

Faceswap bridges the gap between its Keras-based model architecture and PyTorch's distributed training capabilities through a specialized wrapping mechanism.

### Wrapping Keras Models for PyTorch

In `_set_distributed` (lines 70-76), the framework performs a two-step wrapping process:

1. **Inner Wrapping**: The Keras model is first wrapped in a custom `WrappedModel` class that inherits from `torch.nn.Module`. This adapter provides a single-input forward method compatible with PyTorch's expectations.

2. **Parallel Wrapping**: The `WrappedModel` is then passed to `torch.nn.DataParallel`, which handles the distribution of inputs across multiple GPUs and the aggregation of gradients during backpropagation.

This architecture allows Faceswap to maintain its original Keras model definitions (located in [`tools/model/model.py`](https://github.com/deepfakes/faceswap/blob/main/tools/model/model.py)) while leveraging PyTorch's mature multi-GPU infrastructure.

## Distributed Forward Pass and Loss Calculation

During each training iteration, the `_forward` method (lines 111-118) coordinates the distributed computation. The `DataParallel` wrapper automatically splits input tensors across available GPUs, executes the forward pass in parallel, and collects the results.

The loss computation follows this workflow:

1. Each GPU calculates losses independently for its subset of the batch
2. The losses are summed across all devices
3. The total is averaged by the GPU count to maintain consistent gradient magnitudes regardless of the number of devices

Additionally, the `_handle_torch_gpu_mismatch_warning` method (lines 44-58) intercepts PyTorch's internal warnings about GPU utilization imbalances, cleaning them up and presenting user-friendly notifications about potential performance issues.

## Running Distributed Training

To initiate distributed training, simply run the standard training command on a multi-GPU system:

```bash
python scripts/train.py \
    -i /path/to/training_data \
    -o /path/to/model_output \
    -m plugins.train.model.original \
    -b 64

```

The framework automatically handles the rest. Under the hood, the execution flow follows this pattern:

```python

# Logic inside scripts/train.py

import torch
from plugins.train.trainer import distributed, original

if torch.cuda.device_count() > 1:
    trainer = distributed.Trainer(model, batch_size=batch_size)
else:
    trainer = original.Trainer(model, batch_size=batch_size)

trainer.train()

```

When the distributed trainer initializes, it executes the full pipeline: validating the batch size, wrapping the model for PyTorch compatibility, and configuring `DataParallel` for the detected GPU count.

## Summary

- **Automatic Detection**: Faceswap counts available GPUs using `torch.cuda.device_count()` and selects the appropriate trainer without user intervention.
- **Batch Validation**: The `_validate_batch_size` function ensures batch sizes are divisible by GPU count and emits warnings for invalid configurations.
- **Model Adaptation**: Keras models are wrapped in `WrappedModel` (a PyTorch Module) and then in `torch.nn.DataParallel` to enable distributed execution.
- **Loss Aggregation**: The `_forward` method averages losses across all GPUs to maintain training stability.
- **Zero Configuration**: Users simply run [`scripts/train.py`](https://github.com/deepfakes/faceswap/blob/main/scripts/train.py) and the framework handles GPU distribution automatically.

## Frequently Asked Questions

### Does Faceswap support multi-GPU training out of the box?

Yes. As implemented in the deepfakes/faceswap repository, the training pipeline automatically detects multiple GPUs via `torch.cuda.device_count()` and switches to the distributed trainer without requiring configuration changes. You only need to ensure PyTorch is installed with CUDA support.

### What happens if my batch size isn't divisible by the number of GPUs?

The `_validate_batch_size` method in [`plugins/train/trainer/distributed.py`](https://github.com/deepfakes/faceswap/blob/main/plugins/train/trainer/distributed.py) detects this condition and automatically adjusts your batch size to the nearest valid value that is evenly divisible by the GPU count. It also emits a warning informing you of the adjustment.

### Does Faceswap use PyTorch DistributedDataParallel or DataParallel?

Faceswap uses `torch.nn.DataParallel` rather than `DistributedDataParallel`. The framework wraps Keras models in a PyTorch-compatible `WrappedModel` class and then applies `DataParallel` in the `_set_distributed` method (lines 70-76) to handle multi-GPU training within a single process.

### Can I force single-GPU training even if multiple GPUs are available?

The current implementation in [`scripts/train.py`](https://github.com/deepfakes/faceswap/blob/main/scripts/train.py) automatically selects the distributed trainer when multiple GPUs are detected. To force single-GPU training, you would need to modify the selection logic or set `CUDA_VISIBLE_DEVICES` to expose only one GPU to the process before running the training script.