How Faceswap Handles Distributed Training Across Multiple GPUs

Faceswap automatically detects multiple GPUs and switches to distributed training mode using PyTorch's torch.nn.DataParallel, automatically adjusting batch sizes and wrapping Keras models for parallel execution.

The deepfakes/faceswap repository implements a seamless multi-GPU training pipeline that abstracts away the complexity of distributed computing. When you launch a training session, the framework detects your hardware configuration and dynamically selects between single-GPU and distributed training modes without requiring manual configuration changes.

Automatic GPU Detection and Trainer Selection

Faceswap determines the training strategy at runtime based on GPU availability. In scripts/train.py, the entry point checks torch.cuda.device_count() to count available GPUs before selecting the appropriate trainer class.

If the system detects more than one GPU, Faceswap instantiates the distributed trainer from plugins/train/trainer/distributed.py. For single-GPU or CPU-only environments, it falls back to the OriginalTrainer class in plugins/train/trainer/original.py. This automatic selection ensures that users do not need to modify code or configuration files to take advantage of multiple GPUs.

The detection occurs at line 99 in distributed.py, where self._gpu_count = torch.cuda.device_count() captures the number of available CUDA devices.

Batch Size Validation for Multi-GPU Setups

Before initializing the distributed model, Faceswap validates that the requested batch size can be efficiently distributed across all detected GPUs.

Ensuring Divisibility Across GPUs

The _validate_batch_size method (lines 20-32 in distributed.py) enforces two critical constraints:

  • The batch size must be at least equal to the number of GPUs
  • The batch size must be evenly divisible by the GPU count

If your requested batch size violates either constraint, Faceswap emits a warning and automatically adjusts the value to the nearest valid configuration. For example, with 4 GPUs and a requested batch size of 30, the framework would adjust to 32 or 28 to ensure equal distribution across devices.

Model Wrapping for Distributed Execution

Faceswap bridges the gap between its Keras-based model architecture and PyTorch's distributed training capabilities through a specialized wrapping mechanism.

Wrapping Keras Models for PyTorch

In _set_distributed (lines 70-76), the framework performs a two-step wrapping process:

  1. Inner Wrapping: The Keras model is first wrapped in a custom WrappedModel class that inherits from torch.nn.Module. This adapter provides a single-input forward method compatible with PyTorch's expectations.

  2. Parallel Wrapping: The WrappedModel is then passed to torch.nn.DataParallel, which handles the distribution of inputs across multiple GPUs and the aggregation of gradients during backpropagation.

This architecture allows Faceswap to maintain its original Keras model definitions (located in tools/model/model.py) while leveraging PyTorch's mature multi-GPU infrastructure.

Distributed Forward Pass and Loss Calculation

During each training iteration, the _forward method (lines 111-118) coordinates the distributed computation. The DataParallel wrapper automatically splits input tensors across available GPUs, executes the forward pass in parallel, and collects the results.

The loss computation follows this workflow:

  1. Each GPU calculates losses independently for its subset of the batch
  2. The losses are summed across all devices
  3. The total is averaged by the GPU count to maintain consistent gradient magnitudes regardless of the number of devices

Additionally, the _handle_torch_gpu_mismatch_warning method (lines 44-58) intercepts PyTorch's internal warnings about GPU utilization imbalances, cleaning them up and presenting user-friendly notifications about potential performance issues.

Running Distributed Training

To initiate distributed training, simply run the standard training command on a multi-GPU system:

python scripts/train.py \
    -i /path/to/training_data \
    -o /path/to/model_output \
    -m plugins.train.model.original \
    -b 64

The framework automatically handles the rest. Under the hood, the execution flow follows this pattern:


# Logic inside scripts/train.py

import torch
from plugins.train.trainer import distributed, original

if torch.cuda.device_count() > 1:
    trainer = distributed.Trainer(model, batch_size=batch_size)
else:
    trainer = original.Trainer(model, batch_size=batch_size)

trainer.train()

When the distributed trainer initializes, it executes the full pipeline: validating the batch size, wrapping the model for PyTorch compatibility, and configuring DataParallel for the detected GPU count.

Summary

  • Automatic Detection: Faceswap counts available GPUs using torch.cuda.device_count() and selects the appropriate trainer without user intervention.
  • Batch Validation: The _validate_batch_size function ensures batch sizes are divisible by GPU count and emits warnings for invalid configurations.
  • Model Adaptation: Keras models are wrapped in WrappedModel (a PyTorch Module) and then in torch.nn.DataParallel to enable distributed execution.
  • Loss Aggregation: The _forward method averages losses across all GPUs to maintain training stability.
  • Zero Configuration: Users simply run scripts/train.py and the framework handles GPU distribution automatically.

Frequently Asked Questions

Does Faceswap support multi-GPU training out of the box?

Yes. As implemented in the deepfakes/faceswap repository, the training pipeline automatically detects multiple GPUs via torch.cuda.device_count() and switches to the distributed trainer without requiring configuration changes. You only need to ensure PyTorch is installed with CUDA support.

What happens if my batch size isn't divisible by the number of GPUs?

The _validate_batch_size method in plugins/train/trainer/distributed.py detects this condition and automatically adjusts your batch size to the nearest valid value that is evenly divisible by the GPU count. It also emits a warning informing you of the adjustment.

Does Faceswap use PyTorch DistributedDataParallel or DataParallel?

Faceswap uses torch.nn.DataParallel rather than DistributedDataParallel. The framework wraps Keras models in a PyTorch-compatible WrappedModel class and then applies DataParallel in the _set_distributed method (lines 70-76) to handle multi-GPU training within a single process.

Can I force single-GPU training even if multiple GPUs are available?

The current implementation in scripts/train.py automatically selects the distributed trainer when multiple GPUs are detected. To force single-GPU training, you would need to modify the selection logic or set CUDA_VISIBLE_DEVICES to expose only one GPU to the process before running the training script.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →