GPU Detection and Monitoring in Faceswap: A Complete Technical Guide

Faceswap provides a backend-agnostic GPU detection and monitoring API that dynamically supports Nvidia CUDA, AMD ROCm, Apple Silicon, and CPU-only environments through a unified GPUStats interface.

The deepfakes/faceswap repository implements a sophisticated hardware abstraction layer in lib/gpu_stats that enables seamless GPU detection and monitoring across diverse deep learning backends. This system automatically detects the available hardware platform and exposes consistent VRAM statistics, driver information, and device enumeration regardless of whether you are running on Nvidia RTX cards, AMD GPUs, Apple M-series chips, or CPU-only infrastructure.

Architecture Overview

The GPU detection and monitoring system centers on the GPUStats class, which serves as the primary entry point located in lib/gpu_stats/__init__.py. This module dynamically imports the appropriate backend implementation—nvidia.py, apple_silicon.py, rocm.py, or cpu.py—based on the active deep learning framework detected at runtime. All backend implementations inherit from the abstract _GPUStats class defined in lib/gpu_stats/_base.py, ensuring a consistent API surface for device queries and memory statistics.

Core Features and Capabilities

Backend Selection and Abstraction

The system initializes by calling get_backend() to determine whether the environment uses CUDA, ROCm, Metal, or CPU-only mode. This selection logic in lib/gpu_stats/__init__.py returns the concrete implementation class without requiring user intervention. Each backend handles driver-specific APIs internally while exposing uniform methods for cross-platform compatibility.

Device Enumeration and Identification

The _GPUStats base class provides device_count and cli_devices properties that return the total number of detected GPUs and human-readable device names respectively. Backend implementations override _get_device_names() to query hardware-specific APIs—such as pynvml for Nvidia or torch.backends.mps for Apple Silicon—and return model identifiers like "GeForce RTX 3080" or "Apple M1 Max".

VRAM Monitoring and Statistics

Real-time memory metrics are captured through _get_vram() and _get_free_vram() methods, which populate the GPUInfo dataclass with total and available memory in megabytes. The GPUInfo container—defined in lib/gpu_stats/_base.py—exposes fields for vram, vram_free, driver, devices, and devices_active, providing a snapshot of current GPU capacity across all detected hardware.

Active Device Handling and Exclusion

Faceswap respects environment variables like CUDA_VISIBLE_DEVICES while offering programmatic control through exclude_devices() and exclude_all_devices() methods. The _get_active_devices() method in _base.py filters the full device list against exclusion lists, allowing training scripts to hide specific GPUs from TensorFlow or PyTorch without restarting the process.

Backend-Specific Implementations

Nvidia GPUs (CUDA)

The nvidia.py implementation utilizes the pynvml library to query driver versions, device names, and precise VRAM statistics via the Nvidia Management Library. This backend provides the most comprehensive metrics, including free VRAM calculations and multi-GPU enumeration, making it the reference implementation for the abstraction layer.

Apple Silicon (Metal)

For M-series chips, apple_silicon.py leverages torch.backends.mps to detect Metal-capable devices. While it reports accurate total VRAM and driver information (labeled as "Apple"), the Metal API limitations prevent real-time free VRAM queries, returning only static capacity metrics.

AMD GPUs (ROCm)

The rocm.py module interfaces with ROCm runtime libraries to deliver equivalent functionality to the Nvidia backend, including driver version strings and VRAM statistics. This implementation ensures AMD GPU users receive identical device enumeration and memory monitoring capabilities within the Faceswap ecosystem.

CPU Fallback Mode

When no GPU libraries are available, cpu.py provides a compliant fallback that returns zero devices, "N/A" driver strings, and empty lists. This ensures the GPUStats API never raises import errors, allowing CPU-only training pipelines to execute without conditional branching in user code.

Practical Implementation Guide

Retrieving System Information

The simplest way to access GPU data is through the system information module, which automatically aggregates hardware details:

from lib.system.sysinfo import get_sysinfo

# Returns full system diagnostics including GPU details

print(get_sysinfo())

Direct GPU Inspection

For granular control, instantiate GPUStats directly and inspect the GPUInfo dataclass:

from lib.gpu_stats import GPUStats, GPUInfo

if GPUStats is not None:
    stats = GPUStats(log=False)  # Suppress logging during import

    info: GPUInfo = stats.sys_info
    
    print(f"Driver: {info.driver}")
    print(f"Devices: {', '.join(info.devices)}")
    print(f"VRAM (MiB): {', '.join(str(v) for v in info.vram)}")
    print(f"Free VRAM: {', '.join(str(f) for f in info.vram_free)}")
else:
    print("GPU detection library unavailable")

Excluding Specific GPUs

To hide devices from the training process before framework initialization:

from lib.gpu_stats import GPUStats

# Must be called before importing torch/keras

if GPUStats:
    GPUStats().exclude_devices([0])  # Hide GPU 0 from training process

Summary

  • Backend-agnostic design: The lib/gpu_stats module automatically selects Nvidia, AMD, Apple Silicon, or CPU implementations based on runtime environment detection.
  • Comprehensive metrics: The GPUInfo dataclass standardizes access to driver versions, device names, total VRAM, and free VRAM across all supported hardware platforms.
  • Dynamic device control: The exclude_devices() and exclude_all_devices() methods enable runtime GPU hiding without environment variable manipulation.
  • Integration ready: GPU statistics automatically populate SysInfo.full_info() in lib/system/sysinfo.py for unified logging and debugging output.

Frequently Asked Questions

How does Faceswap detect which GPU backend to use?

The detection logic resides in lib/gpu_stats/__init__.py, which calls get_backend() to identify whether the environment uses CUDA, ROCm, Metal, or CPU-only mode. Based on this check, it dynamically imports the corresponding implementation module (nvidia.py, rocm.py, apple_silicon.py, or cpu.py) and returns the appropriate GPUStats class.

Can I monitor free VRAM on Apple Silicon Macs?

No, the Apple Silicon backend in lib/gpu_stats/apple_silicon.py cannot report free VRAM due to Metal API limitations. While it accurately reports total VRAM capacity and device identification, real-time free memory statistics are only available on Nvidia and AMD ROCm platforms using pynvml and ROCm runtime libraries respectively.

How do I prevent Faceswap from using specific GPUs?

Call GPUStats().exclude_devices([device_id_list]) before importing deep learning frameworks like TensorFlow or PyTorch. This method updates the internal active device list used by Faceswap's training scripts. Alternatively, set the CUDA_VISIBLE_DEVICES environment variable before launch, which the _get_active_devices() method respects during initialization.

Where does Faceswap store GPU information for system reports?

The lib/system/sysinfo.py module consumes GPUStats through SysInfo.full_info(), embedding driver versions, device counts, and VRAM statistics into comprehensive system diagnostics. This integration ensures GPU data appears automatically in console output and log files without manual API calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →