Soup CLI GPU Detection and Utilization: A Deep Dive into the Multi-Backend Architecture

Soup CLI automatically detects available GPU hardware across CUDA, MPS, and MLX backends, normalizes device information into a unified schema, and dynamically configures distributed training strategies and memory allocation based on the detected topology.

The MakazhanAlpamys/Soup repository implements a sophisticated three-layer Soup CLI GPU detection pipeline that bridges diverse hardware ecosystems—from NVIDIA data centers to Apple Silicon laptops—into a single coherent interface for machine learning workloads. Unlike simpler tools that rely solely on PyTorch's device queries, Soup CLI orchestrates backend selection, hardware capability enumeration, and topology-aware strategy recommendations to optimize training and inference performance automatically.

How Soup CLI Detects Available GPU Backends

The detection process begins in src/soup_cli/utils/backend_detect.py, where the detect_backend() function performs a hierarchical capability check to determine the appropriate runtime.

The selection logic follows this priority order:

  1. CUDA inspection – The system first checks torch.cuda.is_available() to determine if NVIDIA GPUs are present and accessible.
  2. Apple Metal fallback – If CUDA is unavailable, the code evaluates torch.backends.mps.is_available() to detect Apple Silicon GPUs.
  3. MLX final fallback – When running on Apple devices without standard PyTorch MPS support, the system falls back to the MLX backend via torch.backends.mlx.

This layered approach ensures that Soup CLI GPU detection remains robust across heterogeneous environments without requiring manual configuration. The function returns a standardized backend identifier that downstream modules use to select appropriate hardware queries.

Unified GPU Information Gathering

Once the backend is established, src/soup_cli/utils/gpu.py centralizes all hardware introspection through the get_gpu_info() function. This module normalizes disparate GPU APIs into a consistent dictionary schema containing gpu_count, gpu_name, and gpu_memory_total_bytes.

The implementation delegates to backend-specific helpers:

  • _cuda_gpu_info() – Invokes torch.cuda.device_count() and iterates through devices with torch.cuda.get_device_properties(i), extracting total_memory and name attributes for each NVIDIA GPU.
  • _mps_gpu_info() – Queries torch.backends.mps.device_props to retrieve memory statistics on Apple Silicon, treating unified memory as the available GPU pool.
  • _mlx_gpu_info() – Calls mlx.core.get_device() to access device.type and device.memory attributes for MLX-compatible hardware.

This normalization allows commands like soup train to access hardware capabilities through a single interface regardless of whether the underlying system runs CUDA drivers or Metal Performance Shaders.

Memory-Aware Configuration and Strategy Selection

With hardware data collected, src/soup_cli/utils/topology.py translates raw specifications into actionable deployment configurations. The resolve_num_gpus() function validates user requests against detected hardware, while suggest_strategy() maps model sizes and GPU counts to specific parallelization approaches.

Batch Size Estimation

The estimate_batch_size() function in src/soup_cli/utils/gpu.py applies heuristics from scaling laws research to calculate safe memory allocation. It applies a 0.7 multiplier to the total GPU memory (leaving 30% for optimizer states and activations), converting gpu_memory_total_bytes into recommended batch sizes without risking out-of-memory errors during training.

NCCL Environment Preparation

For multi-GPU configurations, suggest_nccl_env() automatically configures distributed communication parameters. When gpu_count exceeds one, the function sets critical environment variables including NCCL_SOCKET_IFNAME, NCCL_IB_DISABLE, and NCCL_DEBUG based on detected interconnect types (NVLink versus PCIe), optimizing collective communication performance without manual tuning.

Command-Line Integration and Edge Case Handling

The detection pipeline integrates directly with CLI commands through flags like --gpus N, --gpu-memory, and --bf16. In src/soup_cli/commands/train.py, the _hardware_fit_preflight() function validates user-specified GPU counts against the detected topology, aborting with descriptive error messages if the request exceeds available hardware.

Handling Hardware Variations

No GPU environments – When detection finds no accelerators, the system gracefully falls back to CPU training, with detect_device() returning ("cpu", "cpu") and subsequent operations defaulting to torch.device("cpu").

Mixed-precision validation – Functions cuda_supports_bf16() and mps_supports_bf16() in src/soup_cli/utils/advanced_precision.py guard against incompatible precision requests. If a detected GPU lacks native BF16 support, Soup CLI omits the flag and emits a warning rather than failing at runtime.

Apple Silicon unified memory – On M1/M2/M3 devices, get_gpu_info() reports the unified memory size as both system and GPU memory, while profiling utilities treat it as a single pool to prevent double-counting during memory tracking.

Cloud deployment validation – The validate_gpu() function in src/soup_cli/cloud/modal.py ensures requested GPU identifiers match known hardware profiles before provisioning cloud resources, preventing configuration errors in remote environments.

Practical Implementation Examples


# Display detected GPU capabilities

from soup_cli.utils.gpu import get_gpu_info

info = get_gpu_info()
print(f"Detected {info['gpu_count']} GPU(s): {info.get('gpu_name')}")
print(f"Total memory per GPU: {info['gpu_memory_total_bytes'] / (1024**3):.1f} GB")

# Train with specific GPU allocation and automatic strategy selection

soup train --config my.yaml --gpus 2 --bf16

# Internal workflow:

#   1. resolve_num_gpus(2) validates against get_gpu_info()['gpu_count']

#   2. suggest_strategy(2, model_size) selects data-parallel sharding

#   3. suggest_nccl_env(2, "nvlink") configures NCCL variables

# Calculate safe batch size for current hardware

from soup_cli.utils.gpu import estimate_batch_size, get_gpu_info

gpu = get_gpu_info()
batch = estimate_batch_size(gpu["gpu_memory_total_bytes"])
print(f"Recommended batch size: {batch}")

Summary

  • Soup CLI GPU detection operates through three coordinated layers: backend selection (backend_detect.py), information normalization (gpu.py), and topology-aware strategy recommendation (topology.py).
  • The system supports CUDA, MPS (Apple Silicon), and MLX backends through a unified interface that abstracts hardware-specific APIs into consistent data structures.
  • Memory-aware heuristics automatically calculate safe batch sizes while reserving headroom for optimizer states, preventing out-of-memory errors during training.
  • Edge case handling includes graceful CPU fallback, BF16 capability validation, and unified memory accounting for Apple Silicon devices.
  • Command-line integration via --gpus flags validates requests against detected hardware before execution, with automatic NCCL configuration for distributed training.

Frequently Asked Questions

How does Soup CLI handle machines with no GPU available?

When detect_backend() finds no CUDA, MPS, or MLX capabilities, Soup CLI routes execution to CPU-only mode. The system returns ("cpu", "cpu") from device detection functions, and training commands automatically default to torch.device("cpu") without requiring user intervention or configuration changes.

Can Soup CLI detect and utilize multiple different GPU models in the same machine?

According to the source code in src/soup_cli/utils/gpu.py, the detection system enumerates all available devices through backend-specific APIs but consolidates information into a single metadata structure. While get_gpu_info() captures device properties individually, the current suggest_strategy() implementation in topology.py assumes homogeneous GPU configurations for parallel strategies, matching similar constraints in standard PyTorch distributed training.

What happens if I request BF16 precision on a GPU that does not support it?

Soup CLI guards against unsupported precision through validation functions in src/soup_cli/utils/advanced_precision.py. If cuda_supports_bf16() or mps_supports_bf16() determine the detected hardware lacks native support, the system automatically strips the --bf16 flag from the configuration and emits a runtime warning, falling back to FP32 precision to prevent execution failures.

The estimate_batch_size() function applies the heuristic gpu_memory_gb * 0.7 to determine safe allocation limits, based on scaling laws for large language models. This calculation reserves approximately 30% of available VRAM for optimizer states, gradients, and activation checkpoints, ensuring that the recommended batch size utilizes memory aggressively while avoiding out-of-memory crashes during training iterations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →