# Soup CLI GPU Detection and Utilization: A Deep Dive into the Multi-Backend Architecture

> Explore Soup CLI's multi-backend architecture for GPU detection and utilization. It unifies device info across CUDA, MPS, and MLX for optimized distributed training and memory allocation.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: deep-dive
- Published: 2026-09-06

---

**Soup CLI automatically detects available GPU hardware across CUDA, MPS, and MLX backends, normalizes device information into a unified schema, and dynamically configures distributed training strategies and memory allocation based on the detected topology.**

The `MakazhanAlpamys/Soup` repository implements a sophisticated three-layer **Soup CLI GPU detection** pipeline that bridges diverse hardware ecosystems—from NVIDIA data centers to Apple Silicon laptops—into a single coherent interface for machine learning workloads. Unlike simpler tools that rely solely on PyTorch's device queries, Soup CLI orchestrates backend selection, hardware capability enumeration, and topology-aware strategy recommendations to optimize training and inference performance automatically.

## How Soup CLI Detects Available GPU Backends

The detection process begins in [`src/soup_cli/utils/backend_detect.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/backend_detect.py), where the `detect_backend()` function performs a hierarchical capability check to determine the appropriate runtime.

The selection logic follows this priority order:

1. **CUDA inspection** – The system first checks `torch.cuda.is_available()` to determine if NVIDIA GPUs are present and accessible.
2. **Apple Metal fallback** – If CUDA is unavailable, the code evaluates `torch.backends.mps.is_available()` to detect Apple Silicon GPUs.
3. **MLX final fallback** – When running on Apple devices without standard PyTorch MPS support, the system falls back to the MLX backend via `torch.backends.mlx`.

This layered approach ensures that **Soup CLI GPU detection** remains robust across heterogeneous environments without requiring manual configuration. The function returns a standardized backend identifier that downstream modules use to select appropriate hardware queries.

## Unified GPU Information Gathering

Once the backend is established, [`src/soup_cli/utils/gpu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/gpu.py) centralizes all hardware introspection through the `get_gpu_info()` function. This module normalizes disparate GPU APIs into a consistent dictionary schema containing `gpu_count`, `gpu_name`, and `gpu_memory_total_bytes`.

The implementation delegates to backend-specific helpers:

- **`_cuda_gpu_info()`** – Invokes `torch.cuda.device_count()` and iterates through devices with `torch.cuda.get_device_properties(i)`, extracting `total_memory` and `name` attributes for each NVIDIA GPU.
- **`_mps_gpu_info()`** – Queries `torch.backends.mps.device_props` to retrieve memory statistics on Apple Silicon, treating unified memory as the available GPU pool.
- **`_mlx_gpu_info()`** – Calls `mlx.core.get_device()` to access `device.type` and `device.memory` attributes for MLX-compatible hardware.

This normalization allows commands like `soup train` to access hardware capabilities through a single interface regardless of whether the underlying system runs CUDA drivers or Metal Performance Shaders.

## Memory-Aware Configuration and Strategy Selection

With hardware data collected, [`src/soup_cli/utils/topology.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/topology.py) translates raw specifications into actionable deployment configurations. The `resolve_num_gpus()` function validates user requests against detected hardware, while `suggest_strategy()` maps model sizes and GPU counts to specific parallelization approaches.

### Batch Size Estimation

The `estimate_batch_size()` function in [`src/soup_cli/utils/gpu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/gpu.py) applies heuristics from scaling laws research to calculate safe memory allocation. It applies a 0.7 multiplier to the total GPU memory (leaving 30% for optimizer states and activations), converting `gpu_memory_total_bytes` into recommended batch sizes without risking out-of-memory errors during training.

### NCCL Environment Preparation

For multi-GPU configurations, `suggest_nccl_env()` automatically configures distributed communication parameters. When `gpu_count` exceeds one, the function sets critical environment variables including `NCCL_SOCKET_IFNAME`, `NCCL_IB_DISABLE`, and `NCCL_DEBUG` based on detected interconnect types (NVLink versus PCIe), optimizing collective communication performance without manual tuning.

## Command-Line Integration and Edge Case Handling

The detection pipeline integrates directly with CLI commands through flags like `--gpus N`, `--gpu-memory`, and `--bf16`. In [`src/soup_cli/commands/train.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/train.py), the `_hardware_fit_preflight()` function validates user-specified GPU counts against the detected topology, aborting with descriptive error messages if the request exceeds available hardware.

### Handling Hardware Variations

**No GPU environments** – When detection finds no accelerators, the system gracefully falls back to CPU training, with `detect_device()` returning `("cpu", "cpu")` and subsequent operations defaulting to `torch.device("cpu")`.

**Mixed-precision validation** – Functions `cuda_supports_bf16()` and `mps_supports_bf16()` in [`src/soup_cli/utils/advanced_precision.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/advanced_precision.py) guard against incompatible precision requests. If a detected GPU lacks native BF16 support, Soup CLI omits the flag and emits a warning rather than failing at runtime.

**Apple Silicon unified memory** – On M1/M2/M3 devices, `get_gpu_info()` reports the unified memory size as both system and GPU memory, while profiling utilities treat it as a single pool to prevent double-counting during memory tracking.

**Cloud deployment validation** – The `validate_gpu()` function in [`src/soup_cli/cloud/modal.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/cloud/modal.py) ensures requested GPU identifiers match known hardware profiles before provisioning cloud resources, preventing configuration errors in remote environments.

## Practical Implementation Examples

```python

# Display detected GPU capabilities

from soup_cli.utils.gpu import get_gpu_info

info = get_gpu_info()
print(f"Detected {info['gpu_count']} GPU(s): {info.get('gpu_name')}")
print(f"Total memory per GPU: {info['gpu_memory_total_bytes'] / (1024**3):.1f} GB")

```

```bash

# Train with specific GPU allocation and automatic strategy selection

soup train --config my.yaml --gpus 2 --bf16

# Internal workflow:

#   1. resolve_num_gpus(2) validates against get_gpu_info()['gpu_count']

#   2. suggest_strategy(2, model_size) selects data-parallel sharding

#   3. suggest_nccl_env(2, "nvlink") configures NCCL variables

```

```python

# Calculate safe batch size for current hardware

from soup_cli.utils.gpu import estimate_batch_size, get_gpu_info

gpu = get_gpu_info()
batch = estimate_batch_size(gpu["gpu_memory_total_bytes"])
print(f"Recommended batch size: {batch}")

```

## Summary

- **Soup CLI GPU detection** operates through three coordinated layers: backend selection ([`backend_detect.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/backend_detect.py)), information normalization ([`gpu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/gpu.py)), and topology-aware strategy recommendation ([`topology.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/topology.py)).
- The system supports **CUDA, MPS (Apple Silicon), and MLX backends** through a unified interface that abstracts hardware-specific APIs into consistent data structures.
- **Memory-aware heuristics** automatically calculate safe batch sizes while reserving headroom for optimizer states, preventing out-of-memory errors during training.
- **Edge case handling** includes graceful CPU fallback, BF16 capability validation, and unified memory accounting for Apple Silicon devices.
- Command-line integration via `--gpus` flags validates requests against detected hardware before execution, with automatic NCCL configuration for distributed training.

## Frequently Asked Questions

### How does Soup CLI handle machines with no GPU available?

When `detect_backend()` finds no CUDA, MPS, or MLX capabilities, Soup CLI routes execution to CPU-only mode. The system returns `("cpu", "cpu")` from device detection functions, and training commands automatically default to `torch.device("cpu")` without requiring user intervention or configuration changes.

### Can Soup CLI detect and utilize multiple different GPU models in the same machine?

According to the source code in [`src/soup_cli/utils/gpu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/gpu.py), the detection system enumerates all available devices through backend-specific APIs but consolidates information into a single metadata structure. While `get_gpu_info()` captures device properties individually, the current `suggest_strategy()` implementation in [`topology.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/topology.py) assumes homogeneous GPU configurations for parallel strategies, matching similar constraints in standard PyTorch distributed training.

### What happens if I request BF16 precision on a GPU that does not support it?

Soup CLI guards against unsupported precision through validation functions in [`src/soup_cli/utils/advanced_precision.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/advanced_precision.py). If `cuda_supports_bf16()` or `mps_supports_bf16()` determine the detected hardware lacks native support, the system automatically strips the `--bf16` flag from the configuration and emits a runtime warning, falling back to FP32 precision to prevent execution failures.

### How does Soup CLI calculate recommended batch sizes for different GPU memory configurations?

The `estimate_batch_size()` function applies the heuristic `gpu_memory_gb * 0.7` to determine safe allocation limits, based on scaling laws for large language models. This calculation reserves approximately 30% of available VRAM for optimizer states, gradients, and activation checkpoints, ensuring that the recommended batch size utilizes memory aggressively while avoiding out-of-memory crashes during training iterations.