# How to Configure --gpu-devices and --gpu-vram for Multi-GPU Placement in DS4

> Master DS4 multi-GPU placement by configuring --gpu-devices and --gpu-vram. Learn to select CUDA GPUs and set memory budgets for efficient distributed training.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-09

---

**TLDR:** DS4 uses `--gpu-devices` to select specific CUDA GPU indices and `--gpu-vram` to define per-device memory budgets (in GiB) or enable automatic detection via `auto`, with both flags parsed and validated in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) to ensure matching device counts and valid allocation strategies.

DS4 is a high-performance local LLM inference engine developed by antirez. When deploying large language models across multiple NVIDIA GPUs, precise control over **multi-GPU placement** is essential for maximizing throughput and avoiding out-of-memory errors. This article breaks down the exact syntax, validation logic, and runtime behavior of the `--gpu-devices` and `--gpu-vram` flags based on the source code implementation.

## Understanding the Two Key Flags

### --gpu-devices: Selecting CUDA Device Indices

The `--gpu-devices` flag accepts a comma-separated list of CUDA device indices (e.g., `0,2,4,6`) that DS4 is permitted to use. The order you specify matters: DS4 assigns model layers to devices in the sequence provided, making this the primary mechanism for controlling **multi-GPU placement** topology.

### --gpu-vram: Memory Allocation Modes

The `--gpu-vram` flag controls how much video RAM is allocated per device. It accepts three distinct input types:

- **`auto`**: Probe the system for available free VRAM on the selected devices.
- **`0`**: Disable CUDA entirely and run in CPU-only mode.
- **`N,N,...`**: Explicit GiB values for each device (must match `--gpu-devices` length).

## Parsing Logic and Validation in ds4_gpu_args.c

The central parsing routine `parse_gpu_vram_arg` in **[`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c)** handles all four configuration scenarios:

**Case 1: Implicit Auto-Detection**

When you supply `--gpu-devices` without `--gpu-vram`, the parser forces `vram_arg` to `"auto"` (lines 29-37). DS4 then probes only the specified devices for available memory.

**Case 2: CPU-Only Mode (`--gpu-vram 0`)**

Setting `--gpu-vram 0` disables GPU acceleration completely. The code explicitly rejects simultaneous use of `--gpu-devices` in this mode (lines 40-47), as filtering devices is meaningless when CUDA is disabled.

**Case 3: Automatic VRAM Probing (`--gpu-vram auto`)**

In CUDA builds (gated by `#if !defined(DS4_NO_GPU)`), DS4 calls `ds4_gpu_args_probe_auto_cuda` from **`ds4_cuda.cu`** (lines 51-59) to detect free memory on each selected device. If `--gpu-devices` is provided, the probe is limited to that subset.

**Case 4: Explicit Budgets with Strict Validation**

When providing explicit GiB values (e.g., `40,12`), the parser uses `parse_csv_int_list` to tokenize the input. If both flags are present, the code enforces that the list lengths match exactly, emitting the error `"--gpu-devices count … does not match --gpu-vram count …"` if they differ (lines 84-88). The resulting mapping is stored in the **`ds4_gpu_config`** structure (lines 90-95) as `device_indices[i] → vram_bytes[i]`.

## How Layer Placement Works Internally

DS4's runtime uses the validated configuration to distribute model layers across GPUs:

**Sequential Placement**

In standard multi-GPU mode (without tensor parallelism), consecutive layer ranges are assigned to devices in the order specified by `--gpu-devices`. Placing faster GPUs earlier in the list can improve pipeline efficiency.

**Tensor Parallelism Pairing**

When `--cuda-tensor-parallel` is enabled, devices are grouped into pairs: the first two devices split layers 50/50, the next two form another pair, and so on. The `--gpu-devices` flag still controls which physical GPUs participate, but the runtime automatically manages the tensor-parallel grouping.

## Practical Configuration Examples

```bash

# Automatic VRAM detection on eight GPUs (order determines layer placement)

./ds4 -m ds4flash.gguf \
  --gpu-vram auto \
  --gpu-devices 0,2,4,6,1,3,5,7 \
  --ctx 100000

# Explicit VRAM budgets for two GPUs (40GiB on device 0, 12GiB on device 1)

./ds4 -m ds4flash.gguf \
  --gpu-vram 40,12 \
  --gpu-devices 0,1 \
  --ctx 80000

# CPU-only inference (disables GPU entirely)

./ds4 -m ds4flash.gguf \
  --gpu-vram 0 \
  --ctx 50000

# Tensor parallelism across four GPUs (pairs: 0-1 and 2-3)

./ds4 -m ds4flash.gguf \
  --gpu-devices 0,1,2,3 \
  --cuda-tensor-parallel \
  --ctx 200000

```

## Summary

- **`--gpu-devices`** specifies which CUDA indices to use and determines layer placement order.
- **`--gpu-vram`** accepts `auto` for detection, `0` to disable GPUs, or comma-separated GiB values.
- **Length matching** is enforced: explicit VRAM lists must contain one entry per device.
- **Source files**: Parsing logic resides in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c), probing in `ds4_cuda.cu`, and help text in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c).
- **Tensor parallelism** automatically pairs devices from the provided list for distributed layer computation.

## Frequently Asked Questions

### What happens if I specify `--gpu-devices` without `--gpu-vram`?

DS4 implicitly sets `--gpu-vram auto` and probes only the specified devices for available free memory. This is handled in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) by forcing the VRAM argument to `"auto"` when the device argument is present but VRAM is NULL.

### Can I use `--gpu-devices` when running in CPU-only mode?

No. If you set `--gpu-vram 0` to disable CUDA, the parser explicitly rejects any `--gpu-devices` argument and aborts with an error, as device filtering is invalid when GPU acceleration is disabled.

### How does DS4 handle mismatched list lengths between the two flags?

The parser validates that the number of entries in `--gpu-devices` matches `--gpu-vram` when explicit budgets are provided. If they differ, DS4 emits the error `"--gpu-devices count … does not match --gpu-vram count …"` and exits before model loading begins.

### Does the order of devices in `--gpu-devices` affect performance?

Yes. DS4 assigns consecutive model layer ranges to devices in the order listed. In tensor-parallel mode, consecutive pairs form split groups. Placing higher-bandwidth or faster GPUs earlier in the sequence can reduce pipeline bottlenecks during inference.