How to Configure --gpu-devices and --gpu-vram for Multi-GPU Placement in DS4

TLDR: DS4 uses --gpu-devices to select specific CUDA GPU indices and --gpu-vram to define per-device memory budgets (in GiB) or enable automatic detection via auto, with both flags parsed and validated in ds4_gpu_args.c to ensure matching device counts and valid allocation strategies.

DS4 is a high-performance local LLM inference engine developed by antirez. When deploying large language models across multiple NVIDIA GPUs, precise control over multi-GPU placement is essential for maximizing throughput and avoiding out-of-memory errors. This article breaks down the exact syntax, validation logic, and runtime behavior of the --gpu-devices and --gpu-vram flags based on the source code implementation.

Understanding the Two Key Flags

--gpu-devices: Selecting CUDA Device Indices

The --gpu-devices flag accepts a comma-separated list of CUDA device indices (e.g., 0,2,4,6) that DS4 is permitted to use. The order you specify matters: DS4 assigns model layers to devices in the sequence provided, making this the primary mechanism for controlling multi-GPU placement topology.

--gpu-vram: Memory Allocation Modes

The --gpu-vram flag controls how much video RAM is allocated per device. It accepts three distinct input types:

  • auto: Probe the system for available free VRAM on the selected devices.
  • 0: Disable CUDA entirely and run in CPU-only mode.
  • N,N,...: Explicit GiB values for each device (must match --gpu-devices length).

Parsing Logic and Validation in ds4_gpu_args.c

The central parsing routine parse_gpu_vram_arg in ds4_gpu_args.c handles all four configuration scenarios:

Case 1: Implicit Auto-Detection

When you supply --gpu-devices without --gpu-vram, the parser forces vram_arg to "auto" (lines 29-37). DS4 then probes only the specified devices for available memory.

Case 2: CPU-Only Mode (--gpu-vram 0)

Setting --gpu-vram 0 disables GPU acceleration completely. The code explicitly rejects simultaneous use of --gpu-devices in this mode (lines 40-47), as filtering devices is meaningless when CUDA is disabled.

Case 3: Automatic VRAM Probing (--gpu-vram auto)

In CUDA builds (gated by #if !defined(DS4_NO_GPU)), DS4 calls ds4_gpu_args_probe_auto_cuda from ds4_cuda.cu (lines 51-59) to detect free memory on each selected device. If --gpu-devices is provided, the probe is limited to that subset.

Case 4: Explicit Budgets with Strict Validation

When providing explicit GiB values (e.g., 40,12), the parser uses parse_csv_int_list to tokenize the input. If both flags are present, the code enforces that the list lengths match exactly, emitting the error "--gpu-devices count … does not match --gpu-vram count …" if they differ (lines 84-88). The resulting mapping is stored in the ds4_gpu_config structure (lines 90-95) as device_indices[i] → vram_bytes[i].

How Layer Placement Works Internally

DS4's runtime uses the validated configuration to distribute model layers across GPUs:

Sequential Placement

In standard multi-GPU mode (without tensor parallelism), consecutive layer ranges are assigned to devices in the order specified by --gpu-devices. Placing faster GPUs earlier in the list can improve pipeline efficiency.

Tensor Parallelism Pairing

When --cuda-tensor-parallel is enabled, devices are grouped into pairs: the first two devices split layers 50/50, the next two form another pair, and so on. The --gpu-devices flag still controls which physical GPUs participate, but the runtime automatically manages the tensor-parallel grouping.

Practical Configuration Examples


# Automatic VRAM detection on eight GPUs (order determines layer placement)

./ds4 -m ds4flash.gguf \
  --gpu-vram auto \
  --gpu-devices 0,2,4,6,1,3,5,7 \
  --ctx 100000

# Explicit VRAM budgets for two GPUs (40GiB on device 0, 12GiB on device 1)

./ds4 -m ds4flash.gguf \
  --gpu-vram 40,12 \
  --gpu-devices 0,1 \
  --ctx 80000

# CPU-only inference (disables GPU entirely)

./ds4 -m ds4flash.gguf \
  --gpu-vram 0 \
  --ctx 50000

# Tensor parallelism across four GPUs (pairs: 0-1 and 2-3)

./ds4 -m ds4flash.gguf \
  --gpu-devices 0,1,2,3 \
  --cuda-tensor-parallel \
  --ctx 200000

Summary

  • --gpu-devices specifies which CUDA indices to use and determines layer placement order.
  • --gpu-vram accepts auto for detection, 0 to disable GPUs, or comma-separated GiB values.
  • Length matching is enforced: explicit VRAM lists must contain one entry per device.
  • Source files: Parsing logic resides in ds4_gpu_args.c, probing in ds4_cuda.cu, and help text in ds4_help.c.
  • Tensor parallelism automatically pairs devices from the provided list for distributed layer computation.

Frequently Asked Questions

What happens if I specify --gpu-devices without --gpu-vram?

DS4 implicitly sets --gpu-vram auto and probes only the specified devices for available free memory. This is handled in ds4_gpu_args.c by forcing the VRAM argument to "auto" when the device argument is present but VRAM is NULL.

Can I use --gpu-devices when running in CPU-only mode?

No. If you set --gpu-vram 0 to disable CUDA, the parser explicitly rejects any --gpu-devices argument and aborts with an error, as device filtering is invalid when GPU acceleration is disabled.

How does DS4 handle mismatched list lengths between the two flags?

The parser validates that the number of entries in --gpu-devices matches --gpu-vram when explicit budgets are provided. If they differ, DS4 emits the error "--gpu-devices count … does not match --gpu-vram count …" and exits before model loading begins.

Does the order of devices in --gpu-devices affect performance?

Yes. DS4 assigns consecutive model layer ranges to devices in the order listed. In tensor-parallel mode, consecutive pairs form split groups. Placing higher-bandwidth or faster GPUs earlier in the sequence can reduce pipeline bottlenecks during inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →