# How to Troubleshoot CUDA Out-of-Memory Errors in Multi-GPU Setups with DwarfStar (ds4)

> Troubleshoot CUDA out-of-memory errors in multi-GPU setups with DwarfStar. Learn to verify GPU layouts, adjust VRAM margins, set caps, and ensure correct device ordering.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**To resolve CUDA out-of-memory errors in DwarfStar multi-GPU deployments, verify GPU layouts with `--gpu-vram auto`, adjust the 512 MiB safety margin in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c), enforce explicit per-GPU VRAM caps, and ensure correct Tensor-Parallel device ordering.**

DwarfStar (`ds4`) is an open-source inference engine that manages multi-GPU memory through sophisticated budgeting and safety mechanisms implemented in C and CUDA. When running large language models across multiple CUDA devices, out-of-memory errors typically stem from VRAM fragmentation, aggressive safety margins, or incorrect device ordering in Tensor-Parallel configurations. Understanding how the engine allocates memory budgets is essential for stable multi-GPU inference.

## Understanding DwarfStar's GPU Memory Architecture

The `ds4` engine controls CUDA memory through two primary mechanisms defined in [`main/ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/main/ds4_gpu_args.c) and `main/ds4_cuda.cu`.

### GPU-VRAM Budgeting and Safety Margins

The budgeting system parses command-line arguments `--gpu-vram` and `--gpu-devices` to build per-device memory budgets. Each budget automatically applies a safety margin of **512 MiB** defined by the constant `DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN` at line 17 of [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c). This margin prevents driver-level OOM crashes by reserving headroom for CUDA runtime overhead.

### Auto-Probing Mechanism

When `--gpu-vram auto` is specified, the function `ds4_gpu_args_probe_auto_cuda()` in `ds4_cuda.cu` probes each CUDA device and calculates available memory. The implementation discards devices that cannot satisfy the safety margin and returns a layout that fits the model weights and activation buffers.

## Common Causes of CUDA OOM Errors

According to the source code in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) and [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), CUDA out-of-memory errors in multi-GPU setups typically result from four specific configuration issues:

- **Insufficient per-GPU VRAM budgets** – The model or routed MoE experts exceed the allocated memory after accounting for the safety margin.
- **Excessive safety margins** – The default 512 MiB margin may be too aggressive on heavily fragmented GPUs with smaller VRAM pools.
- **Mismatched device lists** – The number of entries in `--gpu-devices` does not equal the number of entries in `--gpu-vram`, causing `parse_csv_int_list()` (lines 45-99) to abort.
- **Incorrect Tensor-Parallel ordering** – Device ordering matters for multi-GPU Tensor-Parallel (TP) runs; placing partner-tier devices before home-tier devices creates uneven memory splits.

## Step-by-Step Troubleshooting Workflow

### Verify GPU Layout Reporting

First, confirm that `ds4` correctly recognizes your hardware configuration. Run the binary with `--gpu-vram auto` and inspect the output generated by `format_gpu_layout_line()` (lines 199-229):

```bash
./ds4 --gpu-vram auto --gpu-devices 0,2,4,6,1,3,5,7 \
      --log-level debug 2>&1 | grep '^ds4: GPU config'

```

A valid configuration prints:

```

ds4: GPU config: 8 devices [0,2,4,6,1,3,5,7] requested, budgets 48,48,48,48,48,48,48,48 GB; auto=true

```

If the line reports **0 devices** or fails to appear, check the syntax of your arguments. The parser `parse_csv_int_list()` requires comma-separated integers without spaces.

### Reduce the Safety Margin

If GPUs have limited free memory due to fragmentation, reduce the safety margin by modifying [`main/ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/main/ds4_gpu_args.c) and recompiling:

```diff
- #define DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN ((size_t)512 * 1024 * 1024)
+ #define DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN ((size_t)128 * 1024 * 1024)

```

Rebuild with `make cuda-generic` or `make cuda-spark` and retry inference. **Warning:** Reducing this margin increases the risk of driver-level OOM crashes if system processes allocate GPU memory during inference.

### Enforce Explicit VRAM Caps

When auto-probing selects inappropriate devices, manually cap per-GPU memory. The validation logic in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) enforces a hard limit of `DS4_GPU_ARGS_MAX_VRAM_GB` (16 TiB) at line 38:

```bash
./ds4 --gpu-vram 24,24,24,24,24,24,24,24 \
      --gpu-devices 0,1,2,3,4,5,6,7

```

Explicit budgets override the auto-probe results and prevent the engine from attempting allocations on devices with insufficient free memory.

### Optimize Tensor-Parallel Placement

For CUDA Tensor-Parallel execution, device order must follow **home-tiers first, then partner-tiers**. Incorrect ordering overloads the first half of the GPU pool. The correct sequence for an 8-GPU host places even indices before odd indices:

```bash
./ds4 --cuda-tensor-parallel \
      --gpu-devices 0,2,4,6,1,3,5,7

```

If OOM persists with correct ordering, reduce the number of active devices or halve the per-GPU budgets.

### Validate CUDA Runtime Versions

The probing function `ds4_gpu_args_probe_auto_cuda()` relies on accurate memory reporting from the CUDA runtime. Outdated drivers may misreport available VRAM. Verify your environment:

```bash
nvcc --version
nvidia-smi

```

Upgrade to drivers supporting your GPU compute capability (e.g., SM 89 for L40S) as specified in the README build instructions.

### Debug Kernel-Level Allocations

For persistent OOM errors, enable verbose allocation logging by recompiling with debug flags. Modify `CFLAGS` in the `Makefile` to include `-DCUDA_DEBUG=1`:

```bash
make clean && make cuda-generic CFLAGS="-DCUDA_DEBUG=1"
./ds4 ... 2> debug.log
grep "cudaMalloc" debug.log

```

This reveals exact allocation sizes for KV-cache, expert-cache, and activation buffers. Use this data to adjust `--ctx`, `--prefill-chunk`, or `--ssd-streaming-cache-experts` values.

## Reducing Model Memory Footprint

### Select Lower Quantization Levels

The repository provides GGUF files with varying bit depths. Switching from 4-bit (`q4`) to 2-bit (`q2`) quantization significantly reduces VRAM requirements:

```bash
./ds4 -m gguf/DeepSeek-V4-Flash-Q2K.gguf \
       --gpu-vram auto --gpu-devices 0,1

```

### Enable Power Throttling

While primarily affecting thermal output, reducing GPU power to 70% can decrease memory pressure from activation buffers:

```bash
./ds4 --cuda-tensor-parallel \
      --gpu-devices 0,2,4,6,1,3,5,7 \
      --power 70

```

## Practical Multi-GPU Configuration Examples

### Auto-Probing with All Visible Devices

```bash
./ds4 --cuda-tensor-parallel \
      --gpu-devices 0,1,2,3,4,5,6,7 \
      --gpu-vram auto \
      -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf \
      --ctx 100000

```

### Fixed 24 GB Budgets on 48 GB Cards

```bash
./ds4 --cuda-tensor-parallel \
      --gpu-devices 0,2,4,6,1,3,5,7 \
      --gpu-vram 24,24,24,24,24,24,24,24 \
      -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf \
      --ctx 100000

```

### Debug Build for Allocation Tracing

```bash
make clean && make cuda-generic CFLAGS="-DCUDA_DEBUG=1"
./ds4 --cuda-tensor-parallel --gpu-vram auto -m model.gguf 2> cuda_debug.log
grep "cudaMalloc" cuda_debug.log | head -20

```

## Summary

- **Verify layouts** using `--gpu-vram auto` and `format_gpu_layout_line()` output before loading models.
- **Adjust safety margins** by modifying `DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN` in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) when running on fragmented GPUs.
- **Use explicit budgets** via `--gpu-vram` to override auto-probe results on shared or partially utilized systems.
- **Order devices correctly** for Tensor-Parallel execution (home-tiers before partner-tiers) to prevent uneven memory distribution.
- **Enable debug builds** with `-DCUDA_DEBUG=1` to inspect exact allocation sizes in `ds4_cuda.cu` and [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c).

## Frequently Asked Questions

### Why does `ds4` report 0 available GPUs when using `--gpu-vram auto`?

This occurs when `parse_csv_int_list()` fails to parse your `--gpu-devices` argument or when `ds4_gpu_args_probe_auto_cuda()` cannot satisfy the 512 MiB safety margin on any visible device. Check for spaces in your comma-separated list and verify that `nvidia-smi` shows free memory exceeding the safety margin plus model requirements.

### How do I calculate the correct safety margin for my specific GPU model?

The default 512 MiB defined in [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) line 17 suits most modern datacenter GPUs. For consumer cards with 8-12 GB VRAM or systems running X11/Wayland compositors, reduce the margin to 128-256 MiB by editing `DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN` and recompiling. Monitor for driver crashes, as insufficient margins cause allocation failures in CUDA driver libraries rather than graceful errors.

### Can I use different VRAM budgets for different GPUs in the same run?

Yes. The `--gpu-vram` argument accepts comma-separated values corresponding to each entry in `--gpu-devices`. For example, `--gpu-vram 48,24,48,24 --gpu-devices 0,1,2,3` assigns 48 GB to devices 0 and 2, and 24 GB to devices 1 and 3. This is useful when mixing GPU models like A100s and L40Ss in heterogeneous clusters.

### What is the difference between `--gpu-vram auto` and manual budgeting?

`--gpu-vram auto` invokes `ds4_gpu_args_probe_auto_cuda()` to query the CUDA runtime for free memory on each specified device, then subtracts the safety margin automatically. Manual budgeting bypasses the probe and trusts your specified values, which is necessary when the CUDA driver incorrectly reports available memory due to persistent processes or hidden allocations.