How to Troubleshoot CUDA Out-of-Memory Errors in Multi-GPU Setups with DwarfStar (ds4)

To resolve CUDA out-of-memory errors in DwarfStar multi-GPU deployments, verify GPU layouts with --gpu-vram auto, adjust the 512 MiB safety margin in ds4_gpu_args.c, enforce explicit per-GPU VRAM caps, and ensure correct Tensor-Parallel device ordering.

DwarfStar (ds4) is an open-source inference engine that manages multi-GPU memory through sophisticated budgeting and safety mechanisms implemented in C and CUDA. When running large language models across multiple CUDA devices, out-of-memory errors typically stem from VRAM fragmentation, aggressive safety margins, or incorrect device ordering in Tensor-Parallel configurations. Understanding how the engine allocates memory budgets is essential for stable multi-GPU inference.

Understanding DwarfStar's GPU Memory Architecture

The ds4 engine controls CUDA memory through two primary mechanisms defined in main/ds4_gpu_args.c and main/ds4_cuda.cu.

GPU-VRAM Budgeting and Safety Margins

The budgeting system parses command-line arguments --gpu-vram and --gpu-devices to build per-device memory budgets. Each budget automatically applies a safety margin of 512 MiB defined by the constant DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN at line 17 of ds4_gpu_args.c. This margin prevents driver-level OOM crashes by reserving headroom for CUDA runtime overhead.

Auto-Probing Mechanism

When --gpu-vram auto is specified, the function ds4_gpu_args_probe_auto_cuda() in ds4_cuda.cu probes each CUDA device and calculates available memory. The implementation discards devices that cannot satisfy the safety margin and returns a layout that fits the model weights and activation buffers.

Common Causes of CUDA OOM Errors

According to the source code in ds4_gpu_args.c and ds4.c, CUDA out-of-memory errors in multi-GPU setups typically result from four specific configuration issues:

  • Insufficient per-GPU VRAM budgets – The model or routed MoE experts exceed the allocated memory after accounting for the safety margin.
  • Excessive safety margins – The default 512 MiB margin may be too aggressive on heavily fragmented GPUs with smaller VRAM pools.
  • Mismatched device lists – The number of entries in --gpu-devices does not equal the number of entries in --gpu-vram, causing parse_csv_int_list() (lines 45-99) to abort.
  • Incorrect Tensor-Parallel ordering – Device ordering matters for multi-GPU Tensor-Parallel (TP) runs; placing partner-tier devices before home-tier devices creates uneven memory splits.

Step-by-Step Troubleshooting Workflow

Verify GPU Layout Reporting

First, confirm that ds4 correctly recognizes your hardware configuration. Run the binary with --gpu-vram auto and inspect the output generated by format_gpu_layout_line() (lines 199-229):

./ds4 --gpu-vram auto --gpu-devices 0,2,4,6,1,3,5,7 \
      --log-level debug 2>&1 | grep '^ds4: GPU config'

A valid configuration prints:


ds4: GPU config: 8 devices [0,2,4,6,1,3,5,7] requested, budgets 48,48,48,48,48,48,48,48 GB; auto=true

If the line reports 0 devices or fails to appear, check the syntax of your arguments. The parser parse_csv_int_list() requires comma-separated integers without spaces.

Reduce the Safety Margin

If GPUs have limited free memory due to fragmentation, reduce the safety margin by modifying main/ds4_gpu_args.c and recompiling:

- #define DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN ((size_t)512 * 1024 * 1024)
+ #define DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN ((size_t)128 * 1024 * 1024)

Rebuild with make cuda-generic or make cuda-spark and retry inference. Warning: Reducing this margin increases the risk of driver-level OOM crashes if system processes allocate GPU memory during inference.

Enforce Explicit VRAM Caps

When auto-probing selects inappropriate devices, manually cap per-GPU memory. The validation logic in ds4_gpu_args.c enforces a hard limit of DS4_GPU_ARGS_MAX_VRAM_GB (16 TiB) at line 38:

./ds4 --gpu-vram 24,24,24,24,24,24,24,24 \
      --gpu-devices 0,1,2,3,4,5,6,7

Explicit budgets override the auto-probe results and prevent the engine from attempting allocations on devices with insufficient free memory.

Optimize Tensor-Parallel Placement

For CUDA Tensor-Parallel execution, device order must follow home-tiers first, then partner-tiers. Incorrect ordering overloads the first half of the GPU pool. The correct sequence for an 8-GPU host places even indices before odd indices:

./ds4 --cuda-tensor-parallel \
      --gpu-devices 0,2,4,6,1,3,5,7

If OOM persists with correct ordering, reduce the number of active devices or halve the per-GPU budgets.

Validate CUDA Runtime Versions

The probing function ds4_gpu_args_probe_auto_cuda() relies on accurate memory reporting from the CUDA runtime. Outdated drivers may misreport available VRAM. Verify your environment:

nvcc --version
nvidia-smi

Upgrade to drivers supporting your GPU compute capability (e.g., SM 89 for L40S) as specified in the README build instructions.

Debug Kernel-Level Allocations

For persistent OOM errors, enable verbose allocation logging by recompiling with debug flags. Modify CFLAGS in the Makefile to include -DCUDA_DEBUG=1:

make clean && make cuda-generic CFLAGS="-DCUDA_DEBUG=1"
./ds4 ... 2> debug.log
grep "cudaMalloc" debug.log

This reveals exact allocation sizes for KV-cache, expert-cache, and activation buffers. Use this data to adjust --ctx, --prefill-chunk, or --ssd-streaming-cache-experts values.

Reducing Model Memory Footprint

Select Lower Quantization Levels

The repository provides GGUF files with varying bit depths. Switching from 4-bit (q4) to 2-bit (q2) quantization significantly reduces VRAM requirements:

./ds4 -m gguf/DeepSeek-V4-Flash-Q2K.gguf \
       --gpu-vram auto --gpu-devices 0,1

Enable Power Throttling

While primarily affecting thermal output, reducing GPU power to 70% can decrease memory pressure from activation buffers:

./ds4 --cuda-tensor-parallel \
      --gpu-devices 0,2,4,6,1,3,5,7 \
      --power 70

Practical Multi-GPU Configuration Examples

Auto-Probing with All Visible Devices

./ds4 --cuda-tensor-parallel \
      --gpu-devices 0,1,2,3,4,5,6,7 \
      --gpu-vram auto \
      -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf \
      --ctx 100000

Fixed 24 GB Budgets on 48 GB Cards

./ds4 --cuda-tensor-parallel \
      --gpu-devices 0,2,4,6,1,3,5,7 \
      --gpu-vram 24,24,24,24,24,24,24,24 \
      -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf \
      --ctx 100000

Debug Build for Allocation Tracing

make clean && make cuda-generic CFLAGS="-DCUDA_DEBUG=1"
./ds4 --cuda-tensor-parallel --gpu-vram auto -m model.gguf 2> cuda_debug.log
grep "cudaMalloc" cuda_debug.log | head -20

Summary

  • Verify layouts using --gpu-vram auto and format_gpu_layout_line() output before loading models.
  • Adjust safety margins by modifying DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN in ds4_gpu_args.c when running on fragmented GPUs.
  • Use explicit budgets via --gpu-vram to override auto-probe results on shared or partially utilized systems.
  • Order devices correctly for Tensor-Parallel execution (home-tiers before partner-tiers) to prevent uneven memory distribution.
  • Enable debug builds with -DCUDA_DEBUG=1 to inspect exact allocation sizes in ds4_cuda.cu and ds4.c.

Frequently Asked Questions

Why does ds4 report 0 available GPUs when using --gpu-vram auto?

This occurs when parse_csv_int_list() fails to parse your --gpu-devices argument or when ds4_gpu_args_probe_auto_cuda() cannot satisfy the 512 MiB safety margin on any visible device. Check for spaces in your comma-separated list and verify that nvidia-smi shows free memory exceeding the safety margin plus model requirements.

How do I calculate the correct safety margin for my specific GPU model?

The default 512 MiB defined in ds4_gpu_args.c line 17 suits most modern datacenter GPUs. For consumer cards with 8-12 GB VRAM or systems running X11/Wayland compositors, reduce the margin to 128-256 MiB by editing DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN and recompiling. Monitor for driver crashes, as insufficient margins cause allocation failures in CUDA driver libraries rather than graceful errors.

Can I use different VRAM budgets for different GPUs in the same run?

Yes. The --gpu-vram argument accepts comma-separated values corresponding to each entry in --gpu-devices. For example, --gpu-vram 48,24,48,24 --gpu-devices 0,1,2,3 assigns 48 GB to devices 0 and 2, and 24 GB to devices 1 and 3. This is useful when mixing GPU models like A100s and L40Ss in heterogeneous clusters.

What is the difference between --gpu-vram auto and manual budgeting?

--gpu-vram auto invokes ds4_gpu_args_probe_auto_cuda() to query the CUDA runtime for free memory on each specified device, then subtracts the safety margin automatically. Manual budgeting bypasses the probe and trusts your specified values, which is necessary when the CUDA driver incorrectly reports available memory due to persistent processes or hidden allocations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →