How to Troubleshoot CUDA Out-of-Memory Errors in Multi-GPU Setups with DwarfStar (ds4)
To resolve CUDA out-of-memory errors in DwarfStar multi-GPU deployments, verify GPU layouts with --gpu-vram auto, adjust the 512 MiB safety margin in ds4_gpu_args.c, enforce explicit per-GPU VRAM caps, and ensure correct Tensor-Parallel device ordering.
DwarfStar (ds4) is an open-source inference engine that manages multi-GPU memory through sophisticated budgeting and safety mechanisms implemented in C and CUDA. When running large language models across multiple CUDA devices, out-of-memory errors typically stem from VRAM fragmentation, aggressive safety margins, or incorrect device ordering in Tensor-Parallel configurations. Understanding how the engine allocates memory budgets is essential for stable multi-GPU inference.
Understanding DwarfStar's GPU Memory Architecture
The ds4 engine controls CUDA memory through two primary mechanisms defined in main/ds4_gpu_args.c and main/ds4_cuda.cu.
GPU-VRAM Budgeting and Safety Margins
The budgeting system parses command-line arguments --gpu-vram and --gpu-devices to build per-device memory budgets. Each budget automatically applies a safety margin of 512 MiB defined by the constant DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN at line 17 of ds4_gpu_args.c. This margin prevents driver-level OOM crashes by reserving headroom for CUDA runtime overhead.
Auto-Probing Mechanism
When --gpu-vram auto is specified, the function ds4_gpu_args_probe_auto_cuda() in ds4_cuda.cu probes each CUDA device and calculates available memory. The implementation discards devices that cannot satisfy the safety margin and returns a layout that fits the model weights and activation buffers.
Common Causes of CUDA OOM Errors
According to the source code in ds4_gpu_args.c and ds4.c, CUDA out-of-memory errors in multi-GPU setups typically result from four specific configuration issues:
- Insufficient per-GPU VRAM budgets – The model or routed MoE experts exceed the allocated memory after accounting for the safety margin.
- Excessive safety margins – The default 512 MiB margin may be too aggressive on heavily fragmented GPUs with smaller VRAM pools.
- Mismatched device lists – The number of entries in
--gpu-devicesdoes not equal the number of entries in--gpu-vram, causingparse_csv_int_list()(lines 45-99) to abort. - Incorrect Tensor-Parallel ordering – Device ordering matters for multi-GPU Tensor-Parallel (TP) runs; placing partner-tier devices before home-tier devices creates uneven memory splits.
Step-by-Step Troubleshooting Workflow
Verify GPU Layout Reporting
First, confirm that ds4 correctly recognizes your hardware configuration. Run the binary with --gpu-vram auto and inspect the output generated by format_gpu_layout_line() (lines 199-229):
./ds4 --gpu-vram auto --gpu-devices 0,2,4,6,1,3,5,7 \
--log-level debug 2>&1 | grep '^ds4: GPU config'
A valid configuration prints:
ds4: GPU config: 8 devices [0,2,4,6,1,3,5,7] requested, budgets 48,48,48,48,48,48,48,48 GB; auto=true
If the line reports 0 devices or fails to appear, check the syntax of your arguments. The parser parse_csv_int_list() requires comma-separated integers without spaces.
Reduce the Safety Margin
If GPUs have limited free memory due to fragmentation, reduce the safety margin by modifying main/ds4_gpu_args.c and recompiling:
- #define DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN ((size_t)512 * 1024 * 1024)
+ #define DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN ((size_t)128 * 1024 * 1024)
Rebuild with make cuda-generic or make cuda-spark and retry inference. Warning: Reducing this margin increases the risk of driver-level OOM crashes if system processes allocate GPU memory during inference.
Enforce Explicit VRAM Caps
When auto-probing selects inappropriate devices, manually cap per-GPU memory. The validation logic in ds4_gpu_args.c enforces a hard limit of DS4_GPU_ARGS_MAX_VRAM_GB (16 TiB) at line 38:
./ds4 --gpu-vram 24,24,24,24,24,24,24,24 \
--gpu-devices 0,1,2,3,4,5,6,7
Explicit budgets override the auto-probe results and prevent the engine from attempting allocations on devices with insufficient free memory.
Optimize Tensor-Parallel Placement
For CUDA Tensor-Parallel execution, device order must follow home-tiers first, then partner-tiers. Incorrect ordering overloads the first half of the GPU pool. The correct sequence for an 8-GPU host places even indices before odd indices:
./ds4 --cuda-tensor-parallel \
--gpu-devices 0,2,4,6,1,3,5,7
If OOM persists with correct ordering, reduce the number of active devices or halve the per-GPU budgets.
Validate CUDA Runtime Versions
The probing function ds4_gpu_args_probe_auto_cuda() relies on accurate memory reporting from the CUDA runtime. Outdated drivers may misreport available VRAM. Verify your environment:
nvcc --version
nvidia-smi
Upgrade to drivers supporting your GPU compute capability (e.g., SM 89 for L40S) as specified in the README build instructions.
Debug Kernel-Level Allocations
For persistent OOM errors, enable verbose allocation logging by recompiling with debug flags. Modify CFLAGS in the Makefile to include -DCUDA_DEBUG=1:
make clean && make cuda-generic CFLAGS="-DCUDA_DEBUG=1"
./ds4 ... 2> debug.log
grep "cudaMalloc" debug.log
This reveals exact allocation sizes for KV-cache, expert-cache, and activation buffers. Use this data to adjust --ctx, --prefill-chunk, or --ssd-streaming-cache-experts values.
Reducing Model Memory Footprint
Select Lower Quantization Levels
The repository provides GGUF files with varying bit depths. Switching from 4-bit (q4) to 2-bit (q2) quantization significantly reduces VRAM requirements:
./ds4 -m gguf/DeepSeek-V4-Flash-Q2K.gguf \
--gpu-vram auto --gpu-devices 0,1
Enable Power Throttling
While primarily affecting thermal output, reducing GPU power to 70% can decrease memory pressure from activation buffers:
./ds4 --cuda-tensor-parallel \
--gpu-devices 0,2,4,6,1,3,5,7 \
--power 70
Practical Multi-GPU Configuration Examples
Auto-Probing with All Visible Devices
./ds4 --cuda-tensor-parallel \
--gpu-devices 0,1,2,3,4,5,6,7 \
--gpu-vram auto \
-m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf \
--ctx 100000
Fixed 24 GB Budgets on 48 GB Cards
./ds4 --cuda-tensor-parallel \
--gpu-devices 0,2,4,6,1,3,5,7 \
--gpu-vram 24,24,24,24,24,24,24,24 \
-m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf \
--ctx 100000
Debug Build for Allocation Tracing
make clean && make cuda-generic CFLAGS="-DCUDA_DEBUG=1"
./ds4 --cuda-tensor-parallel --gpu-vram auto -m model.gguf 2> cuda_debug.log
grep "cudaMalloc" cuda_debug.log | head -20
Summary
- Verify layouts using
--gpu-vram autoandformat_gpu_layout_line()output before loading models. - Adjust safety margins by modifying
DS4_GPU_ARGS_DEFAULT_SAFETY_MARGINinds4_gpu_args.cwhen running on fragmented GPUs. - Use explicit budgets via
--gpu-vramto override auto-probe results on shared or partially utilized systems. - Order devices correctly for Tensor-Parallel execution (home-tiers before partner-tiers) to prevent uneven memory distribution.
- Enable debug builds with
-DCUDA_DEBUG=1to inspect exact allocation sizes inds4_cuda.cuandds4.c.
Frequently Asked Questions
Why does ds4 report 0 available GPUs when using --gpu-vram auto?
This occurs when parse_csv_int_list() fails to parse your --gpu-devices argument or when ds4_gpu_args_probe_auto_cuda() cannot satisfy the 512 MiB safety margin on any visible device. Check for spaces in your comma-separated list and verify that nvidia-smi shows free memory exceeding the safety margin plus model requirements.
How do I calculate the correct safety margin for my specific GPU model?
The default 512 MiB defined in ds4_gpu_args.c line 17 suits most modern datacenter GPUs. For consumer cards with 8-12 GB VRAM or systems running X11/Wayland compositors, reduce the margin to 128-256 MiB by editing DS4_GPU_ARGS_DEFAULT_SAFETY_MARGIN and recompiling. Monitor for driver crashes, as insufficient margins cause allocation failures in CUDA driver libraries rather than graceful errors.
Can I use different VRAM budgets for different GPUs in the same run?
Yes. The --gpu-vram argument accepts comma-separated values corresponding to each entry in --gpu-devices. For example, --gpu-vram 48,24,48,24 --gpu-devices 0,1,2,3 assigns 48 GB to devices 0 and 2, and 24 GB to devices 1 and 3. This is useful when mixing GPU models like A100s and L40Ss in heterogeneous clusters.
What is the difference between --gpu-vram auto and manual budgeting?
--gpu-vram auto invokes ds4_gpu_args_probe_auto_cuda() to query the CUDA runtime for free memory on each specified device, then subtracts the safety margin automatically. Manual budgeting bypasses the probe and trusts your specified values, which is necessary when the CUDA driver incorrectly reports available memory due to persistent processes or hidden allocations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →