Best Practices for Multi-GPU Memory Allocation with Placement Hints in ds4
Use the ds4_gpu_mgpu allocator with explicit ds4_gpu_placement_t hints to control tensor placement across GPUs, falling back to relaxed masks when primary hints fail.
The ds4 deep learning engine by antirez provides a purpose-built multi-GPU memory manager that lets developers control exactly where tensors reside through placement hints. This article explores the architecture, API patterns, and production-tested workflows drawn directly from the ds4 source code.
Core Architecture of ds4 Multi-GPU Memory Management
The ds4 codebase separates concerns across three layers: the low-level multi-GPU allocator, the argument parsing layer, and the engine integration tests.
The Placement Hint Structure
At the heart of the system is ds4_gpu_placement_t, defined in [ds4_gpu_mgpu.h](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h):
typedef struct {
uint32_t mask; /* Bitmap of preferred GPU IDs */
} ds4_gpu_placement_t;
The mask uses bit positions to indicate GPU preferences: bit 0 for GPU 0, bit 1 for GPU 1, and so on. A zero mask allows allocation on any available GPU.
Memory Pool Per GPU
[ds4_gpu_mgpu.c](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.c) maintains separate memory pools for each GPU to eliminate cross-GPU contention. During initialization (ds4_gpu_mgpu_init()), the allocator probes all visible GPUs and records their total capacities. Each pool operates independently with its own free lists and fragmentation management through chunk splitting and coalescing.
Five Best Practices for Production Use
1. Initialize Early with Correct GPU Enumeration
Call ds4_gpu_mgpu_init() before any allocation. The typical entry point is ds4_cli.c:
/* Early in main() */
int num_gpus = ds4_gpu_detect();
if (ds4_gpu_mgpu_init(num_gpus, &config) != 0) {
fprintf(stderr, "GPU init failed: %s\n", ds4_gpu_mgpu_last_error());
return 1;
}
This probes hardware and establishes the per-GPU pools that all subsequent allocations use.
2. Use Explicit Single-GPU Hints for Sharded Data
Model weights sharded across GPUs should use single-bit masks to guarantee locality:
/* Allocate 64 MiB pinned to GPU-0 */
ds4_gpu_placement_t hint = { .mask = 1 << 0 };
void *weights = ds4_gpu_mgpu_alloc(64ULL * 1024 * 1024, &hint);
This prevents the allocator from scattering shards, which would trigger expensive inter-GPU transfers during forward passes.
3. Use Multi-Bit Hints for Replicated Buffers
Activations that can live on multiple GPUs should specify alternatives:
/* 32 MiB buffer acceptable on GPU-0 OR GPU-1 */
ds4_gpu_placement_t hint = { .mask = (1 << 0) | (1 << 1) };
void *activations = ds4_gpu_mgpu_alloc(32ULL * 1024 * 1024, &hint);
The allocator tries GPUs in order of lowest memory pressure among the hinted set.
4. Implement Graceful Fallback on Allocation Failure
When primary hints cannot be satisfied, relax constraints rather than failing:
ds4_gpu_placement_t hint = { .mask = 1 << 2 }; /* Prefer GPU-2 */
size_t bytes = 128ULL * 1024 * 1024;
void *buf = ds4_gpu_mgpu_alloc(bytes, &hint);
if (!buf) {
/* Fallback: any GPU */
hint.mask = 0xFFFFFFFF; /* All 32 possible GPUs */
buf = ds4_gpu_mgpu_alloc(bytes, &hint);
}
This pattern appears in [test_engine_mgpu_refusal.c](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_refusal.c), which validates recovery from out-of-memory conditions.
5. Always Free Through the Multi-GPU API
Use ds4_gpu_mgpu_free() to ensure the correct pool is updated:
ds4_gpu_mgpu_free(buf); /* Updates the originating GPU's pool statistics */
Raw free() would corrupt the allocator's metadata.
CLI Integration: Parsing Placement Hints
The [ds4_gpu_args.c](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) module translates command-line flags into placement masks:
./ds4 --gpu-placement=0,1 --tensor-size=64M train
Internally, this builds mask = (1 << 0) | (1 << 1) and passes it to the allocation path. The parser also supports exclusion syntax (--gpu-exclude=2) for heterogeneous systems.
Validation Through Test Suite
The ds4 test suite provides reference implementations of these patterns.
Placement Correctness Test
[tests/test_engine_mgpu_placement.c](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_placement.c) validates:
- Single-bit hints land on the specified GPU
- Multi-bit hints respect the preference order
- Zero-mask hints distribute load across all GPUs
Refusal and Recovery Test
[tests/test_engine_mgpu_refusal.c](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_refusal.c) exercises:
- Allocation failure when hinted GPU is full
- Correct error reporting via
ds4_gpu_mgpu_last_error() - Successful reallocation with relaxed hints
Performance Characteristics
| Pattern | Latency | Bandwidth Impact | Best For |
|---|---|---|---|
| Single-bit hint | ~1 µs | Zero cross-GPU traffic | Model weights, persistent buffers |
| Multi-bit hint | ~2-5 µs | Minimal if first GPU succeeds | Activations, temporary tensors |
| Zero mask (any GPU) | ~5-10 µs | Potential inter-GPU copies | Opportunistic caching |
Common Pitfalls
- Orphaned hints: Setting bits for non-existent GPUs (e.g.,
mask = 1 << 8on a 4-GPU system) causes immediate fallback to the least-loaded GPU with no warning. Validatenum_gpusagainst your mask. - Hint over-constriction: Single-bit hints on nearly-full GPUs trigger frequent fallbacks. Monitor
ds4_gpu_mgpu_pool_stats()to balance your hints. - Mixed allocation APIs: Never mix
ds4_gpu_mgpu_alloc()with rawcudaMalloc()orhipMalloc()in the same program. The pools become inconsistent.
Summary
- Initialize with
ds4_gpu_mgpu_init()before any tensor operations - Pin sharded data with single-bit placement hints
- Replicate opportunistically with multi-bit masks
- Fallback gracefully by relaxing hints on allocation failure
- Free exclusively through
ds4_gpu_mgpu_free()to maintain pool consistency
Frequently Asked Questions
How does ds4's placement hint differ from CUDA memory pools?
CUDA's memory pools are device-specific with explicit device context switching. ds4's ds4_gpu_placement_t adds a portable abstraction layer that works across CUDA and HIP backends, with automatic fallback across GPUs rather than per-device failure.
Can I change placement hints after allocation?
No. The placement hint is evaluated at allocation time and determines which GPU pool owns the memory. To migrate data, allocate new memory with a different hint and copy explicitly—ds4 does not support transparent migration.
What happens when all hinted GPUs are exhausted?
If no hinted GPU can satisfy the request, ds4_gpu_mgpu_alloc() returns NULL. The caller must either relax the hint mask or propagate the error. The allocator never silently allocates on non-hinted GPUs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →