# Best Practices for Multi-GPU Memory Allocation with Placement Hints in ds4

> Master multi-GPU memory allocation with ds4 best practices. Utilize placement hints for optimal tensor control and ensure fallback strategies.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: best-practices
- Published: 2026-08-05

---

**Use the `ds4_gpu_mgpu` allocator with explicit `ds4_gpu_placement_t` hints to control tensor placement across GPUs, falling back to relaxed masks when primary hints fail.**

The ds4 deep learning engine by antirez provides a purpose-built multi-GPU memory manager that lets developers control exactly where tensors reside through placement hints. This article explores the architecture, API patterns, and production-tested workflows drawn directly from the ds4 source code.

## Core Architecture of ds4 Multi-GPU Memory Management

The ds4 codebase separates concerns across three layers: the low-level multi-GPU allocator, the argument parsing layer, and the engine integration tests.

### The Placement Hint Structure

At the heart of the system is `ds4_gpu_placement_t`, defined in [[`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h)](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h):

```c
typedef struct {
    uint32_t mask;   /* Bitmap of preferred GPU IDs */
} ds4_gpu_placement_t;

```

The **mask** uses bit positions to indicate GPU preferences: bit 0 for GPU 0, bit 1 for GPU 1, and so on. A zero mask allows allocation on any available GPU.

### Memory Pool Per GPU

[[`ds4_gpu_mgpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.c)](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.c) maintains separate memory pools for each GPU to eliminate cross-GPU contention. During initialization (`ds4_gpu_mgpu_init()`), the allocator probes all visible GPUs and records their total capacities. Each pool operates independently with its own free lists and fragmentation management through chunk splitting and coalescing.

## Five Best Practices for Production Use

### 1. Initialize Early with Correct GPU Enumeration

Call `ds4_gpu_mgpu_init()` before any allocation. The typical entry point is [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c):

```c
/* Early in main() */
int num_gpus = ds4_gpu_detect();
if (ds4_gpu_mgpu_init(num_gpus, &config) != 0) {
    fprintf(stderr, "GPU init failed: %s\n", ds4_gpu_mgpu_last_error());
    return 1;
}

```

This probes hardware and establishes the per-GPU pools that all subsequent allocations use.

### 2. Use Explicit Single-GPU Hints for Sharded Data

Model weights sharded across GPUs should use single-bit masks to guarantee locality:

```c
/* Allocate 64 MiB pinned to GPU-0 */
ds4_gpu_placement_t hint = { .mask = 1 << 0 };
void *weights = ds4_gpu_mgpu_alloc(64ULL * 1024 * 1024, &hint);

```

This prevents the allocator from scattering shards, which would trigger expensive inter-GPU transfers during forward passes.

### 3. Use Multi-Bit Hints for Replicated Buffers

Activations that can live on multiple GPUs should specify alternatives:

```c
/* 32 MiB buffer acceptable on GPU-0 OR GPU-1 */
ds4_gpu_placement_t hint = { .mask = (1 << 0) | (1 << 1) };
void *activations = ds4_gpu_mgpu_alloc(32ULL * 1024 * 1024, &hint);

```

The allocator tries GPUs in order of lowest memory pressure among the hinted set.

### 4. Implement Graceful Fallback on Allocation Failure

When primary hints cannot be satisfied, relax constraints rather than failing:

```c
ds4_gpu_placement_t hint = { .mask = 1 << 2 };   /* Prefer GPU-2 */
size_t bytes = 128ULL * 1024 * 1024;
void *buf = ds4_gpu_mgpu_alloc(bytes, &hint);

if (!buf) {
    /* Fallback: any GPU */
    hint.mask = 0xFFFFFFFF;   /* All 32 possible GPUs */
    buf = ds4_gpu_mgpu_alloc(bytes, &hint);
}

```

This pattern appears in [[`test_engine_mgpu_refusal.c`](https://github.com/antirez/ds4/blob/main/test_engine_mgpu_refusal.c)](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_refusal.c), which validates recovery from out-of-memory conditions.

### 5. Always Free Through the Multi-GPU API

Use `ds4_gpu_mgpu_free()` to ensure the correct pool is updated:

```c
ds4_gpu_mgpu_free(buf);   /* Updates the originating GPU's pool statistics */

```

Raw `free()` would corrupt the allocator's metadata.

## CLI Integration: Parsing Placement Hints

The [[`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c)](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) module translates command-line flags into placement masks:

```bash
./ds4 --gpu-placement=0,1 --tensor-size=64M train

```

Internally, this builds `mask = (1 << 0) | (1 << 1)` and passes it to the allocation path. The parser also supports exclusion syntax (`--gpu-exclude=2`) for heterogeneous systems.

## Validation Through Test Suite

The ds4 test suite provides reference implementations of these patterns.

### Placement Correctness Test

[[`tests/test_engine_mgpu_placement.c`](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_placement.c)](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_placement.c) validates:

- Single-bit hints land on the specified GPU
- Multi-bit hints respect the preference order
- Zero-mask hints distribute load across all GPUs

### Refusal and Recovery Test

[[`tests/test_engine_mgpu_refusal.c`](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_refusal.c)](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_refusal.c) exercises:

- Allocation failure when hinted GPU is full
- Correct error reporting via `ds4_gpu_mgpu_last_error()`
- Successful reallocation with relaxed hints

## Performance Characteristics

| Pattern | Latency | Bandwidth Impact | Best For |
|---------|---------|------------------|----------|
| Single-bit hint | ~1 µs | Zero cross-GPU traffic | Model weights, persistent buffers |
| Multi-bit hint | ~2-5 µs | Minimal if first GPU succeeds | Activations, temporary tensors |
| Zero mask (any GPU) | ~5-10 µs | Potential inter-GPU copies | Opportunistic caching |

## Common Pitfalls

- **Orphaned hints**: Setting bits for non-existent GPUs (e.g., `mask = 1 << 8` on a 4-GPU system) causes immediate fallback to the least-loaded GPU with no warning. Validate `num_gpus` against your mask.
- **Hint over-constriction**: Single-bit hints on nearly-full GPUs trigger frequent fallbacks. Monitor `ds4_gpu_mgpu_pool_stats()` to balance your hints.
- **Mixed allocation APIs**: Never mix `ds4_gpu_mgpu_alloc()` with raw `cudaMalloc()` or `hipMalloc()` in the same program. The pools become inconsistent.

## Summary

- **Initialize** with `ds4_gpu_mgpu_init()` before any tensor operations
- **Pin** sharded data with single-bit placement hints
- **Replicate** opportunistically with multi-bit masks
- **Fallback** gracefully by relaxing hints on allocation failure
- **Free** exclusively through `ds4_gpu_mgpu_free()` to maintain pool consistency

## Frequently Asked Questions

### How does ds4's placement hint differ from CUDA memory pools?

CUDA's memory pools are device-specific with explicit device context switching. ds4's `ds4_gpu_placement_t` adds a **portable abstraction layer** that works across CUDA and HIP backends, with automatic fallback across GPUs rather than per-device failure.

### Can I change placement hints after allocation?

No. The placement hint is evaluated at allocation time and determines which GPU pool owns the memory. To migrate data, allocate new memory with a different hint and copy explicitly—ds4 does not support transparent migration.

### What happens when all hinted GPUs are exhausted?

If no hinted GPU can satisfy the request, `ds4_gpu_mgpu_alloc()` returns `NULL`. The caller must either relax the hint mask or propagate the error. The allocator never silently allocates on non-hinted GPUs.