Managing Multi-GPU Memory Placement and Session Allocation in ds4: A Deep Dive into the Layer Packing Algorithm

The ds4 library implements a deterministic monotonic-contiguous placement algorithm through ds4_compute_layer_placement in ds4_layer_pack.c to assign transformer layers across GPU tiers based on configurable memory budgets, while session allocation validates these placements against context-size and concurrency hints to prevent runtime OOM errors.

The ds4 inference engine, maintained by antirez, provides deterministic multi-GPU memory management for transformer inference workloads. Understanding how to manage multi-GPU memory placement and session allocation in ds4 is critical for maximizing hardware utilization while ensuring that KV-cache and activation buffers remain within device memory limits. This article examines the greedy allocation algorithm, configuration hints, and strict validation logic that govern layer distribution across heterogeneous memory tiers.

How ds4 Computes Layer Placement on Multiple GPUs

The core of ds4's multi-GPU strategy resides in the ds4_compute_layer_placement function defined in ds4_layer_pack.c. This routine implements a greedy, monotonic-contiguous packing algorithm that assigns each transformer layer entry to a specific GPU or the CPU tier.

The Monotonic-Contiguous Allocation Algorithm

The placement function receives four key parameters: an array of entry sizes (entry_bytes), the number of entries, a configuration struct (ds4_layer_pack_config) containing the number of GPUs and per-GPU memory budgets, and an output array (device_for_entry) to store device assignments.

The algorithm operates with strict monotonicity: once it advances to a higher-index GPU, it never returns to a lower-index one. This guarantees that each GPU holds a contiguous block of layers. The function first copies GPU budgets locally, then iterates through entries in order. For each entry, it advances to the next GPU until finding a device with sufficient remaining budget. If no GPU can accommodate the entry, the function assigns it to the special constant DS4_LAYER_PACK_CPU.

The core allocation loop (lines 30-43) implements this logic:

while (d < cfg->n_gpus && entry_bytes[e] > budget[d]) d++;
if (d < cfg->n_gpus) {
    device_for_entry[e] = d;
    budget[d] -= entry_bytes[e];
} else {
    device_for_entry[e] = DS4_LAYER_PACK_CPU;
}

This approach ensures deterministic placement that respects user-specified gpu_budget_bytes limits while minimizing fragmentation by maintaining contiguous layer blocks on each device.

Session Allocation and GPU Memory Budgeting

Sessions are created through the engine API in ds4.c, which integrates the placement algorithm with runtime memory allocation. When a session is requested, the engine computes placement using the packing routine and then allocates required memory on each tier.

Configuration Hints for Context and Session Count

Two critical configuration hints influence memory allocation during session creation:

  • placement_ctx_hint: Specifies the expected context size in bytes. The engine propagates this hint to per-layer KV-store sizing, ensuring the memory layout accommodates the intended context length. According to the ds4 source code, this hint is processed in ds4_server.c at line 13042.

  • placement_session_count_hint: Indicates the number of concurrent sessions expected. This allows the engine to pre-reserve sufficient memory across GPUs to avoid runtime OOM errors during high-concurrency scenarios.

Validation and Refusal Logic

The engine validates computed placements before allocating GPU buffers. If any layer would spill to the CPU tier while the session is configured as GPU-only, the allocation is refused via engine_install_gpu_placement. This safety mechanism ensures that GPU-only workloads never accidentally access CPU memory, which would degrade performance.

The test suite in test_engine_mgpu_refusal.c exercises this behavior, confirming that the engine correctly rejects sessions requiring GPU-only placement when the algorithm would otherwise spill layers to CPU. Upon successful validation, the engine records device assignments and allocates GPU buffers (KV cache, attention matrices) such that total bytes allocated never exceed the user-provided gpu_budget_bytes.

Practical Implementation: Configuring Multi-GPU Placement

The following implementation pattern demonstrates how to configure GPU budgets, compute layer placement, and create a session with appropriate hints:

/* 1️⃣ Configure GPU budgets (in bytes). */
ds4_layer_pack_config cfg = {
    .n_gpus = 2,
    .gpu_budget_bytes = { 8ULL * 1024 * 1024 * 1024,   // 8 GB on GPU0
                           6ULL * 1024 * 1024 * 1024 }   // 6 GB on GPU1
};

/* 2️⃣ Prepare entry sizes (e.g., KV‑cache + activation buffers). */
size_t entry_bytes[DS4_N_ENTRY] = { /* … fill with per‑layer sizes … */ };

/* 3️⃣ Compute placement. */
int device_for_entry[DS4_N_ENTRY];
int rc = ds4_compute_layer_placement(entry_bytes,
                                     DS4_N_ENTRY,
                                     &cfg,
                                     device_for_entry);
if (rc != 0) { /* handle error */ }

/* 4️⃣ Create an engine and pass placement hints. */
ds4_engine engine = ds4_engine_create();
engine.cfg.ctx_size = 4096;                 // hint: 4 k context length
engine.cfg.placement_ctx_hint = 4096;
engine.cfg.placement_session_count_hint = 2;

/* 5️⃣ Request a session – the engine will reuse the placement we just computed. */
ds4_session *sess = ds4_engine_create_session(&engine,
                                             entry_bytes,
                                             DS4_N_ENTRY);
if (!sess) { /* allocation failed – likely due to budget overflow */ }

This workflow ensures that layer placement respects hardware constraints while allowing the engine to optimize KV-cache sizing based on expected workload characteristics.

Verification via Test Coverage

The ds4 codebase includes comprehensive test coverage for multi-GPU placement scenarios. The test_layer_pack.c file validates that ds4_compute_layer_placement returns expected device indices for various entry patterns, including edge cases such as null pointers, out-of-range GPU counts, and CPU spill conditions.

Integration testing in test_engine_mgpu_placement.c drives realistic multi-tier placement decisions using the engine API, verifying that placement respects placement_ctx_hint scaling and that GPU budgets are honored during session creation. These tests ensure that the monotonic allocation algorithm behaves correctly under production workloads.

Summary

  • Deterministic placement: The ds4_compute_layer_placement function in ds4_layer_pack.c implements a greedy monotonic-contiguous algorithm that assigns layers to GPUs without backtracking, ensuring contiguous memory blocks per device.
  • CPU fallback: When GPU budgets are exhausted, the algorithm assigns remaining layers to DS4_LAYER_PACK_CPU, though GPU-only sessions can configure the engine to refuse such placements.
  • Configuration hints: The placement_ctx_hint and placement_session_count_hint parameters allow pre-sizing of KV caches and reservation of GPU memory for concurrent sessions.
  • Strict validation: The engine validates placements via engine_install_gpu_placement to prevent GPU-only sessions from spilling to CPU, with test coverage in test_engine_mgpu_refusal.c.

Frequently Asked Questions

How does ds4 handle transformer layers that exceed individual GPU memory budgets?

When a single layer's size exceeds the remaining budget on any available GPU, ds4_compute_layer_placement assigns that layer to DS4_LAYER_PACK_CPU, indicating it should reside in host memory. For GPU-only sessions, the engine detects this CPU spill during validation and refuses the allocation to prevent performance degradation.

What is the purpose of the placement_ctx_hint in ds4 session allocation?

The placement_ctx_hint parameter informs the engine about the expected context size in bytes, allowing it to properly size per-layer KV caches during memory allocation. This hint, processed in ds4_server.c, ensures that the computed placement accounts for the actual memory footprint required to support the target sequence length.

Can ds4 distribute transformer layers non-monotonically across multiple GPUs?

No. The ds4 placement algorithm is strictly monotonic: once it advances from GPU i to GPU i+1, it never returns to GPU i for subsequent layers. This design guarantees contiguous layer blocks on each GPU, simplifying memory management and ensuring predictable access patterns during inference.

How does the ds4 engine validate GPU-only session requests?

The engine validates GPU-only sessions through engine_install_gpu_placement, which checks if the computed placement would assign any layer to DS4_LAYER_PACK_CPU. If CPU spill is detected for a GPU-only session, the engine refuses the allocation immediately. This behavior is verified in test_engine_mgpu_refusal.c, ensuring that budget violations are caught during session creation rather than at runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →