How to Implement Multi-GPU Inference with CUDA Tensor Parallelism in ds4

The ds4 engine enables multi-GPU inference by splitting model tensors across CUDA devices using a tier-based configuration system, allowing tensor parallelism without duplicating kernel implementations.

The antirez/ds4 repository provides a native CUDA backend that scales transformer inference across multiple GPUs through tensor parallelism. By leveraging the multi-GPU plumbing layer in ds4_gpu_mgpu.h and its implementation in ds4_cuda.c, developers can distribute model weights and activations across devices while reusing existing single-GPU kernels. This architecture keeps inference code device-agnostic—whether running on one GPU or eight, only the configuration layer changes.

Configure the Multi-GPU Environment

Multi-GPU inference starts with a ds4_gpu_config structure that enumerates the physical CUDA devices and optionally reserves per-GPU VRAM budgets. Defined in ds4_gpu_mgpu.h at lines 71–84, this structure maps logical tiers to hardware devices.

ds4_gpu_config cfg = {0};
cfg.n_gpus = 2;                     // Number of GPUs to use
cfg.device_indices[0] = 0;          // Logical tier 0 → CUDA device 0
cfg.device_indices[1] = 1;          // Logical tier 1 → CUDA device 1
cfg.vram_bytes[0] = (size_t)4 << 30; // Optional: 4 GiB reserved on GPU 0
cfg.vram_bytes[1] = (size_t)4 << 30; // Optional: 4 GiB reserved on GPU 1

The device_indices array allows arbitrary hardware topologies—you can skip device IDs or reorder them to match your physical layout. The n_gpus field determines how many entries in the array are active.

Initialize the Runtime

After configuration, call ds4_gpu_init_multi() to bootstrap the multi-GPU runtime. This function, declared at lines 105–106 of ds4_gpu_mgpu.h, creates a global g_gpu[] array that holds per-GPU CUDA streams, cuBLAS handles, and bookkeeping buffers.

if (ds4_gpu_init_multi(&cfg) != 0) {
    fprintf(stderr, "Failed to initialize multi-GPU runtime\n");
    return 1;
}

Once initialized, the runtime manages device contexts internally. You do not need to call cudaSetDevice manually—the ds4 abstraction layer handles physical device mapping based on the logical tiers defined in your configuration.

Allocate Tensors on Specific Devices

With the runtime active, allocate tensors on specific logical tiers using ds4_gpu_tensor_alloc_ptr_on(). This function, declared at lines 18–19 of ds4_gpu_mgpu.h, takes a tier index and byte count, returning a ds4_gpu_tensor pointer bound to the corresponding physical GPU.

uint64_t weight_bytes = 512 * 1024 * 1024;  // 512 MiB
uint64_t act_bytes = 1024 * sizeof(float);

ds4_gpu_tensor *weights_tier0 = ds4_gpu_tensor_alloc_ptr_on(0, weight_bytes);
ds4_gpu_tensor *activations_tier1 = ds4_gpu_tensor_alloc_ptr_on(1, act_bytes);

Tier 0 corresponds to cfg.device_indices[0], tier 1 to cfg.device_indices[1], and so on. This tier-indirection allows the same inference code to run on different hardware configurations by only changing the initial config structure.

Route Computation with Device Context Switching

Before invoking kernels that operate on a tensor, switch to the target device using ds4_gpu_set_current_device(). This thin shim, declared at lines 194–200 of ds4_gpu_mgpu.h, translates the logical tier into the physical CUDA device ID stored in g_gpu[tier].device_id and calls cudaSetDevice internally.

ds4_gpu_set_current_device(0);   // Switch to GPU-0
ds4_gpu_matmul_q8_0_tensor(act0, NULL, 0, 0, 1024, 1024, act0, 1);

ds4_gpu_set_current_device(1);   // Switch to GPU-1
ds4_gpu_matmul_q8_0_tensor(act1, NULL, 0, 0, 1024, 1024, act1, 1);

Explicit device switching ensures that CUDA kernels and memory operations execute on the correct hardware. The engine respects these boundaries when placing layers, as demonstrated in test_engine_mgpu_runtime.c at lines 95–103, where the test builds a multi-tier configuration and verifies that logits match a single-GPU baseline.

Cross-Device Tensor Communication

When a layer consumes activations produced on a different GPU, use the asynchronous copy primitives declared at lines 29–38 of ds4_gpu_mgpu.h:

  • ds4_gpu_tensor_copy_xdev – Copies between any two tiers, automatically selecting the fastest path (NVLink peer-to-peer, PCIe P2P, or host bounce).
  • ds4_gpu_tensor_copy_xdev3 – Groups three copies for the "handoff" pattern used by the inference engine to pipeline activations across devices.
// Copy activation tensor from tier 0 to tier 1
ds4_gpu_tensor_copy_xdev(activations_tier1, activations_tier0, act_bytes);

These functions are implemented in ds4_cuda.c and handle stream synchronization internally, ensuring that downstream kernels wait for cross-device transfers to complete before reading the data.

End-to-End Multi-GPU Inference Example

The following self-contained program demonstrates the complete workflow: configuring two GPUs, initializing the runtime, allocating tier-specific tensors, executing kernels, and performing a cross-device copy. Save this as example_mgpu.c and compile it with the ds4 CUDA build target.

/* example_mgpu.c – minimal multi-GPU inference with tensor parallelism */
#include <stdio.h>
#include <stdlib.h>
#include "ds4.h"
#include "ds4_gpu_mgpu.h"

int main(void) {
    /* 1️⃣  Configure two GPUs (device 0 and 1) */
    ds4_gpu_config cfg = {0};
    cfg.n_gpus = 2;
    cfg.device_indices[0] = 0;          // CUDA device 0
    cfg.device_indices[1] = 1;          // CUDA device 1
    cfg.vram_bytes[0] = (size_t)4 << 30; // Reserve 4 GiB per GPU
    cfg.vram_bytes[1] = (size_t)4 << 30;

    if (ds4_gpu_init_multi(&cfg) != 0) {
        fprintf(stderr, "Failed to init multi-GPU runtime\n");
        return 1;
    }

    /* 2️⃣  Open the engine – same flags as single-GPU */
    ds4_engine_options opt = {0};
    opt.model_path = "./gguf/ds4flash.gguf";
    opt.backend    = DS4_BACKEND_CUDA;
    opt.n_threads  = 1;
    
    ds4_engine *engine = NULL;
    if (ds4_engine_open(&engine, &opt) != 0) {
        fprintf(stderr, "engine_open failed\n");
        return 1;
    }

    /* 3️⃣  Allocate tensors on specific tiers */
    uint64_t bytes = 1024 * sizeof(float);
    ds4_gpu_tensor *act0 = ds4_gpu_tensor_alloc_ptr_on(0, bytes);
    ds4_gpu_tensor *act1 = ds4_gpu_tensor_alloc_ptr_on(1, bytes);
    if (!act0 || !act1) return 1;

    /* 4️⃣  Fill with host data */
    float host_buf[1024];
    for (int i = 0; i < 1024; ++i) host_buf[i] = (float)i;
    ds4_gpu_tensor_write(act0, 0, host_buf, bytes);
    ds4_gpu_tensor_write(act1, 0, host_buf, bytes);

    /* 5️⃣  Run kernels on respective devices */
    ds4_gpu_set_current_device(0);
    ds4_gpu_matmul_q8_0_tensor(act0, NULL, 0, 0, 1024, 1024, act0, 1);

    ds4_gpu_set_current_device(1);
    ds4_gpu_matmul_q8_0_tensor(act1, NULL, 0, 0, 1024, 1024, act1, 1);

    /* 6️⃣  Cross-device copy: bring act0 to tier 1 */
    ds4_gpu_tensor_copy_xdev(act1, act0, bytes);

    /* 7️⃣  Cleanup */
    ds4_gpu_tensor_free_in_place(act0);
    ds4_gpu_tensor_free_in_place(act1);
    ds4_engine_close(engine);
    return 0;
}

This example follows the pattern used in tests/test_engine_mgpu_runtime.c, which validates that multi-GPU execution produces identical results to single-GPU baselines.

Summary

  • Configure available GPUs using ds4_gpu_config to map logical tiers to physical CUDA device IDs.
  • Initialize the multi-GPU runtime with ds4_gpu_init_multi(), which creates per-GPU streams and handles in ds4_cuda.c.
  • Allocate tensors on specific devices via ds4_gpu_tensor_alloc_ptr_on(), using tier indices rather than raw CUDA device numbers.
  • Route computation by calling ds4_gpu_set_current_device() before kernel launches to ensure operations execute on the correct GPU.
  • Transfer data between GPUs using ds4_gpu_tensor_copy_xdev(), which automatically selects peer-to-peer or host-bounce paths.
  • Reuse existing single-GPU kernels—no code duplication is required to support multi-GPU inference with CUDA tensor parallelism.

Frequently Asked Questions

What is the maximum number of GPUs supported by ds4 multi-GPU inference?

The ds4_gpu_config structure in ds4_gpu_mgpu.h uses a fixed-size array for device_indices, typically defined by DS4_GPU_MAX_GPUS (usually 8 or 16 depending on the build). You can set cfg.n_gpus up to this limit to utilize all available devices in your system.

Does ds4 automatically shard model layers across GPUs, or require manual placement?

The engine handles layer placement automatically. When you call ds4_engine_open(), the engine internally invokes ds4_gpu_set_current_device_fenced() to assign layers to tiers based on the configuration. As shown in test_engine_mgpu_runtime.c, you provide the GPU list via ds4_gpu_config, and the engine manages the distribution strategy without manual intervention.

How does ds4 handle tensor communication when peer-to-peer access is unavailable?

The ds4_gpu_tensor_copy_xdev() function, implemented in ds4_cuda.c, probes the CUDA topology at initialization and selects the fastest available transfer method. If NVLink or PCIe peer-to-peer is unavailable between two devices, it transparently falls back to pinned host memory bounce buffers, ensuring compatibility across diverse hardware configurations.

Can I use non-sequential CUDA device IDs for multi-GPU inference?

Yes. The device_indices array allows arbitrary mappings. For example, you can set cfg.device_indices[0] = 2 and cfg.device_indices[1] = 0 to use CUDA device 2 as logical tier 0 and device 0 as tier 1. This flexibility enables you to skip problematic GPUs or prioritize specific hardware topologies while keeping the inference code unchanged.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →