# How to Implement Multi-GPU Inference with CUDA Tensor Parallelism in ds4

> Implement multi-GPU inference with CUDA tensor parallelism in ds4. Learn how ds4 splits model tensors across devices for efficient parallel processing without duplicating kernels.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-07

---

**The ds4 engine enables multi-GPU inference by splitting model tensors across CUDA devices using a tier-based configuration system, allowing tensor parallelism without duplicating kernel implementations.**

The antirez/ds4 repository provides a native CUDA backend that scales transformer inference across multiple GPUs through tensor parallelism. By leveraging the multi-GPU plumbing layer in [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h) and its implementation in [`ds4_cuda.c`](https://github.com/antirez/ds4/blob/main/ds4_cuda.c), developers can distribute model weights and activations across devices while reusing existing single-GPU kernels. This architecture keeps inference code device-agnostic—whether running on one GPU or eight, only the configuration layer changes.

## Configure the Multi-GPU Environment

Multi-GPU inference starts with a `ds4_gpu_config` structure that enumerates the physical CUDA devices and optionally reserves per-GPU VRAM budgets. Defined in [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h) at lines 71–84, this structure maps logical tiers to hardware devices.

```c
ds4_gpu_config cfg = {0};
cfg.n_gpus = 2;                     // Number of GPUs to use
cfg.device_indices[0] = 0;          // Logical tier 0 → CUDA device 0
cfg.device_indices[1] = 1;          // Logical tier 1 → CUDA device 1
cfg.vram_bytes[0] = (size_t)4 << 30; // Optional: 4 GiB reserved on GPU 0
cfg.vram_bytes[1] = (size_t)4 << 30; // Optional: 4 GiB reserved on GPU 1

```

The `device_indices` array allows arbitrary hardware topologies—you can skip device IDs or reorder them to match your physical layout. The `n_gpus` field determines how many entries in the array are active.

## Initialize the Runtime

After configuration, call `ds4_gpu_init_multi()` to bootstrap the multi-GPU runtime. This function, declared at lines 105–106 of [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h), creates a global `g_gpu[]` array that holds per-GPU CUDA streams, cuBLAS handles, and bookkeeping buffers.

```c
if (ds4_gpu_init_multi(&cfg) != 0) {
    fprintf(stderr, "Failed to initialize multi-GPU runtime\n");
    return 1;
}

```

Once initialized, the runtime manages device contexts internally. You do not need to call `cudaSetDevice` manually—the ds4 abstraction layer handles physical device mapping based on the logical tiers defined in your configuration.

## Allocate Tensors on Specific Devices

With the runtime active, allocate tensors on specific logical tiers using `ds4_gpu_tensor_alloc_ptr_on()`. This function, declared at lines 18–19 of [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h), takes a tier index and byte count, returning a `ds4_gpu_tensor` pointer bound to the corresponding physical GPU.

```c
uint64_t weight_bytes = 512 * 1024 * 1024;  // 512 MiB
uint64_t act_bytes = 1024 * sizeof(float);

ds4_gpu_tensor *weights_tier0 = ds4_gpu_tensor_alloc_ptr_on(0, weight_bytes);
ds4_gpu_tensor *activations_tier1 = ds4_gpu_tensor_alloc_ptr_on(1, act_bytes);

```

Tier 0 corresponds to `cfg.device_indices[0]`, tier 1 to `cfg.device_indices[1]`, and so on. This tier-indirection allows the same inference code to run on different hardware configurations by only changing the initial config structure.

## Route Computation with Device Context Switching

Before invoking kernels that operate on a tensor, switch to the target device using `ds4_gpu_set_current_device()`. This thin shim, declared at lines 194–200 of [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h), translates the logical tier into the physical CUDA device ID stored in `g_gpu[tier].device_id` and calls `cudaSetDevice` internally.

```c
ds4_gpu_set_current_device(0);   // Switch to GPU-0
ds4_gpu_matmul_q8_0_tensor(act0, NULL, 0, 0, 1024, 1024, act0, 1);

ds4_gpu_set_current_device(1);   // Switch to GPU-1
ds4_gpu_matmul_q8_0_tensor(act1, NULL, 0, 0, 1024, 1024, act1, 1);

```

Explicit device switching ensures that CUDA kernels and memory operations execute on the correct hardware. The engine respects these boundaries when placing layers, as demonstrated in [`test_engine_mgpu_runtime.c`](https://github.com/antirez/ds4/blob/main/test_engine_mgpu_runtime.c) at lines 95–103, where the test builds a multi-tier configuration and verifies that logits match a single-GPU baseline.

## Cross-Device Tensor Communication

When a layer consumes activations produced on a different GPU, use the asynchronous copy primitives declared at lines 29–38 of [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h):

- **`ds4_gpu_tensor_copy_xdev`** – Copies between any two tiers, automatically selecting the fastest path (NVLink peer-to-peer, PCIe P2P, or host bounce).
- **`ds4_gpu_tensor_copy_xdev3`** – Groups three copies for the "handoff" pattern used by the inference engine to pipeline activations across devices.

```c
// Copy activation tensor from tier 0 to tier 1
ds4_gpu_tensor_copy_xdev(activations_tier1, activations_tier0, act_bytes);

```

These functions are implemented in [`ds4_cuda.c`](https://github.com/antirez/ds4/blob/main/ds4_cuda.c) and handle stream synchronization internally, ensuring that downstream kernels wait for cross-device transfers to complete before reading the data.

## End-to-End Multi-GPU Inference Example

The following self-contained program demonstrates the complete workflow: configuring two GPUs, initializing the runtime, allocating tier-specific tensors, executing kernels, and performing a cross-device copy. Save this as [`example_mgpu.c`](https://github.com/antirez/ds4/blob/main/example_mgpu.c) and compile it with the ds4 CUDA build target.

```c
/* example_mgpu.c – minimal multi-GPU inference with tensor parallelism */
#include <stdio.h>
#include <stdlib.h>
#include "ds4.h"
#include "ds4_gpu_mgpu.h"

int main(void) {
    /* 1️⃣  Configure two GPUs (device 0 and 1) */
    ds4_gpu_config cfg = {0};
    cfg.n_gpus = 2;
    cfg.device_indices[0] = 0;          // CUDA device 0
    cfg.device_indices[1] = 1;          // CUDA device 1
    cfg.vram_bytes[0] = (size_t)4 << 30; // Reserve 4 GiB per GPU
    cfg.vram_bytes[1] = (size_t)4 << 30;

    if (ds4_gpu_init_multi(&cfg) != 0) {
        fprintf(stderr, "Failed to init multi-GPU runtime\n");
        return 1;
    }

    /* 2️⃣  Open the engine – same flags as single-GPU */
    ds4_engine_options opt = {0};
    opt.model_path = "./gguf/ds4flash.gguf";
    opt.backend    = DS4_BACKEND_CUDA;
    opt.n_threads  = 1;
    
    ds4_engine *engine = NULL;
    if (ds4_engine_open(&engine, &opt) != 0) {
        fprintf(stderr, "engine_open failed\n");
        return 1;
    }

    /* 3️⃣  Allocate tensors on specific tiers */
    uint64_t bytes = 1024 * sizeof(float);
    ds4_gpu_tensor *act0 = ds4_gpu_tensor_alloc_ptr_on(0, bytes);
    ds4_gpu_tensor *act1 = ds4_gpu_tensor_alloc_ptr_on(1, bytes);
    if (!act0 || !act1) return 1;

    /* 4️⃣  Fill with host data */
    float host_buf[1024];
    for (int i = 0; i < 1024; ++i) host_buf[i] = (float)i;
    ds4_gpu_tensor_write(act0, 0, host_buf, bytes);
    ds4_gpu_tensor_write(act1, 0, host_buf, bytes);

    /* 5️⃣  Run kernels on respective devices */
    ds4_gpu_set_current_device(0);
    ds4_gpu_matmul_q8_0_tensor(act0, NULL, 0, 0, 1024, 1024, act0, 1);

    ds4_gpu_set_current_device(1);
    ds4_gpu_matmul_q8_0_tensor(act1, NULL, 0, 0, 1024, 1024, act1, 1);

    /* 6️⃣  Cross-device copy: bring act0 to tier 1 */
    ds4_gpu_tensor_copy_xdev(act1, act0, bytes);

    /* 7️⃣  Cleanup */
    ds4_gpu_tensor_free_in_place(act0);
    ds4_gpu_tensor_free_in_place(act1);
    ds4_engine_close(engine);
    return 0;
}

```

This example follows the pattern used in [`tests/test_engine_mgpu_runtime.c`](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_runtime.c), which validates that multi-GPU execution produces identical results to single-GPU baselines.

## Summary

- **Configure** available GPUs using `ds4_gpu_config` to map logical tiers to physical CUDA device IDs.
- **Initialize** the multi-GPU runtime with `ds4_gpu_init_multi()`, which creates per-GPU streams and handles in [`ds4_cuda.c`](https://github.com/antirez/ds4/blob/main/ds4_cuda.c).
- **Allocate** tensors on specific devices via `ds4_gpu_tensor_alloc_ptr_on()`, using tier indices rather than raw CUDA device numbers.
- **Route** computation by calling `ds4_gpu_set_current_device()` before kernel launches to ensure operations execute on the correct GPU.
- **Transfer** data between GPUs using `ds4_gpu_tensor_copy_xdev()`, which automatically selects peer-to-peer or host-bounce paths.
- **Reuse** existing single-GPU kernels—no code duplication is required to support multi-GPU inference with CUDA tensor parallelism.

## Frequently Asked Questions

### What is the maximum number of GPUs supported by ds4 multi-GPU inference?

The `ds4_gpu_config` structure in [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h) uses a fixed-size array for `device_indices`, typically defined by `DS4_GPU_MAX_GPUS` (usually 8 or 16 depending on the build). You can set `cfg.n_gpus` up to this limit to utilize all available devices in your system.

### Does ds4 automatically shard model layers across GPUs, or require manual placement?

The engine handles layer placement automatically. When you call `ds4_engine_open()`, the engine internally invokes `ds4_gpu_set_current_device_fenced()` to assign layers to tiers based on the configuration. As shown in [`test_engine_mgpu_runtime.c`](https://github.com/antirez/ds4/blob/main/test_engine_mgpu_runtime.c), you provide the GPU list via `ds4_gpu_config`, and the engine manages the distribution strategy without manual intervention.

### How does ds4 handle tensor communication when peer-to-peer access is unavailable?

The `ds4_gpu_tensor_copy_xdev()` function, implemented in [`ds4_cuda.c`](https://github.com/antirez/ds4/blob/main/ds4_cuda.c), probes the CUDA topology at initialization and selects the fastest available transfer method. If NVLink or PCIe peer-to-peer is unavailable between two devices, it transparently falls back to pinned host memory bounce buffers, ensuring compatibility across diverse hardware configurations.

### Can I use non-sequential CUDA device IDs for multi-GPU inference?

Yes. The `device_indices` array allows arbitrary mappings. For example, you can set `cfg.device_indices[0] = 2` and `cfg.device_indices[1] = 0` to use CUDA device 2 as logical tier 0 and device 0 as tier 1. This flexibility enables you to skip problematic GPUs or prioritize specific hardware topologies while keeping the inference code unchanged.