# Configuring CUDA Tensor Parallelism for Multi-GPU NVIDIA Systems in ds4

> Learn to configure CUDA tensor parallelism for multi-GPU NVIDIA systems in ds4. Explore synchronizing transformer gates and modifying backend validation for enhanced performance.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**ds4 implements tensor-parallel execution across two machines by synchronizing transformer gates via a transport layer in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c), but requires modifying the backend validation to support CUDA instead of the default Metal.**

The ds4 inference engine supports distributed tensor parallelism to split transformer computation across two GPUs, though the current implementation targets Apple Silicon via Metal. Adapting this architecture for multi-GPU NVIDIA systems involves enabling the CUDA backend and relaxing the backend restriction enforced in the tensor-parallel validation layer.

## Understanding ds4's Tensor Parallel Architecture

### The Transport Layer in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c)

The core transport logic resides in [`main/ds4_tp.c`](https://github.com/antirez/ds4/blob/main/main/ds4_tp.c), where the protocol defines a control channel carrying fixed-size frame headers marked with `DS4_TP_MAGIC`. When both endpoints support RDMA, the high-throughput path transports gate vectors directly; otherwise, the system falls back to TCP for 16 KB gate messages. Frame definitions and socket helpers appear at lines 48-52 and 86-90.

### Slab Layout and Memory Synchronization

A shared memory region (`tp->slab`) holds per-gate vectors for input, output, flags, and batch verification. The function `tp_slab_layout` (lines 334-447) calculates deterministic offsets, ensuring each gate's data lands in a known slot on both leader and worker machines. Slab-size computation routines are located at lines 24-32 and 34-48.

## Why CUDA Support Requires Code Changes

By default, `ds4_tp` restricts tensor parallelism to the Metal backend. The validation routine in `ds4_tp_validate_engine_options` explicitly rejects non-Metal configurations at lines 498-500:

```c
if (opt->backend != DS4_BACKEND_METAL) {
    tp_set_err(err, errlen, "tensor parallelism requires the Metal backend");
    return 0;
}

```

To enable NVIDIA GPU support, you must modify this check to accept `DS4_BACKEND_CUDA`.

## Step-by-Step Configuration Guide

### Build ds4 with the CUDA Backend

Compile the project using the CUDA backend by passing the appropriate flag:

```bash
make clean && make backend=cuda

```

Alternatively, set the `DS4_BACKEND` environment variable to `CUDA` at runtime.

### Modify the Backend Validation Check

Edit [`main/ds4_tp.c`](https://github.com/antirez/ds4/blob/main/main/ds4_tp.c) to relax the backend restriction in `ds4_tp_validate_engine_options`. Change the validation logic near line 498 to accept both Metal and CUDA backends, allowing the tensor-parallel transport layer to initialize with NVIDIA hardware.

### Configure GPU Arguments

Ensure [`main/ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/main/ds4_gpu_args.c) receives correct CUDA device IDs via the `--gpu` flag. This file parses device selection arguments and must identify the appropriate NVIDIA GPUs on both nodes.

### Launch Leader and Worker Nodes

Execute the binary on each machine with matching model files and tensor-parallel options. The leader listens for connections while the worker connects to the coordinator:

```bash

# Leader (Node A)

DS4_BACKEND=CUDA ./ds4 \
    --tensor-parallel \
    --role leader \
    --listen 0.0.0.0 12345 \
    --transport auto \
    --model mymodel.gguf

# Worker (Node B)

DS4_BACKEND=CUDA ./ds4 \
    --tensor-parallel \
    --role worker \
    --coordinator <leader-ip> 12345 \
    --transport auto \
    --model mymodel.gguf

```

### Optimize with RDMA over NVLink (Optional)

For systems with NVLink interconnects, force RDMA mode using `--transport rdma` and specify the device via `--rdma-device` and `--rdma-gid-index`. The RDMA path leverages functions prefixed with `tp_rdma_*` in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) for high-bandwidth gate exchange between NVIDIA GPUs.

## Summary

- **ds4** implements tensor parallelism via [`main/ds4_tp.c`](https://github.com/antirez/ds4/blob/main/main/ds4_tp.c), using a slab-based memory layout and RDMA/TCP transport for gate synchronization.
- **CUDA support** requires modifying `ds4_tp_validate_engine_options` to accept `DS4_BACKEND_CUDA` alongside the default Metal backend.
- **Build configuration** uses `backend=cuda` or the `DS4_BACKEND` environment variable to target NVIDIA hardware.
- **CLI options** `--tensor-parallel`, `--role`, and `--transport` control the distributed execution mode across two nodes.
- **RDMA optimization** over NVLink maximizes throughput when available, falling back to TCP automatically if RDMA is unavailable.

## Frequently Asked Questions

### Does ds4 support tensor parallelism on NVIDIA GPUs out of the box?

No. The current implementation in [`main/ds4_tp.c`](https://github.com/antirez/ds4/blob/main/main/ds4_tp.c) explicitly validates against `DS4_BACKEND_METAL` at lines 498-500. You must modify `ds4_tp_validate_engine_options` to accept `DS4_BACKEND_CUDA` and rebuild with the CUDA backend enabled.

### What is the difference between the RDMA and TCP transport modes?

RDMA provides high-throughput, low-latency direct memory access between nodes using InfiniBand or NVLink, while TCP transports 16 KB gate messages over standard network sockets. Use `--transport auto` to prefer RDMA when available, or force RDMA with `--transport rdma` for NVLink-connected NVIDIA GPUs.

### How does the slab layout ensure synchronization between GPUs?

The `tp_slab_layout` function in [`main/ds4_tp.c`](https://github.com/antirez/ds4/blob/main/main/ds4_tp.c) (lines 334-447) calculates deterministic offsets for input vectors, output vectors, and flags within a shared memory region. This deterministic layout ensures that gate data occupies identical memory slots on both the leader and worker, enabling consistent state across the tensor-parallel split.

### Can I use tensor parallelism with more than two GPUs?

The current `ds4_tp` implementation is designed for exactly two machines (leader and worker) synchronizing every transformer gate across a pair of GPUs. Extending beyond two GPUs would require architectural changes to the transport layer and slab allocation logic in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c).