Configuring CUDA Tensor Parallelism for Multi-GPU NVIDIA Systems in ds4

ds4 implements tensor-parallel execution across two machines by synchronizing transformer gates via a transport layer in ds4_tp.c, but requires modifying the backend validation to support CUDA instead of the default Metal.

The ds4 inference engine supports distributed tensor parallelism to split transformer computation across two GPUs, though the current implementation targets Apple Silicon via Metal. Adapting this architecture for multi-GPU NVIDIA systems involves enabling the CUDA backend and relaxing the backend restriction enforced in the tensor-parallel validation layer.

Understanding ds4's Tensor Parallel Architecture

The Transport Layer in ds4_tp.c

The core transport logic resides in main/ds4_tp.c, where the protocol defines a control channel carrying fixed-size frame headers marked with DS4_TP_MAGIC. When both endpoints support RDMA, the high-throughput path transports gate vectors directly; otherwise, the system falls back to TCP for 16 KB gate messages. Frame definitions and socket helpers appear at lines 48-52 and 86-90.

Slab Layout and Memory Synchronization

A shared memory region (tp->slab) holds per-gate vectors for input, output, flags, and batch verification. The function tp_slab_layout (lines 334-447) calculates deterministic offsets, ensuring each gate's data lands in a known slot on both leader and worker machines. Slab-size computation routines are located at lines 24-32 and 34-48.

Why CUDA Support Requires Code Changes

By default, ds4_tp restricts tensor parallelism to the Metal backend. The validation routine in ds4_tp_validate_engine_options explicitly rejects non-Metal configurations at lines 498-500:

if (opt->backend != DS4_BACKEND_METAL) {
    tp_set_err(err, errlen, "tensor parallelism requires the Metal backend");
    return 0;
}

To enable NVIDIA GPU support, you must modify this check to accept DS4_BACKEND_CUDA.

Step-by-Step Configuration Guide

Build ds4 with the CUDA Backend

Compile the project using the CUDA backend by passing the appropriate flag:

make clean && make backend=cuda

Alternatively, set the DS4_BACKEND environment variable to CUDA at runtime.

Modify the Backend Validation Check

Edit main/ds4_tp.c to relax the backend restriction in ds4_tp_validate_engine_options. Change the validation logic near line 498 to accept both Metal and CUDA backends, allowing the tensor-parallel transport layer to initialize with NVIDIA hardware.

Configure GPU Arguments

Ensure main/ds4_gpu_args.c receives correct CUDA device IDs via the --gpu flag. This file parses device selection arguments and must identify the appropriate NVIDIA GPUs on both nodes.

Launch Leader and Worker Nodes

Execute the binary on each machine with matching model files and tensor-parallel options. The leader listens for connections while the worker connects to the coordinator:


# Leader (Node A)

DS4_BACKEND=CUDA ./ds4 \
    --tensor-parallel \
    --role leader \
    --listen 0.0.0.0 12345 \
    --transport auto \
    --model mymodel.gguf

# Worker (Node B)

DS4_BACKEND=CUDA ./ds4 \
    --tensor-parallel \
    --role worker \
    --coordinator <leader-ip> 12345 \
    --transport auto \
    --model mymodel.gguf

For systems with NVLink interconnects, force RDMA mode using --transport rdma and specify the device via --rdma-device and --rdma-gid-index. The RDMA path leverages functions prefixed with tp_rdma_* in ds4_tp.c for high-bandwidth gate exchange between NVIDIA GPUs.

Summary

  • ds4 implements tensor parallelism via main/ds4_tp.c, using a slab-based memory layout and RDMA/TCP transport for gate synchronization.
  • CUDA support requires modifying ds4_tp_validate_engine_options to accept DS4_BACKEND_CUDA alongside the default Metal backend.
  • Build configuration uses backend=cuda or the DS4_BACKEND environment variable to target NVIDIA hardware.
  • CLI options --tensor-parallel, --role, and --transport control the distributed execution mode across two nodes.
  • RDMA optimization over NVLink maximizes throughput when available, falling back to TCP automatically if RDMA is unavailable.

Frequently Asked Questions

Does ds4 support tensor parallelism on NVIDIA GPUs out of the box?

No. The current implementation in main/ds4_tp.c explicitly validates against DS4_BACKEND_METAL at lines 498-500. You must modify ds4_tp_validate_engine_options to accept DS4_BACKEND_CUDA and rebuild with the CUDA backend enabled.

What is the difference between the RDMA and TCP transport modes?

RDMA provides high-throughput, low-latency direct memory access between nodes using InfiniBand or NVLink, while TCP transports 16 KB gate messages over standard network sockets. Use --transport auto to prefer RDMA when available, or force RDMA with --transport rdma for NVLink-connected NVIDIA GPUs.

How does the slab layout ensure synchronization between GPUs?

The tp_slab_layout function in main/ds4_tp.c (lines 334-447) calculates deterministic offsets for input vectors, output vectors, and flags within a shared memory region. This deterministic layout ensures that gate data occupies identical memory slots on both the leader and worker, enabling consistent state across the tensor-parallel split.

Can I use tensor parallelism with more than two GPUs?

The current ds4_tp implementation is designed for exactly two machines (leader and worker) synchronizing every transformer gate across a pair of GPUs. Extending beyond two GPUs would require architectural changes to the transport layer and slab allocation logic in ds4_tp.c.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →