# How to Configure CUDA Tensor Parallelism on Multi-GPU Systems Like DGX Spark

> Configure CUDA tensor parallelism on multi-GPU systems like DGX Spark. Build DS4 with cuda-spark and launch processes using tensor parallel flags to distribute model layers.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-05

---

**Enable DS4's tensor parallelism mode by building with `make cuda-spark`, then launching multiple GPU-bound processes with `--tensor-parallel`, `--tp-rank`, and `--tp-world` flags to distribute model layers across devices.**

DS4 implements tensor parallelism (TP) for CUDA-enabled systems to split DeepSeek V4 model layers across multiple GPUs, with each device processing a disjoint subset of the network. This article explains the complete configuration process for multi-GPU deployments including DGX Spark nodes, based on the implementation in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c), and supporting files from the antirez/ds4 repository.

## Prerequisites and Hardware Validation

The CUDA tensor parallel implementation enforces strict requirements before initializing. In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) around line 16917, the engine validates that your system meets two conditions:

- **Even GPU count** — The number of CUDA devices must be even (2, 4, 8, etc.)
- **Even-expert model** — The DeepSeek V4 checkpoint must have a routable MoE layout with even expert distribution

If either check fails, DS4 aborts with explicit error messages. The validation logic rejects odd-world configurations because the layer partitioning scheme assigns consecutive half-ranges to each rank.

## Key Limitation: Resident Weights Required

Tensor parallelism **cannot** be combined with SSD streaming. As enforced at line 503 in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c):

```c
/* TP requires resident weights — each GPU holds a permanent slice */
if (Flags.tensor_parallel && Flags.ssd_streaming) {
    fprintf(stderr, "error: --tensor-parallel incompatible with --ssd\n");
    exit(1);
}

```

This restriction exists because the TP transport protocol assumes immediate memory access to local layer parameters. With SSD streaming, weight pages would fault across NVMe boundaries during cross-GPU synchronization, destroying the lock-step timing that TP requires.

## Building DS4 with CUDA Tensor Parallel Support

Use the dedicated DGX Spark target to compile with correct libraries and flags:

```bash

# Build CUDA-Spark binary (enables TP for CUDA)

make cuda-spark

```

This Makefile target automatically:
- Links CUDA runtime and driver libraries
- Defines `DS4_CUDA_TP` compile-time flag
- Excludes ROCm and Metal code paths (which explicitly reject TP)

## Launching Tensor Parallel Processes

DS4 uses **environment variables plus CLI flags** to establish the TP group. Each process binds to one GPU and communicates its rank/world size to peers through the transport layer in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c).

### Two-GPU Example

```bash

# GPU 0 (rank 0)

DS4_TP_RANK=0 DS4_TP_WORLD=2 ./ds4 \
    -m gguf/ds4f-q2.gguf \
    --tensor-parallel \
    --tp-rank 0 \
    --tp-world 2 \
    --gpu 0

# GPU 1 (rank 1) — run in separate terminal or job slot

DS4_TP_RANK=1 DS4_TP_WORLD=2 ./ds4 \
    -m gguf/ds4f-q2.gguf \
    --tensor-parallel \
    --tp-rank 1 \
    --tp-world 2 \
    --gpu 1

```

### Eight-GPU DGX Spark Configuration

For a full DGX Spark node (8 × L40S):

```bash
#!/bin/bash

# launch_tp_8gpu.sh — launch 8-way tensor parallel on DGX Spark

MODEL="gguf/ds4f-q2.gguf"
WORLD=8

for RANK in {0..7}; do
    CUDA_VISIBLE_DEVICES=$RANK \
    DS4_TP_RANK=$RANK \
    DS4_TP_WORLD=$WORLD \
    ./ds4 \
        -m "$MODEL" \
        --tensor-parallel \
        --tp-rank $RANK \
        --tp-world $WORLD \
        --gpu 0 \
        > logs/rank${RANK}.log 2>&1 &
done

wait

```

The `CUDA_VISIBLE_DEVICES` isolation ensures each process sees only its assigned GPU as device 0, while `--gpu 0` selects that virtual device.

## Understanding the TP Transport Layer

The lock-step protocol in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) and [`ds4_tp.h`](https://github.com/antirez/ds4/blob/main/ds4_tp.h) handles:

1. **Layer-home offset exchange** — At startup, each rank broadcasts its assigned layer range (computed in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) near line 17036)
2. **Gate sequence synchronization** — MoE routing decisions propagate across GPUs so each device knows which expert activations to expect
3. **Batch boundary alignment** — Timestep tokens wait at barriers until all ranks complete their forward pass

The protocol is **pure CUDA** — ROCm and Metal backends contain explicit rejection code. In `ds4_rocm.cu` at line 139, the engine prints *"tensor parallelism is CUDA-only"* and exits if TP flags are detected on non-NVIDIA hardware.

## Layer Partitioning Strategy

For a 2-GPU run, DS4 assigns the **lower half** of the layer stack to rank 0 and the **upper half** to rank 1. The home range calculation in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (line 17036) computes:

```c
uint32_t layers_per_gpu = model->num_layers / tp_world;
uint32_t my_first_layer = tp_rank * layers_per_gpu;
uint32_t my_last_layer = my_first_layer + layers_per_gpu - 1;

```

At boundaries between ranks, activation data shuttles through the TP transport. The comment block at line 15258 in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) describes how kernel launches remain local while cross-partition tensors traverse NVLink or PCIe:

```c
/* TP boundary: rank 0's final layer output becomes rank 1's input.
 * The transport copies the activation buffer via cudaMemcpyAsync
 * peer-to-peer, then signals the receiving rank's semaphore. */

```

## CLI Flag Reference

| Flag | Purpose |
|------|---------|
| `--tensor-parallel` | Activates TP mode and initializes transport |
| `--tp-rank N` | Logical rank of this process (0 ≤ N < tp-world) |
| `--tp-world M` | Total GPU count in TP group (must be even) |
| `--gpu ID` | CUDA device selection for this process |
| `-m <path>` | Path to GGUF model file (same for all ranks) |

Environment variables `DS4_TP_RANK` and `DS4_TP_WORLD` override CLI values when present, enabling easier orchestration with MPI or Slurm.

## Troubleshooting Common Failures

### "odd number of GPUs for tensor parallelism"

Your `--tp-world` value or detected GPU count is odd. Check with `nvidia-smi` and adjust to 2, 4, or 8.

### "tensor parallelism requires resident weights"

Remove `--ssd` or `--ssd-streaming` from your command. TP needs full model residency.

### "MoE expert count not divisible by TP world size"

Your checkpoint has an incompatible expert layout. Use even-expert DeepSeek V4 variants only.

### Peer access errors on multi-node systems

DS4 TP is designed for **single-node multi-GPU** via NVLink/PCIe. Cross-node tensor parallelism is not implemented as of the current [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) transport.

## Summary

- **Build** with `make cuda-spark` to enable CUDA tensor parallelism
- **Validate** even GPU count and even-expert model before launch
- **Launch** one process per GPU with matching `--tp-rank`/`--tp-world` values
- **Avoid** SSD streaming on TP runs — resident weights are mandatory
- **Monitor** layer-home assignments in logs to verify correct partitioning

The implementation in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c), and [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) provides a lightweight, lock-step transport that distributes inference across available CUDA devices without modifying the core engine loop.

## Frequently Asked Questions

### What GPU interconnect does DS4 tensor parallelism require?

DS4 TP uses CUDA peer-to-peer memory copies via `cudaMemcpyAsync`, which automatically utilizes NVLink when available and falls back to PCIe. No InfiniBand or network configuration is needed — the transport is single-node only as implemented in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c).

### Can I use tensor parallelism with quantized models?

Yes. The TP layer operates below quantization — it moves activation tensors (dequantized for computation) between GPUs. Build your GGUF with any supported quantization (Q2, Q4, Q8) and pass the same file path to all ranks.

### Why does DS4 reject odd numbers of GPUs for tensor parallelism?

The layer partitioning scheme in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) divides the model stack into consecutive half-ranges. Odd divisions would leave residual layers unassigned or require asymmetric splits that complicate the transport protocol. The validation at line 16917 enforces even-world sizes to maintain symmetric communication patterns.

### How do I verify that tensor parallelism is active during inference?

Enable verbose logging with `--verbose 2`. You should see messages from [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) showing "TP rank X/Y initialized" and periodic "TP sync" timestamps. The absence of these messages indicates the transport failed to initialize despite `--tensor-parallel` being present.