How to Configure CUDA Tensor Parallelism on Multi-GPU Systems Like DGX Spark
Enable DS4's tensor parallelism mode by building with make cuda-spark, then launching multiple GPU-bound processes with --tensor-parallel, --tp-rank, and --tp-world flags to distribute model layers across devices.
DS4 implements tensor parallelism (TP) for CUDA-enabled systems to split DeepSeek V4 model layers across multiple GPUs, with each device processing a disjoint subset of the network. This article explains the complete configuration process for multi-GPU deployments including DGX Spark nodes, based on the implementation in ds4.c, ds4_tp.c, and supporting files from the antirez/ds4 repository.
Prerequisites and Hardware Validation
The CUDA tensor parallel implementation enforces strict requirements before initializing. In ds4.c around line 16917, the engine validates that your system meets two conditions:
- Even GPU count — The number of CUDA devices must be even (2, 4, 8, etc.)
- Even-expert model — The DeepSeek V4 checkpoint must have a routable MoE layout with even expert distribution
If either check fails, DS4 aborts with explicit error messages. The validation logic rejects odd-world configurations because the layer partitioning scheme assigns consecutive half-ranges to each rank.
Key Limitation: Resident Weights Required
Tensor parallelism cannot be combined with SSD streaming. As enforced at line 503 in ds4.c:
/* TP requires resident weights — each GPU holds a permanent slice */
if (Flags.tensor_parallel && Flags.ssd_streaming) {
fprintf(stderr, "error: --tensor-parallel incompatible with --ssd\n");
exit(1);
}
This restriction exists because the TP transport protocol assumes immediate memory access to local layer parameters. With SSD streaming, weight pages would fault across NVMe boundaries during cross-GPU synchronization, destroying the lock-step timing that TP requires.
Building DS4 with CUDA Tensor Parallel Support
Use the dedicated DGX Spark target to compile with correct libraries and flags:
# Build CUDA-Spark binary (enables TP for CUDA)
make cuda-spark
This Makefile target automatically:
- Links CUDA runtime and driver libraries
- Defines
DS4_CUDA_TPcompile-time flag - Excludes ROCm and Metal code paths (which explicitly reject TP)
Launching Tensor Parallel Processes
DS4 uses environment variables plus CLI flags to establish the TP group. Each process binds to one GPU and communicates its rank/world size to peers through the transport layer in ds4_tp.c.
Two-GPU Example
# GPU 0 (rank 0)
DS4_TP_RANK=0 DS4_TP_WORLD=2 ./ds4 \
-m gguf/ds4f-q2.gguf \
--tensor-parallel \
--tp-rank 0 \
--tp-world 2 \
--gpu 0
# GPU 1 (rank 1) — run in separate terminal or job slot
DS4_TP_RANK=1 DS4_TP_WORLD=2 ./ds4 \
-m gguf/ds4f-q2.gguf \
--tensor-parallel \
--tp-rank 1 \
--tp-world 2 \
--gpu 1
Eight-GPU DGX Spark Configuration
For a full DGX Spark node (8 × L40S):
#!/bin/bash
# launch_tp_8gpu.sh — launch 8-way tensor parallel on DGX Spark
MODEL="gguf/ds4f-q2.gguf"
WORLD=8
for RANK in {0..7}; do
CUDA_VISIBLE_DEVICES=$RANK \
DS4_TP_RANK=$RANK \
DS4_TP_WORLD=$WORLD \
./ds4 \
-m "$MODEL" \
--tensor-parallel \
--tp-rank $RANK \
--tp-world $WORLD \
--gpu 0 \
> logs/rank${RANK}.log 2>&1 &
done
wait
The CUDA_VISIBLE_DEVICES isolation ensures each process sees only its assigned GPU as device 0, while --gpu 0 selects that virtual device.
Understanding the TP Transport Layer
The lock-step protocol in ds4_tp.c and ds4_tp.h handles:
- Layer-home offset exchange — At startup, each rank broadcasts its assigned layer range (computed in
ds4.cnear line 17036) - Gate sequence synchronization — MoE routing decisions propagate across GPUs so each device knows which expert activations to expect
- Batch boundary alignment — Timestep tokens wait at barriers until all ranks complete their forward pass
The protocol is pure CUDA — ROCm and Metal backends contain explicit rejection code. In ds4_rocm.cu at line 139, the engine prints "tensor parallelism is CUDA-only" and exits if TP flags are detected on non-NVIDIA hardware.
Layer Partitioning Strategy
For a 2-GPU run, DS4 assigns the lower half of the layer stack to rank 0 and the upper half to rank 1. The home range calculation in ds4.c (line 17036) computes:
uint32_t layers_per_gpu = model->num_layers / tp_world;
uint32_t my_first_layer = tp_rank * layers_per_gpu;
uint32_t my_last_layer = my_first_layer + layers_per_gpu - 1;
At boundaries between ranks, activation data shuttles through the TP transport. The comment block at line 15258 in ds4.c describes how kernel launches remain local while cross-partition tensors traverse NVLink or PCIe:
/* TP boundary: rank 0's final layer output becomes rank 1's input.
* The transport copies the activation buffer via cudaMemcpyAsync
* peer-to-peer, then signals the receiving rank's semaphore. */
CLI Flag Reference
| Flag | Purpose |
|---|---|
--tensor-parallel |
Activates TP mode and initializes transport |
--tp-rank N |
Logical rank of this process (0 ≤ N < tp-world) |
--tp-world M |
Total GPU count in TP group (must be even) |
--gpu ID |
CUDA device selection for this process |
-m <path> |
Path to GGUF model file (same for all ranks) |
Environment variables DS4_TP_RANK and DS4_TP_WORLD override CLI values when present, enabling easier orchestration with MPI or Slurm.
Troubleshooting Common Failures
"odd number of GPUs for tensor parallelism"
Your --tp-world value or detected GPU count is odd. Check with nvidia-smi and adjust to 2, 4, or 8.
"tensor parallelism requires resident weights"
Remove --ssd or --ssd-streaming from your command. TP needs full model residency.
"MoE expert count not divisible by TP world size"
Your checkpoint has an incompatible expert layout. Use even-expert DeepSeek V4 variants only.
Peer access errors on multi-node systems
DS4 TP is designed for single-node multi-GPU via NVLink/PCIe. Cross-node tensor parallelism is not implemented as of the current ds4_tp.c transport.
Summary
- Build with
make cuda-sparkto enable CUDA tensor parallelism - Validate even GPU count and even-expert model before launch
- Launch one process per GPU with matching
--tp-rank/--tp-worldvalues - Avoid SSD streaming on TP runs — resident weights are mandatory
- Monitor layer-home assignments in logs to verify correct partitioning
The implementation in ds4.c, ds4_tp.c, and ds4_gpu_args.c provides a lightweight, lock-step transport that distributes inference across available CUDA devices without modifying the core engine loop.
Frequently Asked Questions
What GPU interconnect does DS4 tensor parallelism require?
DS4 TP uses CUDA peer-to-peer memory copies via cudaMemcpyAsync, which automatically utilizes NVLink when available and falls back to PCIe. No InfiniBand or network configuration is needed — the transport is single-node only as implemented in ds4_tp.c.
Can I use tensor parallelism with quantized models?
Yes. The TP layer operates below quantization — it moves activation tensors (dequantized for computation) between GPUs. Build your GGUF with any supported quantization (Q2, Q4, Q8) and pass the same file path to all ranks.
Why does DS4 reject odd numbers of GPUs for tensor parallelism?
The layer partitioning scheme in ds4.c divides the model stack into consecutive half-ranges. Odd divisions would leave residual layers unassigned or require asymmetric splits that complicate the transport protocol. The validation at line 16917 enforces even-world sizes to maintain symmetric communication patterns.
How do I verify that tensor parallelism is active during inference?
Enable verbose logging with --verbose 2. You should see messages from ds4_tp.c showing "TP rank X/Y initialized" and periodic "TP sync" timestamps. The absence of these messages indicates the transport failed to initialize despite --tensor-parallel being present.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →