How to Configure CUDA Multi-GPU Tensor Parallelism in ds4
To configure CUDA multi-GPU tensor parallelism in ds4, compile with CUDA support, ensure an even number of GPUs, use a DeepSeek-4 model with even expert counts, disable conflicting flags like --ssd-streaming, and launch with the --tensor-parallel flag.
Tensor parallelism in the antirez/ds4 repository enables a single DeepSeek-4 model to be split across multiple CUDA GPUs, allowing each device to hold a portion of the model’s layers while cooperatively performing pre-fill and decode operations. This configuration requires satisfying specific runtime checks and CLI options before the engine initializes a multi-GPU tensor-parallel session. Below is a step-by-step guide referencing the exact source locations that enforce each requirement.
Prerequisites and Build Configuration
Compile with CUDA Support
Before enabling tensor parallelism, you must build ds4 with CUDA support enabled. The default Makefile automatically detects the CUDA toolkit and compiles the necessary back-end components.
The CUDA implementation resides in ds4.c and ds4_gpu.c. Without these components, the tensor-parallel branch is never compiled into the binary, and multi-GPU placement logic remains inaccessible.
Verify GPU Count Requirements
Ds4 requires an even number of available GPUs to perform tensor parallelism. The runtime checks the global variable g_n_gpus (the detected GPU count) and validates that an even placement exists.
In ds4.c at lines 16921‑16927, the engine aborts with an error if the GPU count is odd, ensuring symmetric tier pairing is possible.
Model and Compatibility Validation
Supported Model Architectures
Tensor parallelism is restricted to DeepSeek-4 models that contain an even number of experts. The constant DS4_N_EXPERT must be a multiple of 2 for the model to qualify for splitting across tiers.
At lines 16929‑16935 in ds4.c, the engine validates both the model family and expert parity before proceeding with tensor-parallel initialization.
Disable Incompatible Runtime Flags
Certain features conflict with tensor parallelism and must be explicitly disabled. You cannot enable SSD streaming (--ssd-streaming) because tensor parallelism requires resident weights in GPU memory. Additionally, the distributed role (--role distributed) is incompatible with local tensor-parallel configurations.
The validation logic at lines 499‑507 in ds4.c explicitly reports these incompatibilities and prevents engine startup if conflicting options are detected.
Enabling Tensor Parallelism
CLI Configuration
To activate tensor parallelism, pass the --tensor-parallel flag (or the short form -t) when starting the server or CLI. This option is registered in ds4_help.c at lines 246‑247.
./ds4_server -m /path/to/deepseek4.q4_0.gguf -t
If you previously enabled SSD streaming in your configuration file, explicitly disable it:
./ds4_server -m model.gguf --no-ssd-streaming -t
Internal Tier Pairing Logic
Once the CLI flags are validated, ds4 internally pairs GPUs into lower-half and upper-half tiers. The engine maps tier 0 through half-1 with tier half through g_n_gpus-1. The system will reject any placement that already utilizes an upper-half tier to prevent resource conflicts.
This pairing logic and its associated guard clauses are implemented in ds4.c at lines 16937‑16950.
Per-Tier Tensor Allocation
For every tier participating in the tensor-parallel group, ds4 allocates necessary buffers including KV-cache, hidden-state, and attention tensors. The function ds4_gpu_tensor_alloc_ptr_on handles these allocations.
The tier-wise allocation loop appears in ds4.c at lines 17056‑17078, ensuring each GPU receives its partitioned portion of the model state.
Advanced Configuration Options
Environment Variables for KV-Cache Tuning
For extremely long contexts or specific performance requirements, you can tune kernel fusion behavior using environment variables:
DS4_CUDA_NO_QKV_PAIR: Disables QKV-pair kernel optimizationDS4_CUDA_TP_ATTN_OUT_HC_FUSE: Enables attention-output half-cache fusion
These variables are consulted during initialization in ds4.c around lines 16998‑17004.
export DS4_CUDA_TP_ATTN_OUT_HC_FUSE=1
export DS4_CUDA_NO_QKV_PAIR=1
./ds4_server -m model.gguf -t
Engine Initialization
After all checks pass, the function engine_classify_multi_tier() creates a ds4_engine instance configured for tensor-parallel execution. This launch sequence occurs in ds4.c at lines 56312‑56320.
Usage Examples
Basic Server Launch
Launch a tensor-parallel server on a Linux machine with two CUDA GPUs:
./ds4_server -m /path/to/deepseek4.q4_0.gguf -t
Python Client Interaction
Once the server is running, interact with it via the HTTP API:
import requests
import json
url = "http://localhost:8080/v1/chat/completions"
payload = {
"model": "deepseek4",
"messages": [{"role": "user", "content": "Explain tensor parallelism"}],
"max_tokens": 256,
}
resp = requests.post(url, json=payload)
print(json.dumps(resp.json(), indent=2))
Key Source Files
The following files implement the tensor-parallel configuration and runtime logic:
ds4.c: Core engine containing runtime checks for GPU count, model compatibility, tier pairing, and per-tier tensor allocation.ds4_gpu.h/ds4_gpu.c: CUDA GPU abstraction layer definingds4_gpu_tensor_*APIs for tier-wise memory management.ds4_tp.c: Tensor-parallelism specific helpers and error message generation (e.g.,tp_set_err).ds4_help.c: CLI option registration, including the--tensor-parallelflag.tests/test_engine_mgpu_placement.c: Unit tests validating correct tier pairing and placement logic for multi-GPU setups.
Summary
- Build requirement: Compile ds4 with CUDA support via the default Makefile to include
ds4.candds4_gpu.cback-ends. - Hardware requirement: Provide an even number of GPUs;
g_n_gpusvalidation occurs at lines 16921‑16927. - Model requirement: Use DeepSeek-4 models with even expert counts (
DS4_N_EXPERT % 2 == 0), verified at lines 16929‑16935. - Flag conflicts: Disable
--ssd-streamingand avoid--role distributedto prevent launch aborts at lines 499‑507. - Activation: Pass
--tensor-parallelor-t, parsed inds4_help.cat lines 246‑247. - Allocation: Per-tier tensors are allocated via
ds4_gpu_tensor_alloc_ptr_onin the loop at lines 17056‑17078. - Launch: The engine initializes through
engine_classify_multi_tier()at lines 56312‑56320.
Frequently Asked Questions
What happens if I try to use tensor parallelism with an odd number of GPUs?
The engine will detect the asymmetry during initialization and abort with a clear error message. Specifically, the check at ds4.c lines 16921‑16927 validates that g_n_gpus is even, as tensor parallelism requires symmetric pairing between lower and upper tier groups.
Can I use SSD streaming with tensor parallelism enabled?
No. SSD streaming requires weights to reside on disk with partial resident memory, while tensor parallelism requires full resident weights on GPU tiers. The incompatibility check at ds4.c lines 499‑507 explicitly prevents launching with both features enabled.
Which ds4 models support tensor parallelism?
Only DeepSeek-4 models with an even number of experts support tensor parallelism. The validation logic at ds4.c lines 16929‑16935 checks that DS4_N_EXPERT is a multiple of 2 and that the model family is compatible before allowing tier pairing.
How does ds4 pair GPUs for tensor parallelism?
Ds4 automatically pairs the lower half of detected GPUs with the upper half. If you have 4 GPUs, tiers 0‑1 are paired with tiers 2‑3. The pairing logic at ds4.c lines 16937‑16950 creates these associations and aborts if any upper-half tier is already occupied by another process or placement configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →