# How to Configure CUDA Multi-GPU Tensor Parallelism in ds4

> Configure CUDA multi-GPU tensor parallelism in ds4. Follow steps: compile with CUDA, ensure even GPUs, use DeepSeek-4 with even experts, disable SSD streaming, and launch with the tensor parallel flag.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-09

---

**To configure CUDA multi-GPU tensor parallelism in ds4, compile with CUDA support, ensure an even number of GPUs, use a DeepSeek-4 model with even expert counts, disable conflicting flags like `--ssd-streaming`, and launch with the `--tensor-parallel` flag.**

Tensor parallelism in the antirez/ds4 repository enables a single DeepSeek-4 model to be split across multiple CUDA GPUs, allowing each device to hold a portion of the model’s layers while cooperatively performing pre-fill and decode operations. This configuration requires satisfying specific runtime checks and CLI options before the engine initializes a multi-GPU tensor-parallel session. Below is a step-by-step guide referencing the exact source locations that enforce each requirement.

## Prerequisites and Build Configuration

### Compile with CUDA Support

Before enabling tensor parallelism, you must build ds4 with CUDA support enabled. The default Makefile automatically detects the CUDA toolkit and compiles the necessary back-end components.

The CUDA implementation resides in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) and [`ds4_gpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu.c). Without these components, the tensor-parallel branch is never compiled into the binary, and multi-GPU placement logic remains inaccessible.

### Verify GPU Count Requirements

Ds4 requires an even number of available GPUs to perform tensor parallelism. The runtime checks the global variable `g_n_gpus` (the detected GPU count) and validates that an even placement exists.

In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) at lines [16921‑16927](https://github.com/antirez/ds4/blob/main/ds4.c#L16921), the engine aborts with an error if the GPU count is odd, ensuring symmetric tier pairing is possible.

## Model and Compatibility Validation

### Supported Model Architectures

Tensor parallelism is restricted to DeepSeek-4 models that contain an **even number of experts**. The constant `DS4_N_EXPERT` must be a multiple of 2 for the model to qualify for splitting across tiers.

At lines [16929‑16935](https://github.com/antirez/ds4/blob/main/ds4.c#L16929) in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), the engine validates both the model family and expert parity before proceeding with tensor-parallel initialization.

### Disable Incompatible Runtime Flags

Certain features conflict with tensor parallelism and must be explicitly disabled. You cannot enable **SSD streaming** (`--ssd-streaming`) because tensor parallelism requires resident weights in GPU memory. Additionally, the **distributed role** (`--role distributed`) is incompatible with local tensor-parallel configurations.

The validation logic at lines [499‑507](https://github.com/antirez/ds4/blob/main/ds4.c#L499) in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) explicitly reports these incompatibilities and prevents engine startup if conflicting options are detected.

## Enabling Tensor Parallelism

### CLI Configuration

To activate tensor parallelism, pass the `--tensor-parallel` flag (or the short form `-t`) when starting the server or CLI. This option is registered in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) at lines [246‑247](https://github.com/antirez/ds4/blob/main/ds4_help.c#L246).

```bash
./ds4_server -m /path/to/deepseek4.q4_0.gguf -t

```

If you previously enabled SSD streaming in your configuration file, explicitly disable it:

```bash
./ds4_server -m model.gguf --no-ssd-streaming -t

```

### Internal Tier Pairing Logic

Once the CLI flags are validated, ds4 internally pairs GPUs into lower-half and upper-half tiers. The engine maps tier `0` through `half-1` with tier `half` through `g_n_gpus-1`. The system will reject any placement that already utilizes an upper-half tier to prevent resource conflicts.

This pairing logic and its associated guard clauses are implemented in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) at lines [16937‑16950](https://github.com/antirez/ds4/blob/main/ds4.c#L16937).

### Per-Tier Tensor Allocation

For every tier participating in the tensor-parallel group, ds4 allocates necessary buffers including KV-cache, hidden-state, and attention tensors. The function `ds4_gpu_tensor_alloc_ptr_on` handles these allocations.

The tier-wise allocation loop appears in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) at lines [17056‑17078](https://github.com/antirez/ds4/blob/main/ds4.c#L17056), ensuring each GPU receives its partitioned portion of the model state.

## Advanced Configuration Options

### Environment Variables for KV-Cache Tuning

For extremely long contexts or specific performance requirements, you can tune kernel fusion behavior using environment variables:

- `DS4_CUDA_NO_QKV_PAIR`: Disables QKV-pair kernel optimization
- `DS4_CUDA_TP_ATTN_OUT_HC_FUSE`: Enables attention-output half-cache fusion

These variables are consulted during initialization in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) around lines [16998‑17004](https://github.com/antirez/ds4/blob/main/ds4.c#L16998).

```bash
export DS4_CUDA_TP_ATTN_OUT_HC_FUSE=1
export DS4_CUDA_NO_QKV_PAIR=1
./ds4_server -m model.gguf -t

```

### Engine Initialization

After all checks pass, the function `engine_classify_multi_tier()` creates a `ds4_engine` instance configured for tensor-parallel execution. This launch sequence occurs in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) at lines [56312‑56320](https://github.com/antirez/ds4/blob/main/ds4.c#L56312).

## Usage Examples

### Basic Server Launch

Launch a tensor-parallel server on a Linux machine with two CUDA GPUs:

```bash
./ds4_server -m /path/to/deepseek4.q4_0.gguf -t

```

### Python Client Interaction

Once the server is running, interact with it via the HTTP API:

```python
import requests
import json

url = "http://localhost:8080/v1/chat/completions"
payload = {
    "model": "deepseek4",
    "messages": [{"role": "user", "content": "Explain tensor parallelism"}],
    "max_tokens": 256,
}
resp = requests.post(url, json=payload)
print(json.dumps(resp.json(), indent=2))

```

## Key Source Files

The following files implement the tensor-parallel configuration and runtime logic:

- **[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)**: Core engine containing runtime checks for GPU count, model compatibility, tier pairing, and per-tier tensor allocation.
- **[`ds4_gpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu.h) / [`ds4_gpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu.c)**: CUDA GPU abstraction layer defining `ds4_gpu_tensor_*` APIs for tier-wise memory management.
- **[`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c)**: Tensor-parallelism specific helpers and error message generation (e.g., `tp_set_err`).
- **[`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c)**: CLI option registration, including the `--tensor-parallel` flag.
- **[`tests/test_engine_mgpu_placement.c`](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_placement.c)**: Unit tests validating correct tier pairing and placement logic for multi-GPU setups.

## Summary

- **Build requirement**: Compile ds4 with CUDA support via the default Makefile to include [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) and [`ds4_gpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu.c) back-ends.
- **Hardware requirement**: Provide an even number of GPUs; `g_n_gpus` validation occurs at lines 16921‑16927.
- **Model requirement**: Use DeepSeek-4 models with even expert counts (`DS4_N_EXPERT % 2 == 0`), verified at lines 16929‑16935.
- **Flag conflicts**: Disable `--ssd-streaming` and avoid `--role distributed` to prevent launch aborts at lines 499‑507.
- **Activation**: Pass `--tensor-parallel` or `-t`, parsed in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) at lines 246‑247.
- **Allocation**: Per-tier tensors are allocated via `ds4_gpu_tensor_alloc_ptr_on` in the loop at lines 17056‑17078.
- **Launch**: The engine initializes through `engine_classify_multi_tier()` at lines 56312‑56320.

## Frequently Asked Questions

### What happens if I try to use tensor parallelism with an odd number of GPUs?

The engine will detect the asymmetry during initialization and abort with a clear error message. Specifically, the check at [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 16921‑16927 validates that `g_n_gpus` is even, as tensor parallelism requires symmetric pairing between lower and upper tier groups.

### Can I use SSD streaming with tensor parallelism enabled?

No. SSD streaming requires weights to reside on disk with partial resident memory, while tensor parallelism requires full resident weights on GPU tiers. The incompatibility check at [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 499‑507 explicitly prevents launching with both features enabled.

### Which ds4 models support tensor parallelism?

Only DeepSeek-4 models with an even number of experts support tensor parallelism. The validation logic at [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 16929‑16935 checks that `DS4_N_EXPERT` is a multiple of 2 and that the model family is compatible before allowing tier pairing.

### How does ds4 pair GPUs for tensor parallelism?

Ds4 automatically pairs the lower half of detected GPUs with the upper half. If you have 4 GPUs, tiers 0‑1 are paired with tiers 2‑3. The pairing logic at [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 16937‑16950 creates these associations and aborts if any upper-half tier is already occupied by another process or placement configuration.