# How to Deploy DeepSeek V4 Flash Vision-Exp on DGX Spark with Two-Node Tensor Parallelism

> Deploy DeepSeek V4 Flash Vision-Exp on DGX Spark using two-node tensor parallelism. Learn to configure NCCL, share checkpoints via NFS, and launch the DSpark vLLM runtime for efficient LLM deployment.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: how-to-guide
- Published: 2026-09-09

---

**Deploy DeepSeek V4 Flash Vision-Exp across two DGX Spark nodes by configuring NCCL over RoCE, sharing the 157 GiB checkpoint via NFS, and launching the DSpark vLLM runtime with tensor parallelism TP=2 and KV-cache dtype `nvfp4_ds_mla`.**

The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository provides a complete deployment stack for running the DeepSeek V4 Flash Vision-Exp model on a dual-node DGX Spark cluster. This configuration leverages **DSpark**—a specialized vLLM + FlashInfer runtime—to distribute model weights across two nodes using tensor parallelism while maintaining a unified KV cache for long-context inference up to 1,048,576 tokens.

## Prerequisites and Cluster Architecture

Before deploying, ensure both DGX Spark servers are connected via **RoCE/NCCL** using ConnectX adapters. Each node must run the identical container image to prevent version mismatches during distributed initialization.

- **Container Image**: `ghcr.io/anemll/dspark-vllm-gx10:0.1.1` (pull on both nodes)
- **Network**: RoCE enabled with specific NIC identifiers (e.g., `rocep1s0f1`)
- **Storage**: Head node requires ~157 GiB for the Vision-Exp checkpoint; worker can access this via NFSv4 to avoid duplicate downloads

Pull the runtime image on both systems before proceeding:

```bash
docker pull ghcr.io/anemll/dspark-vllm-gx10:0.1.1

```

## Configuring the Environment with `.env.dspark`

All deployment parameters are centralized in `.env.dspark`, sourced from the template at `.env.dspark.example`. This file defines the NCCL fabric, IP addressing, and vLLM performance knobs critical for two-node tensor parallelism.

Create and edit the environment file on the head node:

```bash
cp .env.dspark.example .env.dspark

```

Key variables for the two-node topology include:

```env
WORKER_HOST=10.0.0.2
MASTER_ADDR=10.0.0.1
VLLM_HOST_IP=10.0.0.1
WORKER_VLLM_HOST_IP=10.0.0.2
NCCL_IB_HCA=rocep1s0f1
NCCL_SOCKET_IFNAME=enp1s0f1np1
DSPARK_VLLM_IMAGE=ghcr.io/anemll/dspark-vllm-gx10:0.1.1

# Performance and model settings

MAX_MODEL_LEN=1048576
MAX_NUM_SEQS=6
MAX_NUM_BATCHED_TOKENS=8192
MTP_NUM_TOKENS=6
VLLM_USE_BREAKABLE_CUDAGRAPH=0

```

The `NCCL_IB_HCA` and `NCCL_SOCKET_IFNAME` values must match your specific InfiniBand/RoCE interface names as reported by `ibstat` or `ip link` on the DGX Spark nodes.

## Preparing the Model Checkpoint

The Vision-Exp weights must be available on the head node before starting the distributed runtime. The repository provides [`prepare-dspark-model-cache.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/prepare-dspark-model-cache.sh) to handle downloading and optional NFS export to the worker.

Download the official checkpoint on the head node:

```bash
./prepare-dspark-model-cache.sh --official

```

For the gated "abliterated" variant, use the `--abliterated` flag instead. If `DSPARK_WORKER_HF_NFS=1` is set in `.env.dspark`, the script configures the head node's HuggingFace cache as an NFSv4 export, allowing the worker to mount the model without storing a second 157 GiB copy.

## Starting the Distributed Service

The deployment uses [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) orchestrated by [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh). This script implements the correct startup order: worker first, then head, ensuring the NCCL mesh initializes properly across the TP=2 topology.

Launch the cluster from the head node:

```bash
./start-deepseek-v4-flash-dspark.sh

```

This wrapper performs the following actions:
1. Exports all `.env.dspark` variables into the Compose context
2. Launches the worker container with appropriate NCCL environment variables
3. Starts the head node vLLM server with `--tensor-parallel-size 2` spanning both nodes
4. Exposes the OpenAI-compatible API on port `8888` (`VLLM_HOST=0.0.0.0`)

The runtime initializes with **KV-cache datatype `nvfp4_ds_mla`**, block size 256, and speculative decoding enabled via `MTP_NUM_TOKENS=6`. With TP=2, each GPU holds a slice of the model weights, yielding a total KV pool of approximately 2,331,430 tokens (~17 GiB) shared across the cluster.

## Verification and API Testing

Validate the deployment by querying the health endpoint and running the smoke test:

```bash

# Check model availability

curl -fsS http://127.0.0.1:8888/v1/models

# Run comprehensive smoke test

./smoke-deepseek-v4-flash-dspark.sh

```

The expected response shows model ID `deepseek-v4-flash-vision-exp` with `max_model_len` of `1048576`.

Send a vision-language request to the endpoint:

```bash
curl http://10.0.0.1:8888/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash-vision-exp",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}},
        {"type": "text", "text": "Describe this image in detail."}
      ]
    }],
    "max_tokens": 4096,
    "temperature": 0.7
  }'

```

## Key Deployment Files

Understanding these source files helps troubleshoot and customize the deployment:

- **`.env.dspark.example`**: Template defining all cluster variables including NCCL IB HCA names and IP addresses
- **[`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml)**: Compose specification mounting the model cache and exposing port 8888
- **[`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh)**: Orchestration script that sequences the worker and head startup
- **[`prepare-dspark-model-cache.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/prepare-dspark-model-cache.sh)**: Downloads checkpoints and configures NFS sharing when `DSPARK_WORKER_HF_NFS=1`
- **[`scripts/overlay-vision-exp-ablit-cache.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/scripts/overlay-vision-exp-ablit-cache.py)**: Utility for switching between official and abliterated checkpoints by hard-linking shared blobs

## Summary

- **Two-node tensor parallelism** requires identical Docker images (`ghcr.io/anemll/dspark-vllm-gx10:0.1.1`) and matching NCCL configurations on both DGX Spark nodes
- The `.env.dspark` file controls the RoCE/NCCL fabric, IP addressing, and vLLM performance limits (1M token context, 6 concurrent sequences)
- Model weights (157 GiB) are downloaded via [`prepare-dspark-model-cache.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/prepare-dspark-model-cache.sh) and can be shared over NFSv4 to conserve worker storage
- The DSpark runtime combines vLLM with FlashInfer, using `nvfp4_ds_mla` KV-cache format and speculative decoding (`MTP_NUM_TOKENS=6`) for efficient inference
- Launch order matters: execute [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) from the head node to automatically sequence worker and head initialization

## Frequently Asked Questions

### What network interfaces should I specify for NCCL on DGX Spark?

Set `NCCL_IB_HCA` to your RoCE device name (e.g., `rocep1s0f1`) and `NCCL_SOCKET_IFNAME` to the corresponding Ethernet interface (e.g., `enp1s0f1np1`). These values must match the output of `ibstat` and `ip link` on your specific DGX Spark servers. Incorrect interface names will cause the TP=2 initialization to hang during `ncclCommInitRank`.

### Can I avoid downloading the 157 GiB checkpoint twice?

Yes. Set `DSPARK_WORKER_HF_NFS=1` in `.env.dspark` before running [`prepare-dspark-model-cache.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/prepare-dspark-model-cache.sh). This configures the head node to export its HuggingFace cache via NFSv4, allowing the worker to mount the directory read-only. The worker container will access the model weights remotely without requiring local storage for the full checkpoint.

### Why is the maximum context length limited to 1,048,576 tokens?

The `MAX_MODEL_LEN=1048576` setting in `.env.dspark` protects against KV-cache overallocation. With TP=2 and `nvfp4_ds_mla` quantization, the total available KV pool is approximately 2.3M tokens shared across all requests. Limiting individual requests to 1M tokens and `MAX_NUM_SEQS=6` ensures the aggregate working set fits safely within the allocated GPU memory while leaving headroom for the KV cache manager.

### How do I switch between the official and abliterated model variants?

Use the `--official` or `--abliterated` flags with [`prepare-dspark-model-cache.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/prepare-dspark-model-cache.sh) to download your desired checkpoint. If switching after initial deployment, run [`scripts/overlay-vision-exp-ablit-cache.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/scripts/overlay-vision-exp-ablit-cache.py) to hard-link shared blobs and copy the 26 abliterated-specific shards into the cache directory. Restart the containers via [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) to load the new weights.