How to Deploy DeepSeek V4 Flash Vision-Exp on DGX Spark with Two-Node Tensor Parallelism

Deploy DeepSeek V4 Flash Vision-Exp across two DGX Spark nodes by configuring NCCL over RoCE, sharing the 157 GiB checkpoint via NFS, and launching the DSpark vLLM runtime with tensor parallelism TP=2 and KV-cache dtype nvfp4_ds_mla.

The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository provides a complete deployment stack for running the DeepSeek V4 Flash Vision-Exp model on a dual-node DGX Spark cluster. This configuration leverages DSpark—a specialized vLLM + FlashInfer runtime—to distribute model weights across two nodes using tensor parallelism while maintaining a unified KV cache for long-context inference up to 1,048,576 tokens.

Prerequisites and Cluster Architecture

Before deploying, ensure both DGX Spark servers are connected via RoCE/NCCL using ConnectX adapters. Each node must run the identical container image to prevent version mismatches during distributed initialization.

  • Container Image: ghcr.io/anemll/dspark-vllm-gx10:0.1.1 (pull on both nodes)
  • Network: RoCE enabled with specific NIC identifiers (e.g., rocep1s0f1)
  • Storage: Head node requires ~157 GiB for the Vision-Exp checkpoint; worker can access this via NFSv4 to avoid duplicate downloads

Pull the runtime image on both systems before proceeding:

docker pull ghcr.io/anemll/dspark-vllm-gx10:0.1.1

Configuring the Environment with .env.dspark

All deployment parameters are centralized in .env.dspark, sourced from the template at .env.dspark.example. This file defines the NCCL fabric, IP addressing, and vLLM performance knobs critical for two-node tensor parallelism.

Create and edit the environment file on the head node:

cp .env.dspark.example .env.dspark

Key variables for the two-node topology include:

WORKER_HOST=10.0.0.2
MASTER_ADDR=10.0.0.1
VLLM_HOST_IP=10.0.0.1
WORKER_VLLM_HOST_IP=10.0.0.2
NCCL_IB_HCA=rocep1s0f1
NCCL_SOCKET_IFNAME=enp1s0f1np1
DSPARK_VLLM_IMAGE=ghcr.io/anemll/dspark-vllm-gx10:0.1.1

# Performance and model settings

MAX_MODEL_LEN=1048576
MAX_NUM_SEQS=6
MAX_NUM_BATCHED_TOKENS=8192
MTP_NUM_TOKENS=6
VLLM_USE_BREAKABLE_CUDAGRAPH=0

The NCCL_IB_HCA and NCCL_SOCKET_IFNAME values must match your specific InfiniBand/RoCE interface names as reported by ibstat or ip link on the DGX Spark nodes.

Preparing the Model Checkpoint

The Vision-Exp weights must be available on the head node before starting the distributed runtime. The repository provides prepare-dspark-model-cache.sh to handle downloading and optional NFS export to the worker.

Download the official checkpoint on the head node:

./prepare-dspark-model-cache.sh --official

For the gated "abliterated" variant, use the --abliterated flag instead. If DSPARK_WORKER_HF_NFS=1 is set in .env.dspark, the script configures the head node's HuggingFace cache as an NFSv4 export, allowing the worker to mount the model without storing a second 157 GiB copy.

Starting the Distributed Service

The deployment uses docker-compose.dspark.yml orchestrated by start-deepseek-v4-flash-dspark.sh. This script implements the correct startup order: worker first, then head, ensuring the NCCL mesh initializes properly across the TP=2 topology.

Launch the cluster from the head node:

./start-deepseek-v4-flash-dspark.sh

This wrapper performs the following actions:

  1. Exports all .env.dspark variables into the Compose context
  2. Launches the worker container with appropriate NCCL environment variables
  3. Starts the head node vLLM server with --tensor-parallel-size 2 spanning both nodes
  4. Exposes the OpenAI-compatible API on port 8888 (VLLM_HOST=0.0.0.0)

The runtime initializes with KV-cache datatype nvfp4_ds_mla, block size 256, and speculative decoding enabled via MTP_NUM_TOKENS=6. With TP=2, each GPU holds a slice of the model weights, yielding a total KV pool of approximately 2,331,430 tokens (~17 GiB) shared across the cluster.

Verification and API Testing

Validate the deployment by querying the health endpoint and running the smoke test:


# Check model availability

curl -fsS http://127.0.0.1:8888/v1/models

# Run comprehensive smoke test

./smoke-deepseek-v4-flash-dspark.sh

The expected response shows model ID deepseek-v4-flash-vision-exp with max_model_len of 1048576.

Send a vision-language request to the endpoint:

curl http://10.0.0.1:8888/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash-vision-exp",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}},
        {"type": "text", "text": "Describe this image in detail."}
      ]
    }],
    "max_tokens": 4096,
    "temperature": 0.7
  }'

Key Deployment Files

Understanding these source files helps troubleshoot and customize the deployment:

Summary

  • Two-node tensor parallelism requires identical Docker images (ghcr.io/anemll/dspark-vllm-gx10:0.1.1) and matching NCCL configurations on both DGX Spark nodes
  • The .env.dspark file controls the RoCE/NCCL fabric, IP addressing, and vLLM performance limits (1M token context, 6 concurrent sequences)
  • Model weights (157 GiB) are downloaded via prepare-dspark-model-cache.sh and can be shared over NFSv4 to conserve worker storage
  • The DSpark runtime combines vLLM with FlashInfer, using nvfp4_ds_mla KV-cache format and speculative decoding (MTP_NUM_TOKENS=6) for efficient inference
  • Launch order matters: execute start-deepseek-v4-flash-dspark.sh from the head node to automatically sequence worker and head initialization

Frequently Asked Questions

What network interfaces should I specify for NCCL on DGX Spark?

Set NCCL_IB_HCA to your RoCE device name (e.g., rocep1s0f1) and NCCL_SOCKET_IFNAME to the corresponding Ethernet interface (e.g., enp1s0f1np1). These values must match the output of ibstat and ip link on your specific DGX Spark servers. Incorrect interface names will cause the TP=2 initialization to hang during ncclCommInitRank.

Can I avoid downloading the 157 GiB checkpoint twice?

Yes. Set DSPARK_WORKER_HF_NFS=1 in .env.dspark before running prepare-dspark-model-cache.sh. This configures the head node to export its HuggingFace cache via NFSv4, allowing the worker to mount the directory read-only. The worker container will access the model weights remotely without requiring local storage for the full checkpoint.

Why is the maximum context length limited to 1,048,576 tokens?

The MAX_MODEL_LEN=1048576 setting in .env.dspark protects against KV-cache overallocation. With TP=2 and nvfp4_ds_mla quantization, the total available KV pool is approximately 2.3M tokens shared across all requests. Limiting individual requests to 1M tokens and MAX_NUM_SEQS=6 ensures the aggregate working set fits safely within the allocated GPU memory while leaving headroom for the KV cache manager.

How do I switch between the official and abliterated model variants?

Use the --official or --abliterated flags with prepare-dspark-model-cache.sh to download your desired checkpoint. If switching after initial deployment, run scripts/overlay-vision-exp-ablit-cache.py to hard-link shared blobs and copy the 26 abliterated-specific shards into the cache directory. Restart the containers via start-deepseek-v4-flash-dspark.sh to load the new weights.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →