How to Deploy DeepSeek V4 Flash Vision-Exp on DGX Spark with Two-Node Tensor Parallelism
Deploy DeepSeek V4 Flash Vision-Exp across two DGX Spark nodes by configuring NCCL over RoCE, sharing the 157 GiB checkpoint via NFS, and launching the DSpark vLLM runtime with tensor parallelism TP=2 and KV-cache dtype nvfp4_ds_mla.
The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository provides a complete deployment stack for running the DeepSeek V4 Flash Vision-Exp model on a dual-node DGX Spark cluster. This configuration leverages DSpark—a specialized vLLM + FlashInfer runtime—to distribute model weights across two nodes using tensor parallelism while maintaining a unified KV cache for long-context inference up to 1,048,576 tokens.
Prerequisites and Cluster Architecture
Before deploying, ensure both DGX Spark servers are connected via RoCE/NCCL using ConnectX adapters. Each node must run the identical container image to prevent version mismatches during distributed initialization.
- Container Image:
ghcr.io/anemll/dspark-vllm-gx10:0.1.1(pull on both nodes) - Network: RoCE enabled with specific NIC identifiers (e.g.,
rocep1s0f1) - Storage: Head node requires ~157 GiB for the Vision-Exp checkpoint; worker can access this via NFSv4 to avoid duplicate downloads
Pull the runtime image on both systems before proceeding:
docker pull ghcr.io/anemll/dspark-vllm-gx10:0.1.1
Configuring the Environment with .env.dspark
All deployment parameters are centralized in .env.dspark, sourced from the template at .env.dspark.example. This file defines the NCCL fabric, IP addressing, and vLLM performance knobs critical for two-node tensor parallelism.
Create and edit the environment file on the head node:
cp .env.dspark.example .env.dspark
Key variables for the two-node topology include:
WORKER_HOST=10.0.0.2
MASTER_ADDR=10.0.0.1
VLLM_HOST_IP=10.0.0.1
WORKER_VLLM_HOST_IP=10.0.0.2
NCCL_IB_HCA=rocep1s0f1
NCCL_SOCKET_IFNAME=enp1s0f1np1
DSPARK_VLLM_IMAGE=ghcr.io/anemll/dspark-vllm-gx10:0.1.1
# Performance and model settings
MAX_MODEL_LEN=1048576
MAX_NUM_SEQS=6
MAX_NUM_BATCHED_TOKENS=8192
MTP_NUM_TOKENS=6
VLLM_USE_BREAKABLE_CUDAGRAPH=0
The NCCL_IB_HCA and NCCL_SOCKET_IFNAME values must match your specific InfiniBand/RoCE interface names as reported by ibstat or ip link on the DGX Spark nodes.
Preparing the Model Checkpoint
The Vision-Exp weights must be available on the head node before starting the distributed runtime. The repository provides prepare-dspark-model-cache.sh to handle downloading and optional NFS export to the worker.
Download the official checkpoint on the head node:
./prepare-dspark-model-cache.sh --official
For the gated "abliterated" variant, use the --abliterated flag instead. If DSPARK_WORKER_HF_NFS=1 is set in .env.dspark, the script configures the head node's HuggingFace cache as an NFSv4 export, allowing the worker to mount the model without storing a second 157 GiB copy.
Starting the Distributed Service
The deployment uses docker-compose.dspark.yml orchestrated by start-deepseek-v4-flash-dspark.sh. This script implements the correct startup order: worker first, then head, ensuring the NCCL mesh initializes properly across the TP=2 topology.
Launch the cluster from the head node:
./start-deepseek-v4-flash-dspark.sh
This wrapper performs the following actions:
- Exports all
.env.dsparkvariables into the Compose context - Launches the worker container with appropriate NCCL environment variables
- Starts the head node vLLM server with
--tensor-parallel-size 2spanning both nodes - Exposes the OpenAI-compatible API on port
8888(VLLM_HOST=0.0.0.0)
The runtime initializes with KV-cache datatype nvfp4_ds_mla, block size 256, and speculative decoding enabled via MTP_NUM_TOKENS=6. With TP=2, each GPU holds a slice of the model weights, yielding a total KV pool of approximately 2,331,430 tokens (~17 GiB) shared across the cluster.
Verification and API Testing
Validate the deployment by querying the health endpoint and running the smoke test:
# Check model availability
curl -fsS http://127.0.0.1:8888/v1/models
# Run comprehensive smoke test
./smoke-deepseek-v4-flash-dspark.sh
The expected response shows model ID deepseek-v4-flash-vision-exp with max_model_len of 1048576.
Send a vision-language request to the endpoint:
curl http://10.0.0.1:8888/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash-vision-exp",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}},
{"type": "text", "text": "Describe this image in detail."}
]
}],
"max_tokens": 4096,
"temperature": 0.7
}'
Key Deployment Files
Understanding these source files helps troubleshoot and customize the deployment:
.env.dspark.example: Template defining all cluster variables including NCCL IB HCA names and IP addressesdocker-compose.dspark.yml: Compose specification mounting the model cache and exposing port 8888start-deepseek-v4-flash-dspark.sh: Orchestration script that sequences the worker and head startupprepare-dspark-model-cache.sh: Downloads checkpoints and configures NFS sharing whenDSPARK_WORKER_HF_NFS=1scripts/overlay-vision-exp-ablit-cache.py: Utility for switching between official and abliterated checkpoints by hard-linking shared blobs
Summary
- Two-node tensor parallelism requires identical Docker images (
ghcr.io/anemll/dspark-vllm-gx10:0.1.1) and matching NCCL configurations on both DGX Spark nodes - The
.env.dsparkfile controls the RoCE/NCCL fabric, IP addressing, and vLLM performance limits (1M token context, 6 concurrent sequences) - Model weights (157 GiB) are downloaded via
prepare-dspark-model-cache.shand can be shared over NFSv4 to conserve worker storage - The DSpark runtime combines vLLM with FlashInfer, using
nvfp4_ds_mlaKV-cache format and speculative decoding (MTP_NUM_TOKENS=6) for efficient inference - Launch order matters: execute
start-deepseek-v4-flash-dspark.shfrom the head node to automatically sequence worker and head initialization
Frequently Asked Questions
What network interfaces should I specify for NCCL on DGX Spark?
Set NCCL_IB_HCA to your RoCE device name (e.g., rocep1s0f1) and NCCL_SOCKET_IFNAME to the corresponding Ethernet interface (e.g., enp1s0f1np1). These values must match the output of ibstat and ip link on your specific DGX Spark servers. Incorrect interface names will cause the TP=2 initialization to hang during ncclCommInitRank.
Can I avoid downloading the 157 GiB checkpoint twice?
Yes. Set DSPARK_WORKER_HF_NFS=1 in .env.dspark before running prepare-dspark-model-cache.sh. This configures the head node to export its HuggingFace cache via NFSv4, allowing the worker to mount the directory read-only. The worker container will access the model weights remotely without requiring local storage for the full checkpoint.
Why is the maximum context length limited to 1,048,576 tokens?
The MAX_MODEL_LEN=1048576 setting in .env.dspark protects against KV-cache overallocation. With TP=2 and nvfp4_ds_mla quantization, the total available KV pool is approximately 2.3M tokens shared across all requests. Limiting individual requests to 1M tokens and MAX_NUM_SEQS=6 ensures the aggregate working set fits safely within the allocated GPU memory while leaving headroom for the KV cache manager.
How do I switch between the official and abliterated model variants?
Use the --official or --abliterated flags with prepare-dspark-model-cache.sh to download your desired checkpoint. If switching after initial deployment, run scripts/overlay-vision-exp-ablit-cache.py to hard-link shared blobs and copy the 26 abliterated-specific shards into the cache directory. Restart the containers via start-deepseek-v4-flash-dspark.sh to load the new weights.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →