DeepSeek V4 Flash on DGX Spark System Architecture: Distributed Inference Guide
DeepSeek V4 Flash Vision-Exp deploys as a distributed inference service across NVIDIA DGX Spark nodes using Tensor Parallelism, speculative decoding, and NVFP4 KV caching to support million-token context windows.
The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository provides the complete software stack for running DeepSeek V4 Flash on DGX Spark hardware. This system architecture leverages a customized vLLM inference engine with DSpark extensions to enable high-throughput generation through multi-token prediction and low-precision memory management.
Core Components of the DeepSeek V4 Flash Architecture
The system orchestrates model serving across multiple DGX Spark servers through specialized components that handle everything from checkpoint loading to high-speed inter-node communication.
- vLLM (Anemll fork): Serves as the primary inference engine, handling speculative decoding and GPU scheduling. The deployment uses the container image
ghcr.io/anemll/dspark-vllm-gx10:0.1.1as specified inREADME.md. - DSpark Speculative Decoder: Implements draft-token generation via Multi-Token Prediction (MTP), configured by default to generate
num_speculative_tokens=6(set via theMTP_NUM_TOKENSenvironment variable). - NVFP4 DS MLA KV Cache: Provides low-precision FP4 storage for key-value tensors, accommodating up to 2,331,430 tokens (approximately 17 GiB) per rank using the
--kv-cache-dtype nvfp4_ds_mlaflag. - Vision-Exp Encoder: Integrates vision transformer (ViT) and aligner components to enable image input processing (JPEG/PNG/GIF), bundled within the container image.
- NFS-Shared HF Cache: Eliminates storage duplication by allowing worker nodes to mount the head node's HuggingFace cache, controlled via the
DSPARK_WORKER_HF_NFSenvironment variable. - NCCL/RoCE Fabric: Enables high-bandwidth inter-node communication through RoCE network adapters, configured using
NCCL_IB_HCAandNCCL_SOCKET_IFNAMEvariables defined in.env.dspark.example. - Docker-Compose Orchestration: Defines container networking and volume bindings in
docker-compose.dspark.ymlfor consistent multi-node deployment.
Two-Node Tensor-Parallel Deployment (TP = 2)
The default configuration utilizes two DGX Spark nodes arranged in a head-worker topology with Tensor Parallelism set to 2.
Head and Worker Node Configuration
According to README.md (lines 95-100), the head node runs the primary vLLM container and loads the Vision-Exp checkpoint. The worker node operates a secondary vLLM container that either maintains its own checkpoint copy or mounts the head's cache via NFS when DSPARK_WORKER_HF_NFS=1 is configured.
Both containers communicate over the RoCE fabric using the NICs specified by NCCL_IB_HCA and NCCL_SOCKET_IFNAME, ensuring high-speed data transfer for tensor parallelism operations.
KV Cache and Memory Layout
The KV pool spans both ranks, with each node contributing approximately 17 GiB of NVFP4-compressed cache, yielding a combined 34 GiB capacity. This configuration supports a per-request context ceiling of 1,048,576 tokens while maintaining 6 concurrent sequences (MAX_NUM_SEQS=6).
As documented in README.md (lines 17-21), the FP4 quantization enables the 157 GiB model to serve extended contexts without exceeding the DGX Spark's memory constraints.
Speculative Decoding Implementation
Each rank executes the DSpark speculative decoder to generate draft tokens ahead of the main model pass. With MTP_NUM_TOKENS=6, the system produces six draft tokens per forward step, reducing per-token latency during autoregressive generation.
The dispatcher coordinates token verification across the head and worker nodes, merging draft predictions with final model outputs before returning responses via the HTTP API on port 8888.
Scaling DeepSeek V4 Flash to Three Nodes (TP = 3)
For higher throughput scenarios, the architecture supports a three-node configuration via start-tp3.sh.
Attention Group Padding
The TP = 3 implementation pads the model's 8 attention groups to 9, allowing each of the three ranks to handle an even distribution of attention heads. This modification is detailed in docs/TP3.md (lines 73-83).
Throughput and Latency Trade-offs
Expanding to three nodes doubles the KV cache capacity per rank to approximately 35 GiB, enabling 16 concurrent slots (TP3_MAX_NUM_SEQS=16). According to docs/TP3.md (lines 18-33), this configuration achieves roughly 200 tokens per second aggregate throughput, though it introduces higher pre-fill latency for very long prompts.
Deployment Workflow
Deploying the system requires preparing the model cache and launching the distributed containers in sequence.
First, prepare the Vision-Exp checkpoint on the head node using the provided helper script:
./prepare-dspark-model-cache.sh --official
Launch the two-node service (start the head node first, then the worker):
./start-deepseek-v4-flash-dspark.sh
Verify the service availability using the smoke test and API endpoint:
curl -fsS http://127.0.0.1:8888/v1/models
./smoke-deepseek-v4-flash-dspark.sh
For the optional three-node configuration, define the additional worker in .env.dspark:
WORKER2_HOST=10.0.0.3
# Additional WORKER2_* NCCL variables as specified in docs/TP3.md
Then launch with extended concurrency:
./start-tp3.sh --max-num-seqs 16
Summary
- DeepSeek V4 Flash deploys as a distributed inference service on NVIDIA DGX Spark nodes using Tensor Parallelism (TP = 2 or 3).
- The architecture combines vLLM (Anemll fork), DSpark speculative decoding (MTP), and NVFP4 FP4 KV caching to maximize throughput and context length.
- The two-node default configuration provides 34 GiB of aggregate KV cache (17 GiB per rank), supporting 1,048,576 tokens per request and 6 concurrent sequences.
- Optional three-node scaling increases capacity to 16 concurrent slots with approximately 35 GiB cache per rank, yielding ~200 tok/s throughput.
- NFS checkpoint sharing via
DSPARK_WORKER_HF_NFSeliminates redundant storage of the 157 GiB model across worker nodes.
Frequently Asked Questions
What hardware is required for DeepSeek V4 Flash on DGX Spark?
The system requires two NVIDIA DGX Spark nodes for the default configuration (head plus worker), with an optional third node for TP = 3 deployments. Each node requires RoCE-capable network interfaces configured via NCCL_IB_HCA and NCCL_SOCKET_IFNAME for high-speed inter-GPU communication.
How does the NVFP4 KV cache improve performance?
The NVFP4 DS MLA KV cache stores key-value tensors in 4-bit floating-point precision, compressing approximately 2.3 million tokens into 17 GiB per rank. This quantization, enabled via --kv-cache-dtype nvfp4_ds_mla, allows the system to maintain a million-token context window within the DGX Spark's memory constraints while minimizing bandwidth pressure during attention operations.
What is the difference between TP=2 and TP=3 configurations?
TP = 2 (default) splits the model across two DGX Spark nodes with 6 concurrent sequence slots and moderate latency. TP = 3 adds a third node, pads attention groups from 8 to 9 for even distribution, and increases concurrency to 16 slots with roughly 35 GiB KV cache per rank. The three-node configuration delivers higher throughput (~200 tok/s) but exhibits increased pre-fill latency for long prompts.
How is the model checkpoint shared between nodes?
Worker nodes can access the 157 GiB Vision-Exp checkpoint without local duplication by mounting the head node's HuggingFace cache over NFS. Set DSPARK_WORKER_HF_NFS=1 in the environment configuration to enable this sharing, significantly reducing startup time and storage requirements on worker nodes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →