# DeepSeek V4 Flash on DGX Spark System Architecture: Distributed Inference Guide

> Explore the DeepSeek V4 Flash system architecture on DGX Spark. Learn how to deploy distributed inference with Tensor Parallelism and speculative decoding for million-token contexts.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: architecture
- Published: 2026-09-09

---

**DeepSeek V4 Flash Vision-Exp deploys as a distributed inference service across NVIDIA DGX Spark nodes using Tensor Parallelism, speculative decoding, and NVFP4 KV caching to support million-token context windows.**

The [MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark) repository provides the complete software stack for running DeepSeek V4 Flash on DGX Spark hardware. This system architecture leverages a customized vLLM inference engine with DSpark extensions to enable high-throughput generation through multi-token prediction and low-precision memory management.

## Core Components of the DeepSeek V4 Flash Architecture

The system orchestrates model serving across multiple DGX Spark servers through specialized components that handle everything from checkpoint loading to high-speed inter-node communication.

- **vLLM (Anemll fork)**: Serves as the primary inference engine, handling speculative decoding and GPU scheduling. The deployment uses the container image `ghcr.io/anemll/dspark-vllm-gx10:0.1.1` as specified in [`README.md`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/README.md).
- **DSpark Speculative Decoder**: Implements draft-token generation via Multi-Token Prediction (MTP), configured by default to generate `num_speculative_tokens=6` (set via the `MTP_NUM_TOKENS` environment variable).
- **NVFP4 DS MLA KV Cache**: Provides low-precision FP4 storage for key-value tensors, accommodating up to **2,331,430 tokens** (approximately **17 GiB**) per rank using the `--kv-cache-dtype nvfp4_ds_mla` flag.
- **Vision-Exp Encoder**: Integrates vision transformer (ViT) and aligner components to enable image input processing (JPEG/PNG/GIF), bundled within the container image.
- **NFS-Shared HF Cache**: Eliminates storage duplication by allowing worker nodes to mount the head node's HuggingFace cache, controlled via the `DSPARK_WORKER_HF_NFS` environment variable.
- **NCCL/RoCE Fabric**: Enables high-bandwidth inter-node communication through RoCE network adapters, configured using `NCCL_IB_HCA` and `NCCL_SOCKET_IFNAME` variables defined in `.env.dspark.example`.
- **Docker-Compose Orchestration**: Defines container networking and volume bindings in [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) for consistent multi-node deployment.

## Two-Node Tensor-Parallel Deployment (TP = 2)

The default configuration utilizes two DGX Spark nodes arranged in a head-worker topology with Tensor Parallelism set to 2.

### Head and Worker Node Configuration

According to [`README.md`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/README.md) (lines 95-100), the head node runs the primary vLLM container and loads the Vision-Exp checkpoint. The worker node operates a secondary vLLM container that either maintains its own checkpoint copy or mounts the head's cache via NFS when `DSPARK_WORKER_HF_NFS=1` is configured.

Both containers communicate over the RoCE fabric using the NICs specified by `NCCL_IB_HCA` and `NCCL_SOCKET_IFNAME`, ensuring high-speed data transfer for tensor parallelism operations.

### KV Cache and Memory Layout

The KV pool spans both ranks, with each node contributing approximately **17 GiB** of NVFP4-compressed cache, yielding a combined **34 GiB** capacity. This configuration supports a **per-request context ceiling of 1,048,576 tokens** while maintaining **6 concurrent sequences** (`MAX_NUM_SEQS=6`).

As documented in [`README.md`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/README.md) (lines 17-21), the FP4 quantization enables the 157 GiB model to serve extended contexts without exceeding the DGX Spark's memory constraints.

### Speculative Decoding Implementation

Each rank executes the DSpark speculative decoder to generate draft tokens ahead of the main model pass. With `MTP_NUM_TOKENS=6`, the system produces six draft tokens per forward step, reducing per-token latency during autoregressive generation.

The dispatcher coordinates token verification across the head and worker nodes, merging draft predictions with final model outputs before returning responses via the HTTP API on port 8888.

## Scaling DeepSeek V4 Flash to Three Nodes (TP = 3)

For higher throughput scenarios, the architecture supports a three-node configuration via [`start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-tp3.sh).

### Attention Group Padding

The TP = 3 implementation pads the model's 8 attention groups to 9, allowing each of the three ranks to handle an even distribution of attention heads. This modification is detailed in [`docs/TP3.md`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docs/TP3.md) (lines 73-83).

### Throughput and Latency Trade-offs

Expanding to three nodes doubles the KV cache capacity per rank to approximately **35 GiB**, enabling **16 concurrent slots** (`TP3_MAX_NUM_SEQS=16`). According to [`docs/TP3.md`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docs/TP3.md) (lines 18-33), this configuration achieves roughly **200 tokens per second** aggregate throughput, though it introduces higher pre-fill latency for very long prompts.

## Deployment Workflow

Deploying the system requires preparing the model cache and launching the distributed containers in sequence.

First, prepare the Vision-Exp checkpoint on the head node using the provided helper script:

```bash
./prepare-dspark-model-cache.sh --official

```

Launch the two-node service (start the head node first, then the worker):

```bash
./start-deepseek-v4-flash-dspark.sh

```

Verify the service availability using the smoke test and API endpoint:

```bash
curl -fsS http://127.0.0.1:8888/v1/models
./smoke-deepseek-v4-flash-dspark.sh

```

For the optional three-node configuration, define the additional worker in `.env.dspark`:

```bash
WORKER2_HOST=10.0.0.3

# Additional WORKER2_* NCCL variables as specified in docs/TP3.md

```

Then launch with extended concurrency:

```bash
./start-tp3.sh --max-num-seqs 16

```

## Summary

- **DeepSeek V4 Flash** deploys as a distributed inference service on **NVIDIA DGX Spark** nodes using Tensor Parallelism (TP = 2 or 3).
- The architecture combines **vLLM (Anemll fork)**, **DSpark speculative decoding** (MTP), and **NVFP4 FP4 KV caching** to maximize throughput and context length.
- The two-node default configuration provides **34 GiB** of aggregate KV cache (17 GiB per rank), supporting **1,048,576 tokens** per request and **6 concurrent sequences**.
- Optional three-node scaling increases capacity to **16 concurrent slots** with approximately **35 GiB** cache per rank, yielding ~200 tok/s throughput.
- **NFS checkpoint sharing** via `DSPARK_WORKER_HF_NFS` eliminates redundant storage of the 157 GiB model across worker nodes.

## Frequently Asked Questions

### What hardware is required for DeepSeek V4 Flash on DGX Spark?

The system requires **two NVIDIA DGX Spark nodes** for the default configuration (head plus worker), with an optional third node for TP = 3 deployments. Each node requires RoCE-capable network interfaces configured via `NCCL_IB_HCA` and `NCCL_SOCKET_IFNAME` for high-speed inter-GPU communication.

### How does the NVFP4 KV cache improve performance?

The **NVFP4 DS MLA KV cache** stores key-value tensors in 4-bit floating-point precision, compressing approximately **2.3 million tokens** into **17 GiB** per rank. This quantization, enabled via `--kv-cache-dtype nvfp4_ds_mla`, allows the system to maintain a million-token context window within the DGX Spark's memory constraints while minimizing bandwidth pressure during attention operations.

### What is the difference between TP=2 and TP=3 configurations?

**TP = 2** (default) splits the model across two DGX Spark nodes with **6 concurrent sequence slots** and moderate latency. **TP = 3** adds a third node, pads attention groups from 8 to 9 for even distribution, and increases concurrency to **16 slots** with roughly **35 GiB** KV cache per rank. The three-node configuration delivers higher throughput (~200 tok/s) but exhibits increased pre-fill latency for long prompts.

### How is the model checkpoint shared between nodes?

Worker nodes can access the **157 GiB Vision-Exp checkpoint** without local duplication by mounting the head node's HuggingFace cache over NFS. Set `DSPARK_WORKER_HF_NFS=1` in the environment configuration to enable this sharing, significantly reducing startup time and storage requirements on worker nodes.