# DeepSeek-v4-Flash DSpark Multi-Node Launch Sequence: Head and Worker Node Setup

> Learn the DeepSeek-v4-Flash-DSpark multi-node launch sequence. Set up your head node and worker nodes efficiently for seamless cluster operation and GPU registration.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: how-to-guide
- Published: 2026-09-09

---

**The DeepSeek-v4-Flash-DSpark multi-node launch sequence orchestrates a head node running [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) to initialize the controller stack, followed by worker nodes executing [`start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-tp3.sh) to register GPU workers with the cluster.**

This guide explains the complete startup procedure for distributing the DeepSeek-v4-Flash inference engine across multiple DGX nodes using the DSpark orchestration framework. The MiaAI-Lab repository provides purpose-built shell scripts and Docker Compose configurations that automate the coordination between head (orchestrator) and worker (GPU compute) nodes.

---

## Multi-Node Launch Sequence Overview

The repository implements a three-phase launch protocol designed for 2x DGX-Spark clusters. Each phase targets a specific node type with dedicated scripts that handle container orchestration, network binding, and service discovery.

| Phase | Node Type | Action | Script |
|-------|-----------|--------|--------|
| 1 | **Head** | Initialize controller, VLLM runtime, and auxiliary services | [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) |
| 2 | **Worker** | Launch GPU-bound VLLM worker containers bound to head network | [`start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-tp3.sh) |
| 3 | **Both** | Verify deployment health and run smoke tests | [`status-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/status-deepseek-v4-flash-dspark.sh), [`smoke-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/smoke-deepseek-v4-flash-dspark.sh) |

---

## Phase 1: Head Node Initialization

The head node serves as the cluster orchestrator, exposing the API endpoint and managing worker registration. In [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh), the script wraps `docker compose -f docker-compose.dspark.yml up -d` to bring up three critical services.

**Services started on the head node:**

- **`dspark-controller`** — DSpark master that schedules and discovers workers
- **`vllm-runtime`** — VLLM inference server with OpenAI-compatible `/v1/completions` endpoint
- **`lmcache`** — Optional high-performance KV cache for accelerated serving

### Head Node Start Command

```bash

# On the orchestrator machine

git clone https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark.git
cd DeepSeek-v4-Flash-DSpark-2x-DGX-Spark

# Configure environment

cp .env.dspark.example .env

# Edit .env: set DSK_CLUSTER_NAME, NFS_SHARE, head node IP, etc.

# Optional: pre-populate model weights on shared storage

./prepare-dspark-model-cache.sh

# Launch the head stack

./start-deepseek-v4-flash-dspark.sh

```

The [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) defines a shared Docker network that enables automatic service discovery between head and worker containers across physical nodes.

---

## Phase 2: Worker Node Registration

Worker nodes execute [`start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-tp3.sh) to join the cluster. This script configures the node as a VLLM worker instance and establishes GPU-to-GPU communication with the head.

### Worker Node Start Command

```bash

# On each GPU worker machine

git clone https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark.git
cd DeepSeek-v4-Flash-DSpark-2x-DGX-Spark

# Use identical environment configuration

cp .env.dspark.example .env

# Ensure NFS_SHARE and DSK_CLUSTER_NAME match head node values

# Launch worker and register with head

./start-tp3.sh

```

### Worker Configuration Details

The [`start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-tp3.sh) script performs these operations:

1. **Sets role identifier**: Exports `DSK_CLUSTER_ROLE=worker` for runtime discovery
2. **Mounts model cache**: Binds `/mnt/model-cache` via NFS for shared weight access
3. **Configures transport**: Sources [`dspark-numeric-knobs.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/dspark-numeric-knobs.sh) to set NCCL and InfiniBand parameters
4. **Network attachment**: Joins the Docker network created by the head node's Compose stack

The [`dspark-numeric-knobs.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/dspark-numeric-knobs.sh) file contains environment-specific tuning variables for high-throughput GPU communication across the DGX interconnect.

---

## Phase 3: Health Verification and Testing

After launching all nodes, run the verification scripts to confirm cluster readiness.

### Check Service Status

```bash
./status-deepseek-v4-flash-dspark.sh

```

This script inspects Docker container health, logs, and network connectivity across the head and registered workers.

### Run Smoke Test

```bash
./smoke-deepseek-v4-flash-dspark.sh

```

The smoke test issues a minimal inference request to the VLLM endpoint and validates JSON response format.

### Manual API Test

```python
import requests
import json

url = "http://<head-node-ip>:8000/v1/completions"

payload = {
    "model": "deepseek-v4-flash",
    "prompt": "Explain the DeepSeek-v4-Flash DSpark multi-node launch sequence.",
    "max_tokens": 128,
    "temperature": 0.7
}

response = requests.post(url, json=payload)
print(json.dumps(response.json(), indent=2))

```

Successful execution confirms end-to-end communication through the full stack: API → VLLM runtime → DSpark controller → distributed workers.

---

## Scaling the Cluster

To add workers beyond the initial deployment:

1. Provision additional GPU nodes with identical environment configuration
2. Run [`./start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/./start-tp3.sh) on each new node
3. The `dspark-controller` automatically discovers new workers through the shared Docker network

No head node reconfiguration is required. The DSpark service discovery mechanism handles dynamic worker registration.

---

## Key Configuration Files

| File | Purpose |
|------|---------|
| [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) | Head node orchestration script |
| [`start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-tp3.sh) | Worker node bootstrap script |
| [`status-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/status-deepseek-v4-flash-dspark.sh) | Container health inspection |
| [`smoke-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/smoke-deepseek-v4-flash-dspark.sh) | End-to-end inference validation |
| [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) | Complete service topology definition |
| [`dspark-numeric-knobs.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/dspark-numeric-knobs.sh) | NCCL/IB performance tuning |
| `.env.dspark.example` | Template for environment variables |

---

## Environment Prerequisites

Before executing the multi-node launch sequence, ensure these conditions:

- **Shared storage**: NFS mount accessible at identical path on all nodes (`/mnt/model-cache` by default)
- **Network connectivity**: Head node IP reachable from all workers on port 8000
- **Docker runtime**: Docker and Docker Compose installed with GPU runtime support
- **InfiniBand**: Physical IB connectivity between DGX nodes (handled by [`dspark-numeric-knobs.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/dspark-numeric-knobs.sh))

---

## Summary

- **Start the head first** with [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) to initialize the DSpark controller and VLLM runtime
- **Launch workers second** with [`start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-tp3.sh) on each GPU node, ensuring matching `DSK_CLUSTER_NAME` and NFS configuration
- **Verify with built-in scripts**: [`status-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/status-deepseek-v4-flash-dspark.sh) for health checks, [`smoke-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/smoke-deepseek-v4-flash-dspark.sh) for functional validation
- **Scale dynamically** by running the worker script on additional GPU nodes without head node changes

---

## Frequently Asked Questions

### What order must I start the head and worker nodes?

Always start the head node first. The [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) script creates the Docker network and DSpark controller that workers attempt to join. Starting workers before the head causes registration failures until the controller becomes available.

### Can I run the head and worker on the same physical machine?

Yes. Execute [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) first, then run [`start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-tp3.sh) on the same node. The worker container will bind to local GPU resources while communicating with the co-located controller through the Docker network.

### What happens if a worker node fails during operation?

The DSpark controller implemented in [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) tracks worker health through container status. Failed workers automatically unregister from the inference pool. Restart the worker with [`./start-tp3.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/./start-tp3.sh) to rejoin the cluster without head node intervention.

### How do I configure InfiniBand for multi-node GPU communication?

The [`dspark-numeric-knobs.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/dspark-numeric-knobs.sh) script sets NCCL_IB_DISABLE, NCCL_SOCKET_IFNAME, and related environment variables. Review and modify this file before launching workers if your IB topology requires custom routing or partition key settings.