DeepSeek-v4-Flash DSpark Multi-Node Launch Sequence: Head and Worker Node Setup
The DeepSeek-v4-Flash-DSpark multi-node launch sequence orchestrates a head node running start-deepseek-v4-flash-dspark.sh to initialize the controller stack, followed by worker nodes executing start-tp3.sh to register GPU workers with the cluster.
This guide explains the complete startup procedure for distributing the DeepSeek-v4-Flash inference engine across multiple DGX nodes using the DSpark orchestration framework. The MiaAI-Lab repository provides purpose-built shell scripts and Docker Compose configurations that automate the coordination between head (orchestrator) and worker (GPU compute) nodes.
Multi-Node Launch Sequence Overview
The repository implements a three-phase launch protocol designed for 2x DGX-Spark clusters. Each phase targets a specific node type with dedicated scripts that handle container orchestration, network binding, and service discovery.
| Phase | Node Type | Action | Script |
|---|---|---|---|
| 1 | Head | Initialize controller, VLLM runtime, and auxiliary services | start-deepseek-v4-flash-dspark.sh |
| 2 | Worker | Launch GPU-bound VLLM worker containers bound to head network | start-tp3.sh |
| 3 | Both | Verify deployment health and run smoke tests | status-deepseek-v4-flash-dspark.sh, smoke-deepseek-v4-flash-dspark.sh |
Phase 1: Head Node Initialization
The head node serves as the cluster orchestrator, exposing the API endpoint and managing worker registration. In start-deepseek-v4-flash-dspark.sh, the script wraps docker compose -f docker-compose.dspark.yml up -d to bring up three critical services.
Services started on the head node:
dspark-controller— DSpark master that schedules and discovers workersvllm-runtime— VLLM inference server with OpenAI-compatible/v1/completionsendpointlmcache— Optional high-performance KV cache for accelerated serving
Head Node Start Command
# On the orchestrator machine
git clone https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark.git
cd DeepSeek-v4-Flash-DSpark-2x-DGX-Spark
# Configure environment
cp .env.dspark.example .env
# Edit .env: set DSK_CLUSTER_NAME, NFS_SHARE, head node IP, etc.
# Optional: pre-populate model weights on shared storage
./prepare-dspark-model-cache.sh
# Launch the head stack
./start-deepseek-v4-flash-dspark.sh
The docker-compose.dspark.yml defines a shared Docker network that enables automatic service discovery between head and worker containers across physical nodes.
Phase 2: Worker Node Registration
Worker nodes execute start-tp3.sh to join the cluster. This script configures the node as a VLLM worker instance and establishes GPU-to-GPU communication with the head.
Worker Node Start Command
# On each GPU worker machine
git clone https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark.git
cd DeepSeek-v4-Flash-DSpark-2x-DGX-Spark
# Use identical environment configuration
cp .env.dspark.example .env
# Ensure NFS_SHARE and DSK_CLUSTER_NAME match head node values
# Launch worker and register with head
./start-tp3.sh
Worker Configuration Details
The start-tp3.sh script performs these operations:
- Sets role identifier: Exports
DSK_CLUSTER_ROLE=workerfor runtime discovery - Mounts model cache: Binds
/mnt/model-cachevia NFS for shared weight access - Configures transport: Sources
dspark-numeric-knobs.shto set NCCL and InfiniBand parameters - Network attachment: Joins the Docker network created by the head node's Compose stack
The dspark-numeric-knobs.sh file contains environment-specific tuning variables for high-throughput GPU communication across the DGX interconnect.
Phase 3: Health Verification and Testing
After launching all nodes, run the verification scripts to confirm cluster readiness.
Check Service Status
./status-deepseek-v4-flash-dspark.sh
This script inspects Docker container health, logs, and network connectivity across the head and registered workers.
Run Smoke Test
./smoke-deepseek-v4-flash-dspark.sh
The smoke test issues a minimal inference request to the VLLM endpoint and validates JSON response format.
Manual API Test
import requests
import json
url = "http://<head-node-ip>:8000/v1/completions"
payload = {
"model": "deepseek-v4-flash",
"prompt": "Explain the DeepSeek-v4-Flash DSpark multi-node launch sequence.",
"max_tokens": 128,
"temperature": 0.7
}
response = requests.post(url, json=payload)
print(json.dumps(response.json(), indent=2))
Successful execution confirms end-to-end communication through the full stack: API → VLLM runtime → DSpark controller → distributed workers.
Scaling the Cluster
To add workers beyond the initial deployment:
- Provision additional GPU nodes with identical environment configuration
- Run
./start-tp3.shon each new node - The
dspark-controllerautomatically discovers new workers through the shared Docker network
No head node reconfiguration is required. The DSpark service discovery mechanism handles dynamic worker registration.
Key Configuration Files
| File | Purpose |
|---|---|
start-deepseek-v4-flash-dspark.sh |
Head node orchestration script |
start-tp3.sh |
Worker node bootstrap script |
status-deepseek-v4-flash-dspark.sh |
Container health inspection |
smoke-deepseek-v4-flash-dspark.sh |
End-to-end inference validation |
docker-compose.dspark.yml |
Complete service topology definition |
dspark-numeric-knobs.sh |
NCCL/IB performance tuning |
.env.dspark.example |
Template for environment variables |
Environment Prerequisites
Before executing the multi-node launch sequence, ensure these conditions:
- Shared storage: NFS mount accessible at identical path on all nodes (
/mnt/model-cacheby default) - Network connectivity: Head node IP reachable from all workers on port 8000
- Docker runtime: Docker and Docker Compose installed with GPU runtime support
- InfiniBand: Physical IB connectivity between DGX nodes (handled by
dspark-numeric-knobs.sh)
Summary
- Start the head first with
start-deepseek-v4-flash-dspark.shto initialize the DSpark controller and VLLM runtime - Launch workers second with
start-tp3.shon each GPU node, ensuring matchingDSK_CLUSTER_NAMEand NFS configuration - Verify with built-in scripts:
status-deepseek-v4-flash-dspark.shfor health checks,smoke-deepseek-v4-flash-dspark.shfor functional validation - Scale dynamically by running the worker script on additional GPU nodes without head node changes
Frequently Asked Questions
What order must I start the head and worker nodes?
Always start the head node first. The start-deepseek-v4-flash-dspark.sh script creates the Docker network and DSpark controller that workers attempt to join. Starting workers before the head causes registration failures until the controller becomes available.
Can I run the head and worker on the same physical machine?
Yes. Execute start-deepseek-v4-flash-dspark.sh first, then run start-tp3.sh on the same node. The worker container will bind to local GPU resources while communicating with the co-located controller through the Docker network.
What happens if a worker node fails during operation?
The DSpark controller implemented in docker-compose.dspark.yml tracks worker health through container status. Failed workers automatically unregister from the inference pool. Restart the worker with ./start-tp3.sh to rejoin the cluster without head node intervention.
How do I configure InfiniBand for multi-node GPU communication?
The dspark-numeric-knobs.sh script sets NCCL_IB_DISABLE, NCCL_SOCKET_IFNAME, and related environment variables. Review and modify this file before launching workers if your IB topology requires custom routing or partition key settings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →