DeepSeek-v4-Flash DSpark Multi-Node Launch Sequence: Head and Worker Node Setup

The DeepSeek-v4-Flash-DSpark multi-node launch sequence orchestrates a head node running start-deepseek-v4-flash-dspark.sh to initialize the controller stack, followed by worker nodes executing start-tp3.sh to register GPU workers with the cluster.

This guide explains the complete startup procedure for distributing the DeepSeek-v4-Flash inference engine across multiple DGX nodes using the DSpark orchestration framework. The MiaAI-Lab repository provides purpose-built shell scripts and Docker Compose configurations that automate the coordination between head (orchestrator) and worker (GPU compute) nodes.


Multi-Node Launch Sequence Overview

The repository implements a three-phase launch protocol designed for 2x DGX-Spark clusters. Each phase targets a specific node type with dedicated scripts that handle container orchestration, network binding, and service discovery.

Phase Node Type Action Script
1 Head Initialize controller, VLLM runtime, and auxiliary services start-deepseek-v4-flash-dspark.sh
2 Worker Launch GPU-bound VLLM worker containers bound to head network start-tp3.sh
3 Both Verify deployment health and run smoke tests status-deepseek-v4-flash-dspark.sh, smoke-deepseek-v4-flash-dspark.sh

Phase 1: Head Node Initialization

The head node serves as the cluster orchestrator, exposing the API endpoint and managing worker registration. In start-deepseek-v4-flash-dspark.sh, the script wraps docker compose -f docker-compose.dspark.yml up -d to bring up three critical services.

Services started on the head node:

  • dspark-controller — DSpark master that schedules and discovers workers
  • vllm-runtime — VLLM inference server with OpenAI-compatible /v1/completions endpoint
  • lmcache — Optional high-performance KV cache for accelerated serving

Head Node Start Command


# On the orchestrator machine

git clone https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark.git
cd DeepSeek-v4-Flash-DSpark-2x-DGX-Spark

# Configure environment

cp .env.dspark.example .env

# Edit .env: set DSK_CLUSTER_NAME, NFS_SHARE, head node IP, etc.

# Optional: pre-populate model weights on shared storage

./prepare-dspark-model-cache.sh

# Launch the head stack

./start-deepseek-v4-flash-dspark.sh

The docker-compose.dspark.yml defines a shared Docker network that enables automatic service discovery between head and worker containers across physical nodes.


Phase 2: Worker Node Registration

Worker nodes execute start-tp3.sh to join the cluster. This script configures the node as a VLLM worker instance and establishes GPU-to-GPU communication with the head.

Worker Node Start Command


# On each GPU worker machine

git clone https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark.git
cd DeepSeek-v4-Flash-DSpark-2x-DGX-Spark

# Use identical environment configuration

cp .env.dspark.example .env

# Ensure NFS_SHARE and DSK_CLUSTER_NAME match head node values

# Launch worker and register with head

./start-tp3.sh

Worker Configuration Details

The start-tp3.sh script performs these operations:

  1. Sets role identifier: Exports DSK_CLUSTER_ROLE=worker for runtime discovery
  2. Mounts model cache: Binds /mnt/model-cache via NFS for shared weight access
  3. Configures transport: Sources dspark-numeric-knobs.sh to set NCCL and InfiniBand parameters
  4. Network attachment: Joins the Docker network created by the head node's Compose stack

The dspark-numeric-knobs.sh file contains environment-specific tuning variables for high-throughput GPU communication across the DGX interconnect.


Phase 3: Health Verification and Testing

After launching all nodes, run the verification scripts to confirm cluster readiness.

Check Service Status

./status-deepseek-v4-flash-dspark.sh

This script inspects Docker container health, logs, and network connectivity across the head and registered workers.

Run Smoke Test

./smoke-deepseek-v4-flash-dspark.sh

The smoke test issues a minimal inference request to the VLLM endpoint and validates JSON response format.

Manual API Test

import requests
import json

url = "http://<head-node-ip>:8000/v1/completions"

payload = {
    "model": "deepseek-v4-flash",
    "prompt": "Explain the DeepSeek-v4-Flash DSpark multi-node launch sequence.",
    "max_tokens": 128,
    "temperature": 0.7
}

response = requests.post(url, json=payload)
print(json.dumps(response.json(), indent=2))

Successful execution confirms end-to-end communication through the full stack: API → VLLM runtime → DSpark controller → distributed workers.


Scaling the Cluster

To add workers beyond the initial deployment:

  1. Provision additional GPU nodes with identical environment configuration
  2. Run ./start-tp3.sh on each new node
  3. The dspark-controller automatically discovers new workers through the shared Docker network

No head node reconfiguration is required. The DSpark service discovery mechanism handles dynamic worker registration.


Key Configuration Files

File Purpose
start-deepseek-v4-flash-dspark.sh Head node orchestration script
start-tp3.sh Worker node bootstrap script
status-deepseek-v4-flash-dspark.sh Container health inspection
smoke-deepseek-v4-flash-dspark.sh End-to-end inference validation
docker-compose.dspark.yml Complete service topology definition
dspark-numeric-knobs.sh NCCL/IB performance tuning
.env.dspark.example Template for environment variables

Environment Prerequisites

Before executing the multi-node launch sequence, ensure these conditions:

  • Shared storage: NFS mount accessible at identical path on all nodes (/mnt/model-cache by default)
  • Network connectivity: Head node IP reachable from all workers on port 8000
  • Docker runtime: Docker and Docker Compose installed with GPU runtime support
  • InfiniBand: Physical IB connectivity between DGX nodes (handled by dspark-numeric-knobs.sh)

Summary


Frequently Asked Questions

What order must I start the head and worker nodes?

Always start the head node first. The start-deepseek-v4-flash-dspark.sh script creates the Docker network and DSpark controller that workers attempt to join. Starting workers before the head causes registration failures until the controller becomes available.

Can I run the head and worker on the same physical machine?

Yes. Execute start-deepseek-v4-flash-dspark.sh first, then run start-tp3.sh on the same node. The worker container will bind to local GPU resources while communicating with the co-located controller through the Docker network.

What happens if a worker node fails during operation?

The DSpark controller implemented in docker-compose.dspark.yml tracks worker health through container status. Failed workers automatically unregister from the inference pool. Restart the worker with ./start-tp3.sh to rejoin the cluster without head node intervention.

How do I configure InfiniBand for multi-node GPU communication?

The dspark-numeric-knobs.sh script sets NCCL_IB_DISABLE, NCCL_SOCKET_IFNAME, and related environment variables. Review and modify this file before launching workers if your IB topology requires custom routing or partition key settings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →