# How to Orchestrate the DeepSeek‑v4‑Flash DSpark Docker Compose Stack: A Complete Guide to `start‑deepseek‑v4‑flash‑dspark.sh`

> Learn how the start-deepseek-v4-flash-dspark.sh script orchestrates multi-node Docker Compose deployments. This guide covers configuration, environment validation, and SSH coordination for your DSpark stack.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: how-to-guide
- Published: 2026-09-09

---

**The `start‑deepseek‑v4‑flash‑dspark.sh` script orchestrates multi‑node Docker Compose deployments by loading centralized configuration, validating environment variables, resolving NCCL network settings, and executing coordinated `docker compose up` commands across head and worker nodes via SSH.**

This article breaks down exactly how the MiaAI‑Lab `start‑deepseek‑v4‑flash‑dspark.sh` bootstrap script works. You'll learn how it handles everything from HF cache staging to RoCE GID resolution, and how it ultimately launches the distributed DeepSeek‑v4‑Flash inference stack on DGX‑2 × 2 (or × 3) clusters.

---

## Prerequisites and Configuration Loading

The script begins by establishing its execution context and loading environment configuration.

### Script Directory Resolution

At lines 4–8, the script determines its own location to enable relative path references:

```bash
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"

```

This ensures all subsequent file operations—sourcing helper scripts, reading compose files, and staging assets—use consistent absolute paths regardless of where the user invokes the command.

### Environment File Validation

The script requires a per‑node `.env.dspark` file containing cluster‑specific settings. At lines 68–71, it validates this file exists and aborts with a descriptive error if missing:

```bash
ENV_FILE="$SCRIPT_DIR/.env.dspark"
if [ ! -f "$ENV_FILE" ]; then
    echo "ERROR: Environment file $ENV_FILE not found" >&2
    exit 1
fi

```

A sanitized copy `_dspark_env_clean` is generated to prevent variable pollution during remote execution.

---

## Validating Numeric Knobs and API Keys

Before any containers launch, the script enforces configuration correctness.

### Performance Knob Validation

Lines 88–90 source `dspark‑numeric‑knobs.sh` and immediately invoke validation:

```bash
source "$SCRIPT_DIR/dspark-numeric-knobs.sh"
dspark_validate_numeric_knobs || exit 1

```

This helper validates critical parameters like `MAX_NUM_SEQS`, `GPU_MEMORY_UTILIZATION`, and tensor‑parallel sizing—catching misconfigurations before they cause runtime failures.

### API Key Consistency Check

Lines 15–30 prevent conflicting authentication setups by verifying mutual exclusivity between `DSPARK_API_KEYS` and `VLLM_API_KEY`:

```bash
case "${DSPARK_API_KEYS:-}" in
    *[! ,a-zA-Z0-9_-]*)
        echo "ERROR: DSPARK_API_KEYS contains invalid characters" >&2
        exit 1
        ;;
esac

```

If both variables are set, the script aborts to prevent ambiguous authentication behavior in the vLLM service layer.

---

## Optional Feature: Ablation Direction Staging

### Gated Model Asset Preparation

When `ABLATE=1`, the script at lines 78–109 stages the proprietary `direction_r1.pt` file from HuggingFace into the local cache and replicates it to all workers:

```bash
stage_ablation_direction() {
    local direction_url="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/resolve/main/direction_r1.pt"
    local cache_dir="${HF_CACHE}/dspark-ablation"
    # SHA-256 verification, download, and worker distribution

}

```

SHA‑256 verification ensures bit‑exact replication across the cluster, preventing silent corruption that would cause inference divergence.

---

## Determining Tensor Parallelism and Node Topology

### TP Size and Node Role Resolution

Lines 58–64 detect the execution mode by checking `DSPARK_TP3` (set by `./start‑tp3.sh`):

```bash
if [ "${DSPARK_TP3:-0}" = "1" ]; then
    TP_SIZE=3
    NNODES=3
    WORKER2_HOST="${WORKER2_HOST:-worker2}"
else
    TP_SIZE=2
    NNODES=2
fi

```

This conditional export drives compose file selection and remote command generation, enabling seamless switching between 2‑node and 3‑node configurations without code changes.

---

## NCCL RoCE v2 GID Resolution

### Network Interface Index Computation

The `resolve_nccl_gid_indexes()` function (lines 46–121) handles InfiniBand GID selection for optimal NCCL performance:

- **Auto‑detection mode** (`NCCL_IB_GID_AUTO=1`): Validates that selected NICs expose usable RoCE v2 GIDs
- **Static mode**: Pins indexes from explicit `.env.dspark` values

This resolution is critical for DGX‑2 × 2 deployments where incorrect GID selection causes 10–100× slowdown in all‑reduce operations.

---

## NFS Shared Cache Preparation (Optional)

### Conditional Compose Override Generation

When `DSPARK_WORKER_HF_NFS=1` (lines 26–36), the script:

1. Discovers the NFS server head‑node IP via `files/nfs‑share.sh`
2. Generates `docker‑compose.dspark‑nfs.override.yml` with shared volume mounts
3. Appends this override to `WORKER_COMPOSE_FILES`

This eliminates per‑node model downloads, reducing startup time from 30+ minutes to under 2 minutes for 70B+ parameter models.

---

## Remote Command Assembly and SSH Distribution

### Building Worker‑Side Compose Commands

Lines 40–45 construct sanitized remote execution environments:

```bash
REMOTE_COMPOSE="cd $REMOTE_WORKER_DIR && \
    env -u MASTER_ADDR -u MASTER_PORT \
    COMPOSE_DISABLE_ENV_FILE=1 \
    docker compose -f docker-compose.dspark.yml ${WORKER_COMPOSE_FILES} up -d"

```

Key sanitization measures:

| Measure | Purpose |
|---------|---------|
| `env -u MASTER_ADDR -u MASTER_PORT` | Prevents head‑node coordination variables from leaking to workers |
| `COMPOSE_DISABLE_ENV_FILE=1` | Forces use of generated `.env.dspark` over stale files |
| `REMOTE_*` variable prefixing | Transports hot‑fix flags, NCCL settings, and ablation state |

---

## Executing the Docker Compose Stack Launch

### Coordinated Multi‑Node Deployment

The final orchestration phase (approximately lines 1200–1300) executes:

```bash

# Head node

docker compose -f "$COMPOSE_FILE" up -d

# Workers (parallel SSH)

ssh "$WORKER_HOST" "$REMOTE_COMPOSE"
[ "$TP_SIZE" = 3 ] && ssh "$WORKER2_HOST" "$REMOTE_COMPOSE"

```

The actual command pattern in the source:

```bash
docker compose -f "$COMPOSE_FILE" $WORKER_COMPOSE_FILES up -d

```

Post‑launch, the script polls the vLLM health endpoint (`/health`) with configurable retries (`WAIT_ATTEMPTS × WAIT_SECONDS`) before declaring the stack ready.

---

## Post‑Start Diagnostics and Logging

### Service Availability Reporting

Upon successful launch, the script emits:

```bash
echo "API is available at $API_URL"
echo "Chat completions: $CHAT_URL/v1/chat/completions"

```

With `--log-since`, it attaches to head container logs for real‑time startup monitoring:

```bash
docker logs -f --since "$LOG_SINCE" "${PROJECT_NAME}-vllm-1"

```

---

## Complete Usage Examples

### Basic 2‑Node Startup

```bash
./start-deepseek-v4-flash-dspark.sh

```

Uses `.env.dspark` defaults, TP=2, localhost:8888.

### Custom Host and Port

```bash
./start-deepseek-v4-flash-dspark.sh --host 10.1.2.3 --port 9000

```

CLI overrides take precedence over environment files.

### 3‑Node Tensor Parallel

```bash
./start-tp3.sh  # Exports DSPARK_TP3=1

./start-deepseek-v4-flash-dspark.sh

```

Automatically includes `worker2` and adjusts NCCL topology.

### Ablation Mode with NFS Cache

```bash
DSPARK_WORKER_HF_NFS=1 ABLATE=1 \
    ./start-deepseek-v4-flash-dspark.sh --log-since "5m"

```

Enables gated model features with shared model storage and live log tailing.

---

## Key Source Files Reference

| File | Lines | Function |
|------|-------|----------|
| `start‑deepseek‑v4‑flash‑dspark.sh` | 4–8 | Script directory resolution |
| `start‑deepseek‑v4‑flash‑dspark.sh` | 15–30 | API key validation |
| `start‑deepseek‑v4‑flash‑dspark.sh` | 26–36 | NFS override preparation |
| `start‑deepseek‑v4‑flash‑dspark.sh` | 40–45 | Remote command assembly |
| `start‑deepseek‑v4‑flash‑dspark.sh` | 46–121 | `resolve_nccl_gid_indexes()` |
| `start‑deepseek‑v4‑flash‑dspark.sh` | 58–64 | TP size and node detection |
| `start‑deepseek‑v4‑flash‑dspark.sh` | 68–71 | Environment file validation |
| `start‑deepseek‑v4‑flash‑dspark.sh` | 78–109 | `stage_ablation_direction()` |
| `start‑deepseek‑v4‑flash‑dspark.sh` | 88–90 | Numeric knob validation |
| `start‑deepseek‑v4‑flash‑dspark.sh` | ~1200–1300 | `docker compose up` execution |
| `docker‑compose.dspark.yml` | Full | Core service definitions |
| `dspark‑numeric‑knobs.sh` | Full | Performance parameter validation |

---

## Summary

- **Configuration centralization**: `.env.dspark` drives all per‑node behavior with CLI overrides for ad‑hoc changes
- **Validation‑first design**: Numeric knobs, API keys, and file assets are verified before any container starts
- **Network intelligence**: Automatic RoCE v2 GID resolution optimizes DGX‑2 NCCL performance
- **Modular compose assembly**: Conditional overrides (NFS, TP=3, ablations) stack cleanly via `WORKER_COMPOSE_FILES`
- **Sanitized remote execution**: Environment variable isolation prevents configuration leakage across SSH boundaries
- **Health‑aware startup**: Polling‑based readiness detection ensures clients receive a functional API endpoint

---

## Frequently Asked Questions

### What does `COMPOSE_DISABLE_ENV_FILE=1` prevent?

This environment variable stops Docker Compose from automatically loading `.env` files in the working directory. According to the MiaAI‑Lab source code, this ensures workers use the explicitly generated `.env.dspark` rather than stale or conflicting environment files that might exist on remote nodes.

### How does the script handle the third node in TP=3 mode?

When `DSPARK_TP3=1` is exported (via `./start‑tp3.sh`), lines 58–64 set `TP_SIZE=3` and `NNODES=3`, export `WORKER2_HOST`, and conditionally invoke SSH deployment to the third node. The same `REMOTE_COMPOSE` command template is used, maintaining consistency with the 2‑node code path.

### Why separate API key variables (`DSPARK_API_KEYS` vs `VLLM_API_KEY`)?

The mutual exclusivity check at lines 15–30 prevents ambiguous authentication: `DSPARK_API_KEYS` enables multi‑key rotation for load‑balanced deployments, while `VLLM_API_KEY` sets a single key in the underlying vLLM engine. Using both would create undefined behavior in request routing.

### Can I run the script without SSH access to workers?

No. The `start‑deepseek‑v4‑flash‑dspark.sh` architecture assumes passwordless SSH availability to `WORKER_HOST` and optionally `WORKER2_HOST`. The remote compose execution at lines ~1200–1300 depends on this for coordinated multi‑node container lifecycle management.