# What is docker-compose.dspark.yml? The DeepSeek V4 Flash Deployment Manifest

> Understand docker-compose.dspark.yml, the manifest for orchestrating the DeepSeek V4 Flash DSpark inference service. Learn how it manages GPU resources, storage, and multi-node deployments.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: how-to-guide
- Published: 2026-09-09

---

**docker-compose.dspark.yml is the central Docker Compose manifest that defines and orchestrates the DSpark-vLLM inference service for the DeepSeek V4 Flash model, exposing it as an HTTP API while managing GPU resources, persistent storage, and distributed multi-node configuration.**

The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository provides a production-ready containerized inference stack for the DeepSeek V4 Flash model. At its core lies [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml), a specialized Docker Compose file that transforms the `ghcr.io/anemll/dspark-vllm-gx10:0.1.1` container image into a fully configured inference server with support for speculative decoding, selective runtime hot-fixes, and Tensor-Parallel distributed execution.

## Core Responsibilities of docker-compose.dspark.yml

The manifest encapsulates six critical functions that turn a raw container into a production inference engine.

### Service Definition and Container Orchestration

The file declares a single service named `vllm-dspark` that pulls the `ghcr.io/anemll/dspark-vllm-gx10:0.1.1` image. According to the source code at [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) line 1, this service acts as the exclusive runtime for the DSpark-vLLM binary, isolating the inference environment while exposing it to the host network stack.

### GPU and Memory Resource Allocation

The configuration grants the container aggressive hardware access required for large language model inference. It specifies `gpus: all` to expose every available GPU, allocates `shm_size: "64gb"` for substantial shared memory segments needed by NCCL collectives, and sets `network_mode: host` to eliminate container networking overhead (as defined at line 35). These settings are mandatory for multi-GPU Tensor-Parallel communication.

### Persistent Storage and Hot-Fix Mounts

Three critical volume mounts ensure state persistence and runtime flexibility (defined at line 40):

- **HuggingFace cache directory**: Stores downloaded model weights to survive container restarts
- **Stage-C overlay**: Optional model file overrides for experimental weights
- **`patches/` directory**: Runtime hot-fix scripts that are selectively injected into the Python environment based on feature flags

### Environment Configuration and Runtime Tuning

The manifest exposes an extensive environment variable schema prefixed with `DSPARK_*` and `VLLM_*` (defined at line 21). These variables control:

- **Model loading**: `DSPARK_MODEL` specifies the HuggingFace model identifier
- **Speculative decoding**: Token budgeting and draft model configuration
- **Selective hot-fixes**: Boolean toggles like `DSPARK_ENABLE_ISSUE31_GPU_HOTFIX` and `DSPARK_ENABLE_ADAPTIVE_CHUNK` activate specific Python patches from the mounted `patches/` directory without rebuilding the image

### Startup Scripting and Entrypoint Logic

The container executes a Bash entrypoint defined at line 14 that performs four sequential operations:

1. Resolves API key arguments from environment variables
2. Applies selected hot-fix Python scripts by copying them into the vLLM installation path
3. Builds the speculative decoding configuration JSON
4. Executes `vllm serve` with dynamically generated CLI flags derived from the `DSPARK_*` environment variables

This indirection allows operators to modify inference behavior via environment variables rather than manual command-line editing.

### Distributed Inference and NCCL Networking

For multi-node deployments, the manifest passes standard PyTorch distributed environment variables (defined at line 60):

- `NNODES`: Total number of participating nodes
- `NODE_RANK`: Zero-indexed identifier for the current node
- `MASTER_ADDR` and `MASTER_PORT`: Rendezvous coordinates for the NCCL process group

These settings enable Tensor-Parallel inference across multiple DGX or GX10 systems using the host network stack for RDMA-capable communication.

## How to Deploy DSpark Using docker-compose.dspark.yml

Start the inference service with a single command that consumes the manifest:

```bash

# Start with default Tensor-Parallel size 2 across 2 nodes

docker compose -f docker-compose.dspark.yml up -d

```

Override the default model or activate GPU hot-fixes via environment prefixes:

```bash
DSPARK_MODEL=deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
DSPARK_ENABLE_ISSUE31_GPU_HOTFIX=1 \
docker compose -f docker-compose.dspark.yml up -d

```

Query the HTTP API exposed on port 8888:

```bash
curl http://127.0.0.1:8888/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek-v4-flash-vision-exp","messages":[{"role":"user","content":"Explain DSpark"}]}'

```

## Integration with the DSpark Ecosystem

The [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) file functions as the glue binding several repository components into a coherent deployment:

- **`Dockerfile.gb10-dsv4-dspark`**: Provides the base image containing the compiled vLLM binary with DSpark-specific optimizations
- **`patches/`**: Contains Python and Bash hot-fixes mounted at runtime; the entrypoint selectively applies these based on `DSPARK_ENABLE_*` flags
- **`vllm_patch_gb10/`**: Optional GB10 hybrid plugin tree referenced by `VLLM_GB10_PATCH_DIR` for advanced mixed-precision kernels

## Summary

- **docker-compose.dspark.yml** serves as the single-source-of-truth for deploying the DSpark-vLLM inference service.
- It configures the `vllm-dspark` container with full GPU access, 64GB shared memory, and host networking for optimal NCCL performance.
- The manifest supports selective hot-fixes via environment variables like `DSPARK_ENABLE_ISSUE31_GPU_HOTFIX` without requiring image rebuilds.
- Distributed multi-node inference is enabled through standard `NNODES` and `NODE_RANK` environment variable injection.
- A single `docker compose -f docker-compose.dspark.yml up` command launches a production-ready HTTP API on port 8888.

## Frequently Asked Questions

### How do I enable specific GPU hot-fixes when deploying?

Set the corresponding `DSPARK_ENABLE_*` environment variable to `1` before launching the compose stack. For example, `DSPARK_ENABLE_ISSUE31_GPU_HOTFIX=1` activates the Issue 31 GPU patch, while `DSPARK_ENABLE_ADAPTIVE_CHUNK=1` enables adaptive chunking optimizations. The entrypoint script at line 14 detects these flags and copies the relevant Python files from the mounted `patches/` directory into the active vLLM installation.

### What network port does the inference service expose?

The service exposes an OpenAI-compatible HTTP API on **port 8888**. Because the manifest sets `network_mode: host` at line 35, the container binds directly to the host's network interface, making the service accessible at `http://127.0.0.1:8888` without port mapping conflicts or Docker NAT overhead.

### Can I run this configuration on a single GPU instead of multi-node?

Yes, adjust the `DSPARK_TENSOR_PARALLEL_SIZE` and `NNODES` environment variables accordingly. For single-GPU inference, set `NNODES=1` and `DSPARK_TENSOR_PARALLEL_SIZE=1`. The `gpus: all` directive will still expose all GPUs to the container, but vLLM will only utilize the count specified in the tensor parallelism setting.

### Where are the DeepSeek model weights stored between restarts?

The manifest mounts a host directory to the HuggingFace cache path inside the container (defined at line 40). This ensures that downloaded weights for `deepseek-ai/DeepSeek-V4-Flash` or variant models persist on the host filesystem, eliminating redundant downloads and reducing container startup time after the initial pull.