What is docker-compose.dspark.yml? The DeepSeek V4 Flash Deployment Manifest

docker-compose.dspark.yml is the central Docker Compose manifest that defines and orchestrates the DSpark-vLLM inference service for the DeepSeek V4 Flash model, exposing it as an HTTP API while managing GPU resources, persistent storage, and distributed multi-node configuration.

The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository provides a production-ready containerized inference stack for the DeepSeek V4 Flash model. At its core lies docker-compose.dspark.yml, a specialized Docker Compose file that transforms the ghcr.io/anemll/dspark-vllm-gx10:0.1.1 container image into a fully configured inference server with support for speculative decoding, selective runtime hot-fixes, and Tensor-Parallel distributed execution.

Core Responsibilities of docker-compose.dspark.yml

The manifest encapsulates six critical functions that turn a raw container into a production inference engine.

Service Definition and Container Orchestration

The file declares a single service named vllm-dspark that pulls the ghcr.io/anemll/dspark-vllm-gx10:0.1.1 image. According to the source code at docker-compose.dspark.yml line 1, this service acts as the exclusive runtime for the DSpark-vLLM binary, isolating the inference environment while exposing it to the host network stack.

GPU and Memory Resource Allocation

The configuration grants the container aggressive hardware access required for large language model inference. It specifies gpus: all to expose every available GPU, allocates shm_size: "64gb" for substantial shared memory segments needed by NCCL collectives, and sets network_mode: host to eliminate container networking overhead (as defined at line 35). These settings are mandatory for multi-GPU Tensor-Parallel communication.

Persistent Storage and Hot-Fix Mounts

Three critical volume mounts ensure state persistence and runtime flexibility (defined at line 40):

  • HuggingFace cache directory: Stores downloaded model weights to survive container restarts
  • Stage-C overlay: Optional model file overrides for experimental weights
  • patches/ directory: Runtime hot-fix scripts that are selectively injected into the Python environment based on feature flags

Environment Configuration and Runtime Tuning

The manifest exposes an extensive environment variable schema prefixed with DSPARK_* and VLLM_* (defined at line 21). These variables control:

  • Model loading: DSPARK_MODEL specifies the HuggingFace model identifier
  • Speculative decoding: Token budgeting and draft model configuration
  • Selective hot-fixes: Boolean toggles like DSPARK_ENABLE_ISSUE31_GPU_HOTFIX and DSPARK_ENABLE_ADAPTIVE_CHUNK activate specific Python patches from the mounted patches/ directory without rebuilding the image

Startup Scripting and Entrypoint Logic

The container executes a Bash entrypoint defined at line 14 that performs four sequential operations:

  1. Resolves API key arguments from environment variables
  2. Applies selected hot-fix Python scripts by copying them into the vLLM installation path
  3. Builds the speculative decoding configuration JSON
  4. Executes vllm serve with dynamically generated CLI flags derived from the DSPARK_* environment variables

This indirection allows operators to modify inference behavior via environment variables rather than manual command-line editing.

Distributed Inference and NCCL Networking

For multi-node deployments, the manifest passes standard PyTorch distributed environment variables (defined at line 60):

  • NNODES: Total number of participating nodes
  • NODE_RANK: Zero-indexed identifier for the current node
  • MASTER_ADDR and MASTER_PORT: Rendezvous coordinates for the NCCL process group

These settings enable Tensor-Parallel inference across multiple DGX or GX10 systems using the host network stack for RDMA-capable communication.

How to Deploy DSpark Using docker-compose.dspark.yml

Start the inference service with a single command that consumes the manifest:


# Start with default Tensor-Parallel size 2 across 2 nodes

docker compose -f docker-compose.dspark.yml up -d

Override the default model or activate GPU hot-fixes via environment prefixes:

DSPARK_MODEL=deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
DSPARK_ENABLE_ISSUE31_GPU_HOTFIX=1 \
docker compose -f docker-compose.dspark.yml up -d

Query the HTTP API exposed on port 8888:

curl http://127.0.0.1:8888/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek-v4-flash-vision-exp","messages":[{"role":"user","content":"Explain DSpark"}]}'

Integration with the DSpark Ecosystem

The docker-compose.dspark.yml file functions as the glue binding several repository components into a coherent deployment:

  • Dockerfile.gb10-dsv4-dspark: Provides the base image containing the compiled vLLM binary with DSpark-specific optimizations
  • patches/: Contains Python and Bash hot-fixes mounted at runtime; the entrypoint selectively applies these based on DSPARK_ENABLE_* flags
  • vllm_patch_gb10/: Optional GB10 hybrid plugin tree referenced by VLLM_GB10_PATCH_DIR for advanced mixed-precision kernels

Summary

  • docker-compose.dspark.yml serves as the single-source-of-truth for deploying the DSpark-vLLM inference service.
  • It configures the vllm-dspark container with full GPU access, 64GB shared memory, and host networking for optimal NCCL performance.
  • The manifest supports selective hot-fixes via environment variables like DSPARK_ENABLE_ISSUE31_GPU_HOTFIX without requiring image rebuilds.
  • Distributed multi-node inference is enabled through standard NNODES and NODE_RANK environment variable injection.
  • A single docker compose -f docker-compose.dspark.yml up command launches a production-ready HTTP API on port 8888.

Frequently Asked Questions

How do I enable specific GPU hot-fixes when deploying?

Set the corresponding DSPARK_ENABLE_* environment variable to 1 before launching the compose stack. For example, DSPARK_ENABLE_ISSUE31_GPU_HOTFIX=1 activates the Issue 31 GPU patch, while DSPARK_ENABLE_ADAPTIVE_CHUNK=1 enables adaptive chunking optimizations. The entrypoint script at line 14 detects these flags and copies the relevant Python files from the mounted patches/ directory into the active vLLM installation.

What network port does the inference service expose?

The service exposes an OpenAI-compatible HTTP API on port 8888. Because the manifest sets network_mode: host at line 35, the container binds directly to the host's network interface, making the service accessible at http://127.0.0.1:8888 without port mapping conflicts or Docker NAT overhead.

Can I run this configuration on a single GPU instead of multi-node?

Yes, adjust the DSPARK_TENSOR_PARALLEL_SIZE and NNODES environment variables accordingly. For single-GPU inference, set NNODES=1 and DSPARK_TENSOR_PARALLEL_SIZE=1. The gpus: all directive will still expose all GPUs to the container, but vLLM will only utilize the count specified in the tensor parallelism setting.

Where are the DeepSeek model weights stored between restarts?

The manifest mounts a host directory to the HuggingFace cache path inside the container (defined at line 40). This ensures that downloaded weights for deepseek-ai/DeepSeek-V4-Flash or variant models persist on the host filesystem, eliminating redundant downloads and reducing container startup time after the initial pull.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →