What is docker-compose.dspark.yml? The DeepSeek V4 Flash Deployment Manifest
docker-compose.dspark.yml is the central Docker Compose manifest that defines and orchestrates the DSpark-vLLM inference service for the DeepSeek V4 Flash model, exposing it as an HTTP API while managing GPU resources, persistent storage, and distributed multi-node configuration.
The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository provides a production-ready containerized inference stack for the DeepSeek V4 Flash model. At its core lies docker-compose.dspark.yml, a specialized Docker Compose file that transforms the ghcr.io/anemll/dspark-vllm-gx10:0.1.1 container image into a fully configured inference server with support for speculative decoding, selective runtime hot-fixes, and Tensor-Parallel distributed execution.
Core Responsibilities of docker-compose.dspark.yml
The manifest encapsulates six critical functions that turn a raw container into a production inference engine.
Service Definition and Container Orchestration
The file declares a single service named vllm-dspark that pulls the ghcr.io/anemll/dspark-vllm-gx10:0.1.1 image. According to the source code at docker-compose.dspark.yml line 1, this service acts as the exclusive runtime for the DSpark-vLLM binary, isolating the inference environment while exposing it to the host network stack.
GPU and Memory Resource Allocation
The configuration grants the container aggressive hardware access required for large language model inference. It specifies gpus: all to expose every available GPU, allocates shm_size: "64gb" for substantial shared memory segments needed by NCCL collectives, and sets network_mode: host to eliminate container networking overhead (as defined at line 35). These settings are mandatory for multi-GPU Tensor-Parallel communication.
Persistent Storage and Hot-Fix Mounts
Three critical volume mounts ensure state persistence and runtime flexibility (defined at line 40):
- HuggingFace cache directory: Stores downloaded model weights to survive container restarts
- Stage-C overlay: Optional model file overrides for experimental weights
patches/directory: Runtime hot-fix scripts that are selectively injected into the Python environment based on feature flags
Environment Configuration and Runtime Tuning
The manifest exposes an extensive environment variable schema prefixed with DSPARK_* and VLLM_* (defined at line 21). These variables control:
- Model loading:
DSPARK_MODELspecifies the HuggingFace model identifier - Speculative decoding: Token budgeting and draft model configuration
- Selective hot-fixes: Boolean toggles like
DSPARK_ENABLE_ISSUE31_GPU_HOTFIXandDSPARK_ENABLE_ADAPTIVE_CHUNKactivate specific Python patches from the mountedpatches/directory without rebuilding the image
Startup Scripting and Entrypoint Logic
The container executes a Bash entrypoint defined at line 14 that performs four sequential operations:
- Resolves API key arguments from environment variables
- Applies selected hot-fix Python scripts by copying them into the vLLM installation path
- Builds the speculative decoding configuration JSON
- Executes
vllm servewith dynamically generated CLI flags derived from theDSPARK_*environment variables
This indirection allows operators to modify inference behavior via environment variables rather than manual command-line editing.
Distributed Inference and NCCL Networking
For multi-node deployments, the manifest passes standard PyTorch distributed environment variables (defined at line 60):
NNODES: Total number of participating nodesNODE_RANK: Zero-indexed identifier for the current nodeMASTER_ADDRandMASTER_PORT: Rendezvous coordinates for the NCCL process group
These settings enable Tensor-Parallel inference across multiple DGX or GX10 systems using the host network stack for RDMA-capable communication.
How to Deploy DSpark Using docker-compose.dspark.yml
Start the inference service with a single command that consumes the manifest:
# Start with default Tensor-Parallel size 2 across 2 nodes
docker compose -f docker-compose.dspark.yml up -d
Override the default model or activate GPU hot-fixes via environment prefixes:
DSPARK_MODEL=deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
DSPARK_ENABLE_ISSUE31_GPU_HOTFIX=1 \
docker compose -f docker-compose.dspark.yml up -d
Query the HTTP API exposed on port 8888:
curl http://127.0.0.1:8888/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"deepseek-v4-flash-vision-exp","messages":[{"role":"user","content":"Explain DSpark"}]}'
Integration with the DSpark Ecosystem
The docker-compose.dspark.yml file functions as the glue binding several repository components into a coherent deployment:
Dockerfile.gb10-dsv4-dspark: Provides the base image containing the compiled vLLM binary with DSpark-specific optimizationspatches/: Contains Python and Bash hot-fixes mounted at runtime; the entrypoint selectively applies these based onDSPARK_ENABLE_*flagsvllm_patch_gb10/: Optional GB10 hybrid plugin tree referenced byVLLM_GB10_PATCH_DIRfor advanced mixed-precision kernels
Summary
- docker-compose.dspark.yml serves as the single-source-of-truth for deploying the DSpark-vLLM inference service.
- It configures the
vllm-dsparkcontainer with full GPU access, 64GB shared memory, and host networking for optimal NCCL performance. - The manifest supports selective hot-fixes via environment variables like
DSPARK_ENABLE_ISSUE31_GPU_HOTFIXwithout requiring image rebuilds. - Distributed multi-node inference is enabled through standard
NNODESandNODE_RANKenvironment variable injection. - A single
docker compose -f docker-compose.dspark.yml upcommand launches a production-ready HTTP API on port 8888.
Frequently Asked Questions
How do I enable specific GPU hot-fixes when deploying?
Set the corresponding DSPARK_ENABLE_* environment variable to 1 before launching the compose stack. For example, DSPARK_ENABLE_ISSUE31_GPU_HOTFIX=1 activates the Issue 31 GPU patch, while DSPARK_ENABLE_ADAPTIVE_CHUNK=1 enables adaptive chunking optimizations. The entrypoint script at line 14 detects these flags and copies the relevant Python files from the mounted patches/ directory into the active vLLM installation.
What network port does the inference service expose?
The service exposes an OpenAI-compatible HTTP API on port 8888. Because the manifest sets network_mode: host at line 35, the container binds directly to the host's network interface, making the service accessible at http://127.0.0.1:8888 without port mapping conflicts or Docker NAT overhead.
Can I run this configuration on a single GPU instead of multi-node?
Yes, adjust the DSPARK_TENSOR_PARALLEL_SIZE and NNODES environment variables accordingly. For single-GPU inference, set NNODES=1 and DSPARK_TENSOR_PARALLEL_SIZE=1. The gpus: all directive will still expose all GPUs to the container, but vLLM will only utilize the count specified in the tensor parallelism setting.
Where are the DeepSeek model weights stored between restarts?
The manifest mounts a host directory to the HuggingFace cache path inside the container (defined at line 40). This ensures that downloaded weights for deepseek-ai/DeepSeek-V4-Flash or variant models persist on the host filesystem, eliminating redundant downloads and reducing container startup time after the initial pull.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →