DeepSeek-v4-Flash-DSpark MAX_NUM_SEQS: Maximum Concurrency Explained
The DeepSeek-v4-Flash-DSpark inference engine supports up to 16 concurrent sequences when configured with --max-num-seqs 16, with a conservative default of 6 slots to balance memory usage and throughput.
DeepSeek-v4-Flash-DSpark implements sequence-level parallelism through a configurable concurrency limit controlled by MAX_NUM_SEQS. This parameter determines how many requests the vLLM-based scheduler can process simultaneously, directly impacting throughput and GPU memory utilization. Understanding how to tune this value is essential for optimizing inference performance on multi-GPU DGX deployments.
How MAX_NUM_SEQS Controls Concurrency
The MAX_NUM_SEQS parameter defines the number of active sequence slots available in the KV cache manager. Each concurrent sequence requires dedicated GPU memory for key-value tensors, making this setting a critical trade-off between throughput and memory consumption.
In validate-dspark-config.sh, the default value is explicitly set:
# Lines 149-153 in validate-dspark-config.sh
if [ -z "$MAX_NUM_SEQS" ]; then
echo "MAX_NUM_SEQS not set, using default: 6"
MAX_NUM_SEQS=6
else
echo "MAX_NUM_SEQS: $MAX_NUM_SEQS"
fi
This fallback ensures predictable memory usage out of the box, preventing out-of-memory errors on standard configurations.
Default vs. Maximum Supported Values
| Configuration | Value | Source |
|---|---|---|
| Default | 6 | validate-dspark-config.sh (lines 149-153) |
| Documented maximum example | 16 | README.md (lines 534-580) |
| Hard upper bound | None enforced | Runtime-limited by GPU memory |
The repository does not implement a hardcoded maximum. The practical ceiling emerges from KV cache capacity: each additional sequence multiplies memory requirements by sequence length, hidden dimension, and number of layers.
Launching with Custom Concurrency
The start-tp3.sh launcher accepts --max-num-seqs as a passthrough argument to the underlying vLLM engine (lines 14-20):
# Default launch with 6 concurrent sequences
./start-tp3.sh
# Moderate concurrency increase to 12
./start-tp3.sh --max-num-seqs 12
# Maximum documented configuration (16 slots)
./start-tp3.sh --max-num-seqs 16
The launcher script forwards this flag directly to vLLM's AsyncLLMEngine, which initializes the scheduler with the specified capacity.
Environment Variable Configuration
You can also set MAX_NUM_SEQS via environment variable before launch:
import os
import subprocess
# Configure 12 concurrent sequences
os.environ["MAX_NUM_SEQS"] = "12"
# Launch the inference server
subprocess.run(["./start-tp3.sh", "--tensor-parallel-size", "3"])
This approach integrates cleanly with containerized deployments and orchestration systems where command-line flags are less convenient.
Memory Considerations for High Concurrency
Raising MAX_NUM_SEQS linearly increases KV cache allocation. For DeepSeek-v4's architecture, estimate required memory per sequence:
- Per-token KV cache:
2 × num_layers × hidden_size × num_kv_heads × precision_bytes - Total cache:
max_seq_length × max_num_seqs × per_token_size
On 2×DGX configurations with H100 GPUs, the documented maximum of 16 sequences assumes typical sequence lengths (4K-8K tokens). For longer contexts or batching scenarios, reduce concurrency proportionally.
Code Reference: Configuration Propagation
The configuration flows through three key files as implemented in MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark:
validate-dspark-config.sh— Validates and defaults the settingstart-tp3.sh— Accepts and forwards the CLI flag- vLLM engine initialization — Receives via
EngineArgs
No source-level enforcement caps the value; vLLM's scheduler will attempt to allocate requested slots until GPU memory exhaustion triggers runtime errors.
Summary
- Default concurrency: 6 sequences, defined in
validate-dspark-config.shfor safe baseline operation - Demonstrated maximum: 16 sequences, shown in official documentation examples
- Configuration methods:
--max-num-seqsCLI flag orMAX_NUM_SEQSenvironment variable - Practical limit: GPU memory capacity for KV cache, not a software-enforced cap
- Tuning recommendation: Start at 6, profile memory usage, increase incrementally to maximize throughput without OOM errors
Frequently Asked Questions
What happens if I set MAX_NUM_SEQS too high?
The vLLM engine will fail to initialize or crash with a CUDA out-of-memory error during the first forward pass. No automatic fallback reduces concurrency; you must restart with a lower value. Profile memory with nvidia-smi or PyTorch memory stats to find your hardware's practical limit.
Does higher MAX_NUM_SEQS always improve throughput?
Up to a point. Additional concurrent sequences improve GPU utilization through better memory bandwidth hiding, but diminishing returns occur when overhead from context switching and memory fragmentation outweighs gains. Benchmark with your specific prompt length distribution to identify the optimal setting.
Can different GPU configurations support different maximums?
Yes. The 2×DGX reference hardware (16×H100 80GB) supports 16 sequences at typical lengths. Smaller deployments should scale proportionally: single-node A100 systems often stabilize around 8-12 sequences for full-context workloads.
Is MAX_NUM_SEQS the same as batch size?
Conceptually related but operationally distinct. Batch size refers to tokens processed together in a single forward pass, while MAX_NUM_SEQS bounds how many independent sequences remain active in the scheduler. Continuous batching in vLLM dynamically packs tokens across these slots, so effective throughput depends on both parameters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →