DSpark Speculative Decoding Parameters in DeepSeek-v4-Flash: Complete Configuration Guide
DSpark speculative decoding in the MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository is controlled by MTP_NUM_TOKENS and DRAFT_SAMPLE_METHOD, with optional hot-fix flags DSPARK_ENABLE_DSPARK_BLOCK_K and DSPARK_ENABLE_SWA_PREFIX that modify validation rules and attention behavior.
The DeepSeek-v4-Flash-DSpark 2× DGX Spark setup implements a custom speculative decoding pipeline built on vLLM with DSpark-specific optimizations. Understanding these parameters is essential for tuning throughput and acceptance rates in production deployments. This guide breaks down every configurable knob, where each is defined in the source code, and how they flow from environment variables to the vLLM runtime.
Core DSpark Speculative Decoding Parameters
MTP_NUM_TOKENS: Speculative Token Depth
The MTP_NUM_TOKENS environment variable sets how many draft tokens DSpark generates per forward pass. This is the primary lever for trading off latency against acceptance probability.
Default behavior varies by model variant:
- Vision-Exp default:
6 - Validation rule: Must be ≥ 5 and divisible by 3 unless hot-fix is applied
The parameter also directly impacts CUDA memory through graph capture sizing:
# CUDA graph capture size calculation (conceptual)
graph_capture_size = MAX_NUM_SEQS * (MTP_NUM_TOKENS + 1)
Set this in .env.dspark.example at line 424:
# .env.dspark.example
MTP_NUM_TOKENS=6
DRAFT_SAMPLE_METHOD: Draft Model Sampling Strategy
The DRAFT_SAMPLE_METHOD controls how the draft model selects tokens during speculation. Two strategies are supported:
| Value | Behavior | Use Case |
|---|---|---|
probabilistic |
Sample from draft model distribution | Higher diversity, better for creative generation |
greedy |
Always select highest-probability token | Lower variance, faster verification |
This parameter is defined at line 423 in .env.dspark.example and injected into the SPECULATIVE_CONFIG JSON passed to vLLM.
SPECULATIVE_CONFIG: Runtime JSON Construction
The actual configuration reaching vLLM is a generated JSON blob built by the launch infrastructure. It is not set directly by users but assembled from the parameters above.
In docker-compose.dspark.yml at line 479, the composition occurs:
# docker-compose.dspark.yml (line 479)
SPECULATIVE_CONFIG="{\"method\":\"dspark\",\"num_speculative_tokens\":$${MTP_NUM_TOKENS:-6},\"draft_sample_method\":\"$${DRAFT_SAMPLE_METHOD}\"}"
The start-deepseek-v4-flash-dspark.sh launch script at line 1281 echoes the effective configuration for verification:
# start-deepseek-v4-flash-dspark.sh (line 1281)
echo "SPECULATIVE_CONFIG: ${SPECULATIVE_CONFIG}"
Optional Hot-Fix Flags for Advanced Tuning
DSPARK_ENABLE_DSPARK_BLOCK_K: Disable Block-K Constraints
Setting DSPARK_ENABLE_DSPARK_BLOCK_K=1 removes the Vision-Exp restriction that MTP_NUM_TOKENS must be ≥ 5 and divisible by 3. This enables:
- Any
MTP_NUM_TOKENS >= 1 - Original DSpark "stacked-stage" speculative behavior
- Finer-grained latency/throughput tradeoffs
Documented in docs/ENVS.md at line 110:
# docs/ENVS.md (line 110)
DSPARK_ENABLE_DSPARK_BLOCK_K=1 # Disables block-k validation, allows any MTP_NUM_TOKENS >= 1
DSPARK_ENABLE_SWA_PREFIX: Sliding Window Attention Fix
While not directly controlling speculative depth, DSPARK_ENABLE_SWA_PREFIX=1 is required when running Vision-Exp with default configurations. This flag enables the DSpark sliding-window-attention prefix fix, preventing attention computation errors with long context prefixes.
Defined in docs/ENVS.md at line 108.
Parameter Flow: From Environment to vLLM Runtime
Understanding how DSpark speculative decoding parameters propagate helps debug configuration issues:
-
Environment definition – Set
MTP_NUM_TOKENS,DRAFT_SAMPLE_METHOD, and optional flags in.env.dspark -
Validation & logging –
start-deepseek-v4-flash-dspark.shperforms runtime checks:- Vision-Exp enforces
MTP_NUM_TOKENS >= 5and divisibility by 3 (unlessDSPARK_ENABLE_DSPARK_BLOCK_K=1) - Prints effective
SPECULATIVE_CONFIGto stdout at line 1281
- Vision-Exp enforces
-
Container injection –
docker-compose.dspark.ymlconstructs the JSON and binds it into container environment -
vLLM parsing – The
SpeculativeConfigclass inrecipe/overlay/vllm/config/speculative.pydeserializes the JSON, instantiates the DSpark proposer, and drives the speculative decoding loop with the configured token count and sampling method
Configuration Examples
Standard Vision-Exp Deployment
# .env.dspark
MTP_NUM_TOKENS=6
DRAFT_SAMPLE_METHOD=probabilistic
# No hot-fix flags needed for default behavior
Low-Latency with Block-K Hot-Fix
# .env.dspark
MTP_NUM_TOKENS=3 # Below standard minimum
DRAFT_SAMPLE_METHOD=greedy # Deterministic draft sampling
DSPARK_ENABLE_DSPARK_BLOCK_K=1 # Required to allow MTP_NUM_TOKENS=3
Maximum Throughput Configuration
# .env.dspark
MTP_NUM_TOKENS=9 # Higher speculative depth
DRAFT_SAMPLE_METHOD=greedy # Faster draft generation
DSPARK_ENABLE_DSPARK_BLOCK_K=1 # Ensure divisibility constraint doesn't block
DSPARK_ENABLE_SWA_PREFIX=1 # Required for Vision-Exp stability
Summary
MTP_NUM_TOKENScontrols speculative token depth indocker-compose.dspark.ymland.env.dspark.example, defaulting to 6 for Vision-Exp with constraints of ≥ 5 and divisible by 3DRAFT_SAMPLE_METHODselects betweenprobabilisticandgreedydraft sampling, defined at line 423 in.env.dspark.exampleSPECULATIVE_CONFIGis auto-generated JSON combining both parameters, constructed indocker-compose.dspark.ymland logged bystart-deepseek-v4-flash-dspark.shDSPARK_ENABLE_DSPARK_BLOCK_K=1removes token count constraints, enabling anyMTP_NUM_TOKENS >= 1DSPARK_ENABLE_SWA_PREFIX=1stabilizes Vision-Exp attention but does not affect speculative decoding depth- vLLM receives configuration through
recipe/overlay/vllm/config/speculative.pywhich parses the JSON and instantiates the DSpark proposer
Frequently Asked Questions
How do I change the number of speculative tokens in DSpark?
Set the MTP_NUM_TOKENS environment variable in your .env.dspark file. For Vision-Exp models, the default is 6, and values must be at least 5 and divisible by 3 unless you enable DSPARK_ENABLE_DSPARK_BLOCK_K=1. The change takes effect when you restart the deployment via docker-compose.dspark.yml.
What is the difference between probabilistic and greedy draft sampling in DSpark?
probabilistic sampling draws draft tokens from the model's output distribution, matching typical autoregressive generation and improving acceptance rates for diverse outputs. greedy always selects the highest-probability token, reducing draft variance and potentially speeding up verification when the target model agrees with the draft model's confident predictions. Set via DRAFT_SAMPLE_METHOD in .env.dspark.example.
Why does my Vision-Exp deployment fail with MTP_NUM_TOKENS validation errors?
The start-deepseek-v4-flash-dspark.sh script enforces Vision-Exp constraints requiring MTP_NUM_TOKENS >= 5 and divisibility by 3. To use smaller values (such as 1, 2, 3, or 4), set DSPARK_ENABLE_DSPARK_BLOCK_K=1 in your environment. This hot-fix restores original DSpark "stacked-stage" behavior but may require validation for your specific workload.
Where is the DSpark speculative configuration actually consumed by vLLM?
The generated SPECULATIVE_CONFIG JSON is parsed by the SpeculativeConfig class in recipe/overlay/vllm/config/speculative.py. This overlay modifies standard vLLM to instantiate a DSpark-specific proposer using your num_speculative_tokens and draft_sample_method values, driving the speculative decoding loop during inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →