DSpark Speculative Decoding Parameters in DeepSeek-v4-Flash: Complete Configuration Guide

DSpark speculative decoding in the MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository is controlled by MTP_NUM_TOKENS and DRAFT_SAMPLE_METHOD, with optional hot-fix flags DSPARK_ENABLE_DSPARK_BLOCK_K and DSPARK_ENABLE_SWA_PREFIX that modify validation rules and attention behavior.

The DeepSeek-v4-Flash-DSpark 2× DGX Spark setup implements a custom speculative decoding pipeline built on vLLM with DSpark-specific optimizations. Understanding these parameters is essential for tuning throughput and acceptance rates in production deployments. This guide breaks down every configurable knob, where each is defined in the source code, and how they flow from environment variables to the vLLM runtime.

Core DSpark Speculative Decoding Parameters

MTP_NUM_TOKENS: Speculative Token Depth

The MTP_NUM_TOKENS environment variable sets how many draft tokens DSpark generates per forward pass. This is the primary lever for trading off latency against acceptance probability.

Default behavior varies by model variant:

  • Vision-Exp default: 6
  • Validation rule: Must be ≥ 5 and divisible by 3 unless hot-fix is applied

The parameter also directly impacts CUDA memory through graph capture sizing:


# CUDA graph capture size calculation (conceptual)

graph_capture_size = MAX_NUM_SEQS * (MTP_NUM_TOKENS + 1)

Set this in .env.dspark.example at line 424:


# .env.dspark.example

MTP_NUM_TOKENS=6

DRAFT_SAMPLE_METHOD: Draft Model Sampling Strategy

The DRAFT_SAMPLE_METHOD controls how the draft model selects tokens during speculation. Two strategies are supported:

Value Behavior Use Case
probabilistic Sample from draft model distribution Higher diversity, better for creative generation
greedy Always select highest-probability token Lower variance, faster verification

This parameter is defined at line 423 in .env.dspark.example and injected into the SPECULATIVE_CONFIG JSON passed to vLLM.

SPECULATIVE_CONFIG: Runtime JSON Construction

The actual configuration reaching vLLM is a generated JSON blob built by the launch infrastructure. It is not set directly by users but assembled from the parameters above.

In docker-compose.dspark.yml at line 479, the composition occurs:


# docker-compose.dspark.yml (line 479)

SPECULATIVE_CONFIG="{\"method\":\"dspark\",\"num_speculative_tokens\":$${MTP_NUM_TOKENS:-6},\"draft_sample_method\":\"$${DRAFT_SAMPLE_METHOD}\"}"

The start-deepseek-v4-flash-dspark.sh launch script at line 1281 echoes the effective configuration for verification:


# start-deepseek-v4-flash-dspark.sh (line 1281)

echo "SPECULATIVE_CONFIG: ${SPECULATIVE_CONFIG}"

Optional Hot-Fix Flags for Advanced Tuning

DSPARK_ENABLE_DSPARK_BLOCK_K: Disable Block-K Constraints

Setting DSPARK_ENABLE_DSPARK_BLOCK_K=1 removes the Vision-Exp restriction that MTP_NUM_TOKENS must be ≥ 5 and divisible by 3. This enables:

  • Any MTP_NUM_TOKENS >= 1
  • Original DSpark "stacked-stage" speculative behavior
  • Finer-grained latency/throughput tradeoffs

Documented in docs/ENVS.md at line 110:


# docs/ENVS.md (line 110)

DSPARK_ENABLE_DSPARK_BLOCK_K=1  # Disables block-k validation, allows any MTP_NUM_TOKENS >= 1

DSPARK_ENABLE_SWA_PREFIX: Sliding Window Attention Fix

While not directly controlling speculative depth, DSPARK_ENABLE_SWA_PREFIX=1 is required when running Vision-Exp with default configurations. This flag enables the DSpark sliding-window-attention prefix fix, preventing attention computation errors with long context prefixes.

Defined in docs/ENVS.md at line 108.

Parameter Flow: From Environment to vLLM Runtime

Understanding how DSpark speculative decoding parameters propagate helps debug configuration issues:

  1. Environment definition – Set MTP_NUM_TOKENS, DRAFT_SAMPLE_METHOD, and optional flags in .env.dspark

  2. Validation & logging – start-deepseek-v4-flash-dspark.sh performs runtime checks:

    • Vision-Exp enforces MTP_NUM_TOKENS >= 5 and divisibility by 3 (unless DSPARK_ENABLE_DSPARK_BLOCK_K=1)
    • Prints effective SPECULATIVE_CONFIG to stdout at line 1281
  3. Container injection – docker-compose.dspark.yml constructs the JSON and binds it into container environment

  4. vLLM parsing – The SpeculativeConfig class in recipe/overlay/vllm/config/speculative.py deserializes the JSON, instantiates the DSpark proposer, and drives the speculative decoding loop with the configured token count and sampling method

Configuration Examples

Standard Vision-Exp Deployment


# .env.dspark

MTP_NUM_TOKENS=6
DRAFT_SAMPLE_METHOD=probabilistic

# No hot-fix flags needed for default behavior

Low-Latency with Block-K Hot-Fix


# .env.dspark

MTP_NUM_TOKENS=3              # Below standard minimum

DRAFT_SAMPLE_METHOD=greedy    # Deterministic draft sampling

DSPARK_ENABLE_DSPARK_BLOCK_K=1  # Required to allow MTP_NUM_TOKENS=3

Maximum Throughput Configuration


# .env.dspark

MTP_NUM_TOKENS=9              # Higher speculative depth

DRAFT_SAMPLE_METHOD=greedy    # Faster draft generation

DSPARK_ENABLE_DSPARK_BLOCK_K=1  # Ensure divisibility constraint doesn't block

DSPARK_ENABLE_SWA_PREFIX=1    # Required for Vision-Exp stability

Summary

  • MTP_NUM_TOKENS controls speculative token depth in docker-compose.dspark.yml and .env.dspark.example, defaulting to 6 for Vision-Exp with constraints of ≥ 5 and divisible by 3
  • DRAFT_SAMPLE_METHOD selects between probabilistic and greedy draft sampling, defined at line 423 in .env.dspark.example
  • SPECULATIVE_CONFIG is auto-generated JSON combining both parameters, constructed in docker-compose.dspark.yml and logged by start-deepseek-v4-flash-dspark.sh
  • DSPARK_ENABLE_DSPARK_BLOCK_K=1 removes token count constraints, enabling any MTP_NUM_TOKENS >= 1
  • DSPARK_ENABLE_SWA_PREFIX=1 stabilizes Vision-Exp attention but does not affect speculative decoding depth
  • vLLM receives configuration through recipe/overlay/vllm/config/speculative.py which parses the JSON and instantiates the DSpark proposer

Frequently Asked Questions

How do I change the number of speculative tokens in DSpark?

Set the MTP_NUM_TOKENS environment variable in your .env.dspark file. For Vision-Exp models, the default is 6, and values must be at least 5 and divisible by 3 unless you enable DSPARK_ENABLE_DSPARK_BLOCK_K=1. The change takes effect when you restart the deployment via docker-compose.dspark.yml.

What is the difference between probabilistic and greedy draft sampling in DSpark?

probabilistic sampling draws draft tokens from the model's output distribution, matching typical autoregressive generation and improving acceptance rates for diverse outputs. greedy always selects the highest-probability token, reducing draft variance and potentially speeding up verification when the target model agrees with the draft model's confident predictions. Set via DRAFT_SAMPLE_METHOD in .env.dspark.example.

Why does my Vision-Exp deployment fail with MTP_NUM_TOKENS validation errors?

The start-deepseek-v4-flash-dspark.sh script enforces Vision-Exp constraints requiring MTP_NUM_TOKENS >= 5 and divisibility by 3. To use smaller values (such as 1, 2, 3, or 4), set DSPARK_ENABLE_DSPARK_BLOCK_K=1 in your environment. This hot-fix restores original DSpark "stacked-stage" behavior but may require validation for your specific workload.

Where is the DSpark speculative configuration actually consumed by vLLM?

The generated SPECULATIVE_CONFIG JSON is parsed by the SpeculativeConfig class in recipe/overlay/vllm/config/speculative.py. This overlay modifies standard vLLM to instantiate a DSpark-specific proposer using your num_speculative_tokens and draft_sample_method values, driving the speculative decoding loop during inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →