# DSpark Speculative Decoding Parameters in DeepSeek-v4-Flash: Complete Configuration Guide

> Master DSpark speculative decoding with DeepSeek-v4-Flash. Explore MTP_NUM_TOKENS, DRAFT_SAMPLE_METHOD, and optional hot-fix flags for optimal configuration.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: deep-dive
- Published: 2026-09-09

---

**DSpark speculative decoding in the MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository is controlled by `MTP_NUM_TOKENS` and `DRAFT_SAMPLE_METHOD`, with optional hot-fix flags `DSPARK_ENABLE_DSPARK_BLOCK_K` and `DSPARK_ENABLE_SWA_PREFIX` that modify validation rules and attention behavior.**

The DeepSeek-v4-Flash-DSpark 2× DGX Spark setup implements a custom speculative decoding pipeline built on vLLM with DSpark-specific optimizations. Understanding these parameters is essential for tuning throughput and acceptance rates in production deployments. This guide breaks down every configurable knob, where each is defined in the source code, and how they flow from environment variables to the vLLM runtime.

## Core DSpark Speculative Decoding Parameters

### `MTP_NUM_TOKENS`: Speculative Token Depth

The **`MTP_NUM_TOKENS`** environment variable sets how many draft tokens DSpark generates per forward pass. This is the primary lever for trading off latency against acceptance probability.

Default behavior varies by model variant:
- **Vision-Exp default**: `6`
- **Validation rule**: Must be ≥ 5 and divisible by 3 unless hot-fix is applied

The parameter also directly impacts CUDA memory through graph capture sizing:

```python

# CUDA graph capture size calculation (conceptual)

graph_capture_size = MAX_NUM_SEQS * (MTP_NUM_TOKENS + 1)

```

Set this in `.env.dspark.example` at line 424:

```bash

# .env.dspark.example

MTP_NUM_TOKENS=6

```

### `DRAFT_SAMPLE_METHOD`: Draft Model Sampling Strategy

The **`DRAFT_SAMPLE_METHOD`** controls how the draft model selects tokens during speculation. Two strategies are supported:

| Value | Behavior | Use Case |
|-------|----------|----------|
| `probabilistic` | Sample from draft model distribution | Higher diversity, better for creative generation |
| `greedy` | Always select highest-probability token | Lower variance, faster verification |

This parameter is defined at line 423 in `.env.dspark.example` and injected into the `SPECULATIVE_CONFIG` JSON passed to vLLM.

### `SPECULATIVE_CONFIG`: Runtime JSON Construction

The actual configuration reaching vLLM is a **generated JSON blob** built by the launch infrastructure. It is not set directly by users but assembled from the parameters above.

In [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) at line 479, the composition occurs:

```yaml

# docker-compose.dspark.yml (line 479)

SPECULATIVE_CONFIG="{\"method\":\"dspark\",\"num_speculative_tokens\":$${MTP_NUM_TOKENS:-6},\"draft_sample_method\":\"$${DRAFT_SAMPLE_METHOD}\"}"

```

The [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) launch script at line 1281 echoes the effective configuration for verification:

```bash

# start-deepseek-v4-flash-dspark.sh (line 1281)

echo "SPECULATIVE_CONFIG: ${SPECULATIVE_CONFIG}"

```

## Optional Hot-Fix Flags for Advanced Tuning

### `DSPARK_ENABLE_DSPARK_BLOCK_K`: Disable Block-K Constraints

Setting **`DSPARK_ENABLE_DSPARK_BLOCK_K=1`** removes the Vision-Exp restriction that `MTP_NUM_TOKENS` must be ≥ 5 and divisible by 3. This enables:

- Any `MTP_NUM_TOKENS >= 1`
- Original DSpark "stacked-stage" speculative behavior
- Finer-grained latency/throughput tradeoffs

Documented in [`docs/ENVS.md`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docs/ENVS.md) at line 110:

```bash

# docs/ENVS.md (line 110)

DSPARK_ENABLE_DSPARK_BLOCK_K=1  # Disables block-k validation, allows any MTP_NUM_TOKENS >= 1

```

### `DSPARK_ENABLE_SWA_PREFIX`: Sliding Window Attention Fix

While not directly controlling speculative depth, **`DSPARK_ENABLE_SWA_PREFIX=1`** is required when running Vision-Exp with default configurations. This flag enables the DSpark sliding-window-attention prefix fix, preventing attention computation errors with long context prefixes.

Defined in [`docs/ENVS.md`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docs/ENVS.md) at line 108.

## Parameter Flow: From Environment to vLLM Runtime

Understanding how DSpark speculative decoding parameters propagate helps debug configuration issues:

1. **Environment definition** – Set `MTP_NUM_TOKENS`, `DRAFT_SAMPLE_METHOD`, and optional flags in `.env.dspark`

2. **Validation & logging** – [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) performs runtime checks:
   - Vision-Exp enforces `MTP_NUM_TOKENS >= 5` and divisibility by 3 (unless `DSPARK_ENABLE_DSPARK_BLOCK_K=1`)
   - Prints effective `SPECULATIVE_CONFIG` to stdout at line 1281

3. **Container injection** – [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) constructs the JSON and binds it into container environment

4. **vLLM parsing** – The `SpeculativeConfig` class in [`recipe/overlay/vllm/config/speculative.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/recipe/overlay/vllm/config/speculative.py) deserializes the JSON, instantiates the DSpark proposer, and drives the speculative decoding loop with the configured token count and sampling method

## Configuration Examples

### Standard Vision-Exp Deployment

```bash

# .env.dspark

MTP_NUM_TOKENS=6
DRAFT_SAMPLE_METHOD=probabilistic

# No hot-fix flags needed for default behavior

```

### Low-Latency with Block-K Hot-Fix

```bash

# .env.dspark

MTP_NUM_TOKENS=3              # Below standard minimum

DRAFT_SAMPLE_METHOD=greedy    # Deterministic draft sampling

DSPARK_ENABLE_DSPARK_BLOCK_K=1  # Required to allow MTP_NUM_TOKENS=3

```

### Maximum Throughput Configuration

```bash

# .env.dspark

MTP_NUM_TOKENS=9              # Higher speculative depth

DRAFT_SAMPLE_METHOD=greedy    # Faster draft generation

DSPARK_ENABLE_DSPARK_BLOCK_K=1  # Ensure divisibility constraint doesn't block

DSPARK_ENABLE_SWA_PREFIX=1    # Required for Vision-Exp stability

```

## Summary

- **`MTP_NUM_TOKENS`** controls speculative token depth in [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) and `.env.dspark.example`, defaulting to 6 for Vision-Exp with constraints of ≥ 5 and divisible by 3
- **`DRAFT_SAMPLE_METHOD`** selects between `probabilistic` and `greedy` draft sampling, defined at line 423 in `.env.dspark.example`
- **`SPECULATIVE_CONFIG`** is auto-generated JSON combining both parameters, constructed in [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml) and logged by [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh)
- **`DSPARK_ENABLE_DSPARK_BLOCK_K=1`** removes token count constraints, enabling any `MTP_NUM_TOKENS >= 1`
- **`DSPARK_ENABLE_SWA_PREFIX=1`** stabilizes Vision-Exp attention but does not affect speculative decoding depth
- vLLM receives configuration through [`recipe/overlay/vllm/config/speculative.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/recipe/overlay/vllm/config/speculative.py) which parses the JSON and instantiates the DSpark proposer

## Frequently Asked Questions

### How do I change the number of speculative tokens in DSpark?

Set the **`MTP_NUM_TOKENS`** environment variable in your `.env.dspark` file. For Vision-Exp models, the default is 6, and values must be at least 5 and divisible by 3 unless you enable `DSPARK_ENABLE_DSPARK_BLOCK_K=1`. The change takes effect when you restart the deployment via [`docker-compose.dspark.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark.yml).

### What is the difference between probabilistic and greedy draft sampling in DSpark?

**`probabilistic`** sampling draws draft tokens from the model's output distribution, matching typical autoregressive generation and improving acceptance rates for diverse outputs. **`greedy`** always selects the highest-probability token, reducing draft variance and potentially speeding up verification when the target model agrees with the draft model's confident predictions. Set via `DRAFT_SAMPLE_METHOD` in `.env.dspark.example`.

### Why does my Vision-Exp deployment fail with MTP_NUM_TOKENS validation errors?

The [`start-deepseek-v4-flash-dspark.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/start-deepseek-v4-flash-dspark.sh) script enforces Vision-Exp constraints requiring `MTP_NUM_TOKENS >= 5` and divisibility by 3. To use smaller values (such as 1, 2, 3, or 4), set **`DSPARK_ENABLE_DSPARK_BLOCK_K=1`** in your environment. This hot-fix restores original DSpark "stacked-stage" behavior but may require validation for your specific workload.

### Where is the DSpark speculative configuration actually consumed by vLLM?

The generated **`SPECULATIVE_CONFIG`** JSON is parsed by the `SpeculativeConfig` class in [`recipe/overlay/vllm/config/speculative.py`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/recipe/overlay/vllm/config/speculative.py). This overlay modifies standard vLLM to instantiate a DSpark-specific proposer using your `num_speculative_tokens` and `draft_sample_method` values, driving the speculative decoding loop during inference.