# DSpark Speculative Decoding: How to Enable DeepSeek V4 Flash Draft-Verify Acceleration

> Explore DSpark speculative decoding for DeepSeek V4 Flash. Accelerate token generation by using a lightweight draft model for faster verification and increased throughput. Learn how to enable this experimental feature.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-07

---

**DSpark speculative decoding is an experimental auxiliary draft model that generates up to five future tokens using a lightweight ~5.6 GiB checkpoint, which the main DeepSeek V4 Flash model then verifies to reduce full verification passes and increase generation throughput.**

The antirez/ds4 repository implements DSpark speculative decoding to accelerate inference for the DeepSeek V4 Flash model. By employing a separate draft model to predict token sequences before verification, this mechanism aims to minimize expensive forward passes through the main model during predictable text continuations such as code generation.

## What is DSpark Speculative Decoding?

DSpark speculative decoding utilizes an **auxiliary draft model** released by DeepSeek that operates alongside the main DeepSeek V4 Flash model. The system works through a draft-verify loop where the draft model proposes candidate tokens and the main model verifies them in parallel.

The **DSpark support GGUF** stores a lightweight draft checkpoint (approximately 5.6 GiB) capable of proposing **up to five future tokens** by reusing hidden states from the main model. The main Flash model then verifies these proposals, committing only the accepted prefix. If a draft suffix is rejected or falls below a confidence threshold, the runtime **falls back to ordinary target decoding**.

This approach targets **reduction in full verification passes** to increase throughput, particularly for predictable continuations. However, because the draft work incurs overhead, performance gains are prompt-dependent; some inputs may show no speed improvement or may even degrade performance. DSpark speculative decoding is **experimental** and requires explicit activation.

## How DSpark Speculative Decoding Works Under the Hood

### The Draft-Verify Architecture

The architecture relies on several key components working in concert:

- **DSpark support GGUF**: Contains the draft model weights and metadata including block size and Markov rank
- **`--mtp` flag**: Supplies the support GGUF to the engine using the existing Multi-Token Prediction path mechanism
- **`--dspark` flag**: Activates the speculative decoder, triggering the draft-verify loop in the runtime
- **Confidence threshold** (`--dspark-confidence`): Prunes low-probability draft blocks with a default value of 0.7; setting to 0 forces fixed-size five-token blocks for diagnostics
- **Strict mode** (`--dspark-strict`): Loads the DSpark support model but maintains target-only decoding, useful for correctness testing and baseline comparison

### Configuration Flag Implementation

The command-line parser in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) handles DSpark activation through explicit flag detection. At lines 1851-1860, the parser sets the relevant engine configuration fields:

```c
// ds4_cli.c – handling of DSpark flags
} else if (!strcmp(arg, "--dspark")) {
    c.engine.dspark = true;
} else if (!strcmp(arg, "--dspark-confidence")) {
    c.engine.dspark = true;
    c.engine.dspark_confidence_threshold =
        parse_float_range(need_arg(&i, argc, argv, arg), arg, 0.0f, 1.0f);
    c.engine.dspark_confidence_threshold_set = true;
} else if (!strcmp(arg, "--dspark-strict")) {
    c.engine.dspark = true;
    c.engine.dspark_strict = true;
}

```

The same flag handling appears in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) (lines 12851-12858) for the server binary, ensuring consistent behavior across deployment modes.

## How to Enable DSpark Speculative Decoding

### 1. Download the DSpark Support Model

Download the required support checkpoint (approximately 5.6 GiB) using the provided script:

```bash
./download_model.sh ds4f-dspark

```

This command retrieves `DeepSeek-V4-Flash-DSpark-support-0731.gguf` into the `gguf/` directory.

### 2. Run with DSpark Activation

Execute the engine with both the main model and the DSpark support GGUF:

```bash
./ds4 -m ds4flash.gguf \
    --mtp gguf/DeepSeek-V4-Flash-DSpark-support-0731.gguf \
    --dspark --temp 0

```

**Required parameters:**
- `--mtp` supplies the support checkpoint path
- `--dspark` activates the speculative decoder
- `--temp 0` enforces greedy decoding (mandatory for DSpark operation)

### 3. Tune Acceptance Behavior

Adjust the speculative decoding behavior using optional flags:

- **Confidence pruning** (default 0.7): Lower values accept more aggressive drafts
  ```bash
  ./ds4 ... --dspark-confidence 0.5
  ```

  
- **Fixed-size blocks**: Set confidence to 0 for diagnostic five-token blocks
  ```bash
  ./ds4 ... --dspark-confidence 0
  ```

  
- **Strict mode**: Load support model without speculative execution for testing
  ```bash
  ./ds4 ... --dspark-strict
  ```

## Practical Command Examples

```bash

# Basic DSpark run with greedy decoding

./ds4 -m ds4flash.gguf \
    --mtp gguf/DeepSeek-V4-Flash-DSpark-support-0731.gguf \
    --dspark --temp 0

```

```bash

# Aggressive acceptance threshold at 0.6 confidence

./ds4 -m ds4flash.gguf \
    --mtp gguf/DeepSeek-V4-Flash-DSpark-support-0731.gguf \
    --dspark --dspark-confidence 0.6 --temp 0

```

```bash

# Strict mode for baseline performance comparison

./ds4 -m ds4flash.gguf \
    --mtp gguf/DeepSeek-V4-Flash-DSpark-support-0731.gguf \
    --dspark-strict --temp 0

```

These examples assume the main Flash GGUF (`ds4flash.gguf`) is already present, downloadable via `./download_model.sh ds4f-q2`, `ds4f-q2-q4`, or `ds4f-q4`.

## Summary

- **DSpark speculative decoding** employs a ~5.6 GiB auxiliary draft model to predict up to five tokens ahead of the main DeepSeek V4 Flash model, reducing verification passes.
- **Enable the feature** by downloading the support GGUF and invoking `ds4` with `--mtp`, `--dspark`, and `--temp 0` flags.
- **Tune performance** using `--dspark-confidence` (default 0.7) to control acceptance thresholds, or set to 0 for fixed-size diagnostic blocks.
- **Test correctness** using `--dspark-strict` mode, which loads the support model while maintaining target-only decoding for baseline comparison.
- **Performance varies** by prompt predictability; the feature is experimental and most effective on structured content like code.

## Frequently Asked Questions

### What hardware requirements does DSpark speculative decoding have?

DSpark speculative decoding requires sufficient VRAM to load both the main DeepSeek V4 Flash model and the ~5.6 GiB DSpark support GGUF simultaneously. The draft model adds computational overhead during the verify phase, so optimal performance requires GPUs with spare compute capacity to handle the draft-verify loop efficiently.

### Why does DSpark require greedy decoding (--temp 0)?

The draft-verify mechanism relies on deterministic token prediction to maintain consistency between the draft model's proposals and the main model's verification path. Random sampling would desynchronize the draft and target distributions, causing frequent rejections of draft tokens and eliminating the throughput benefits of speculative execution.

### How do I know if DSpark is actually improving my generation speed?

Monitor tokens-per-second metrics during generation and compare against `--dspark-strict` mode, which loads the support model but disables speculative decoding. According to the source implementation, DSpark reduces full verification passes for predictable continuations like code, but unpredictable prompts may show no improvement or minor slowdown due to draft computation overhead.

### Can I use DSpark with quantization formats other than GGUF?

The current antirez/ds4 implementation specifically requires the DSpark support model in GGUF format, as evidenced by the `download_model.sh ds4f-dspark` command and the `--mtp` flag integration. The quantization tool [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) (lines 2642-2650) handles DSpark support GGUF generation through `--dspark-manifest` and `--dspark-support` options, indicating tight format coupling.