# How ds4 Speculative Decoding Works: Implementation Guide and Setup

> Learn how ds4 speculative decoding accelerates GPU generation with its MTP draft model. This guide explains implementation and setup, ensuring identical results to autoregressive decoding.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**TLDR:** ds4 uses a lightweight MTP draft model to propose future tokens that are verified against the main model's greedy output, accepting matches until a mismatch triggers a KV-cache rollback, guaranteeing bit-for-bit identical results to autoregressive decoding while accelerating generation on GPU hardware.

ds4, the high-performance inference engine developed by antirez, implements **speculative decoding** (also called MTP draft decoding) to minimize per-token latency during transformer generation. This technique relies on a smaller draft model to predict token sequences, which the main model then verifies in parallel. The implementation guarantees that final outputs remain identical to a pure greedy autoregressive pass while significantly improving throughput on CUDA, Metal, and ROCm backends.

## Architecture of ds4 Speculative Decoding

The core mechanism resides in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) within the `ds4_session_eval_speculative_argmax` function, which orchestrates a five-stage draft-verify-commit pipeline.

### Draft Generation with the MTP Model

The process begins when the **MTP (Multi-Token Prediction) draft model** generates a short sequence of proposed tokens. As implemented at line 56530 of [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), this lightweight model executes a forward pass to produce a configurable number of draft tokens controlled by the `--mtp-draft N` parameter. The draft suffix represents the model's prediction of the next N tokens in the sequence.

### Token Verification and Commit

The main target model then evaluates the proposed suffix layer-by-layer to verify each token against its own greedy argmax selection. The verification logic at line 44115 of [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) compares the draft token against the target model's chosen token at each position. When tokens match, they are **committed** to the output stream and the KV-cache; upon the first mismatch, verification halts immediately.

### KV-Cache Management and Rollback

Draft tokens are stored in a **speculative KV ring buffer** separate from the main cache. According to the source at line 54117 of [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), when verification succeeds, these entries merge into the regular KV cache to maintain state consistency. If verification fails, the engine executes a rollback by discarding the speculative KV rows and resuming generation from the last committed token, ensuring the state remains equivalent to a non-speculative greedy run.

### Batched Verification for Efficiency

When draft depth exceeds two tokens, ds4 optimizes GPU utilization through batched verification. The code at line 45625 of [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) implements a "Batched output head for speculative verification" that processes multiple speculative tokens simultaneously, reducing kernel launch overhead and improving memory bandwidth efficiency on the target model.

## How to Enable Speculative Decoding in ds4

Enabling this feature requires the optional MTP draft model and GPU support, as the speculative code path is guarded by `#ifndef DS4_NO_GPU` and falls back to standard greedy decoding on CPU-only builds.

### 1. Download the MTP Draft Model

Fetch the lightweight draft model using the provided utility script referenced in [`gguf-tools/imatrix/dataset/rendered_prompts.txt`](https://github.com/antirez/ds4/blob/main/gguf-tools/imatrix/dataset/rendered_prompts.txt) at line 35823:

```bash
./download_model.sh mtp

```

This retrieves the `MTP.gguf` file containing the draft network weights.

### 2. Launch ds4 with the MTP Flag

Activate speculative decoding by passing both the main model and the draft model to the CLI, as parsed in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c):

```bash
ds4 --model DeepSeek-V4-Flash.gguf \
    --mtp MTP.gguf \
    --port 8080

```

The `--mtp` flag instructs the engine to load the draft model and initialize the speculative evaluation loop defined in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c).

### 3. Configure Draft Depth

Control the speculation window size using the `--mtp-draft` parameter. The default value is 1, but increasing this can improve throughput at the cost of higher verification overhead:

```bash
ds4 --model DeepSeek-V4-Flash.gguf \
    --mtp MTP.gguf \
    --mtp-draft 2

```

Higher values trigger the batched verification paths described in the "Batched output head" implementation.

### 4. Verify Correctness

Validate that speculative commits match autoregressive tokens using the test harness flag defined in [`tests/ds4_test.c`](https://github.com/antirez/ds4/blob/main/tests/ds4_test.c) at line 6594:

```bash
ds4 ... --mtp-verify-depth

```

This runs a parity check ensuring that speculative decoding produces identical results to greedy generation.

## Testing and Verification

The test suite in [`tests/ds4_test.c`](https://github.com/antirez/ds4/blob/main/tests/ds4_test.c) includes comprehensive validation of the speculative pipeline. The `--mtp-verify-depth` flag triggers assertions that compare speculative token commits against reference autoregressive outputs, confirming that the KV-cache rollback mechanisms and draft verification logic maintain mathematical equivalence to the non-accelerated path.

## Summary

- **ds4 speculative decoding** uses an MTP draft model to predict token sequences that the main model verifies in parallel.
- The `ds4_session_eval_speculative_argmax` function in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) handles the core draft-verify-commit cycle, with specific logic at lines 44115 and 54117 managing verification and rollback.
- Draft tokens reside in a speculative KV ring buffer that merges on commit or discards on mismatch, preserving greedy decoding equivalence.
- Enable the feature by running `./download_model.sh mtp` and launching with `--mtp MTP.gguf --mtp-draft N`.
- This capability requires GPU support and is disabled in `DS4_NO_GPU` builds.

## Frequently Asked Questions

### What is the MTP model in ds4?

The MTP (Multi-Token Prediction) model is a lightweight draft network shipped separately from the main weights. According to the source in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c), this model predicts future tokens to create a "draft suffix" that the main model verifies, allowing ds4 to accept multiple tokens per forward pass when the draft agrees with the target model's greedy selection.

### Does speculative decoding change the output quality?

No. The implementation guarantees **bit-for-bit identical** output compared to standard greedy decoding. As noted in the rollback logic at line 54117 of [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), any verification failure triggers an immediate KV-cache rollback to the last committed state, ensuring the final sequence matches exactly what autoregressive generation would produce.

### Can I use speculative decoding on CPU-only builds?

No. Speculative decoding requires GPU acceleration. The codebase wraps speculative logic with `#ifndef DS4_NO_GPU` guards, and CPU-only builds automatically fall back to single-step greedy decoding even if `--mtp` flags are provided.

### How do I tune the draft depth for optimal performance?

Use the `--mtp-draft N` flag where N is typically between 1 and 4. The optimal value depends on your draft model's accuracy and GPU memory bandwidth. Higher values engage the batched verification head (line 45625 in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)) but increase wasted computation if the draft model predictions diverge frequently from the target model.