# How MTP Speculative Decoding Works in the DeepSeek V4 Engine (antirez/ds4)

> Understand MTP speculative decoding in DeepSeek V4. Accelerate generation by drafting and verifying token sequences efficiently with a lightweight model and target model integration.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-08

---

**Multi-Token Prediction (MTP) speculative decoding accelerates generation by drafting short token sequences with a lightweight auxiliary model and verifying them in a single forward pass through the target model, committing drafts only when they match the target's argmax within a configurable margin.**

The `antirez/ds4` engine implements MTP speculative decoding as a legacy drafter path that proposes candidate token suffixes during greedy generation. This technique reduces per-token latency by accepting multiple tokens per target model evaluation when the lightweight MTP head accurately predicts the target model's outputs.

## Architecture and Source Locations

The implementation spans several components within the [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) core:

- **MTP Draft Generation** ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), lines 57973–60496): Executes the GLM MTP drafting loop to produce short suffixes using a separate raw cache.
- **Speculative Driver** ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), line 64773): The `ds4_session_eval_speculative_argmax()` function orchestrates drafting, verification, and result merging.
- **Target Verifier** ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), line 34432): Performs layer-major speculative verification over the entire draft suffix.
- **Acceptance Logic** ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), lines 64924–64944): Implements margin-based confidence checks using `DS4_MTP_MIN_MARGIN`.
- **CLI Interface** ([`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c)): Exposes `--mtp`, `--mtp-draft`, `--mtp-margin`, and `--glm-mtp` flags.

## Step-by-Step Execution Flow

### Draft Generation with the MTP Head

When the engine operates in greedy mode (`temperature ≤ 0`), the GLM MTP block ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), lines 42828–43100) receives the current token embedding and hidden state (`enorm`, `hnorm`). The MTP head—a small auxiliary network loaded from a separate GGUF file—runs a few transformer layers (the "nextn" block) to predict the next one or two tokens. These draft tokens populate a dedicated MTP cache, isolated from the main KV cache to prevent side effects (see comment at line 15129).

### Target Model Verification

After drafting, the engine invokes the layer-major speculative target verifier ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), line 34432). Unlike standard per-token decoding, this verifier runs the full target model once over the entire draft suffix, computing logits for each position and storing them in `s->mtp_logits`. This single forward pass evaluates all drafted positions simultaneously, minimizing overhead.

### Acceptance Criteria and Token Commit

The acceptance test compares the draft logits against the target logits at each position:

1. **Argmax Match**: The draft token must equal the target model's argmax at that position.
2. **Margin Check**: The target logit must exceed the draft logit by at least the configured margin (`DS4_MTP_MIN_MARGIN`, default 3.0).

If both conditions hold for a token, the engine commits it by copying the MTP cache state into the main KV cache via `ds4_session_eval_speculative_argmax()`. The process repeats from the new position; if verification fails, the engine falls back to standard per-token decoding.

## Configuration and Usage

### Command-Line Interface

Enable MTP speculative decoding via the CLI flags defined in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c):

```bash
./ds4 --model deepseek_v4.gguf \
      --mtp deepseek_mtp.gguf \
      --glm-mtp \
      --mtp-draft 2 \
      --mtp-margin 3.0

```

- `--mtp FILE`: Loads the auxiliary MTP head GGUF.
- `--glm-mtp`: Forces greedy-only mode (required for MTP).
- `--mtp-draft N`: Sets the maximum draft tokens per step (typically 1–2).
- `--mtp-margin F`: Configures the acceptance threshold (default 3.0).

### C API Integration

Programmatically configure MTP through the `ds4_cfg` structure:

```c
#include "ds4.h"

int main(void) {
    ds4_cfg cfg = ds4_default_cfg();
    cfg.model_path = "deepseek_v4.gguf";
    cfg.mtp_path = "deepseek_mtp.gguf";
    cfg.gen.temperature = 0.0f;  // Required for MTP
    cfg.mtp_draft = 2;
    cfg.mtp_margin = 3.0f;

    ds4_engine *engine = ds4_engine_new(&cfg);
    ds4_session *sess = ds4_engine_open_session(engine);
    
    // Automatically uses speculative decoding when conditions are met
    int token = ds4_session_eval_speculative_argmax(sess, 0, 100);
    
    ds4_engine_free(engine);
    return 0;
}

```

The `ds4_session_eval_speculative_argmax()` function automatically routes through the MTP path when `mtp_path` is set and temperature is zero.

## Key Implementation Details

**Greedy-Only Constraint**: MTP speculative decoding only activates when `cli_greedy_argmax_requested()` returns true ([`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c)). This restriction exists because the acceptance logic assumes deterministic argmax sampling from the target model.

**Separate Caching Strategy**: The engine maintains distinct raw caches for the MTP head and target model. This isolation prevents the drafter's intermediate tensors from contaminating the target model's KV cache during verification.

**Verification Depth**: Under tensor-parallel regimes ([`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c)), the engine falls back to per-token decoding rather than speculative batches, as noted in the TP implementation comments.

## Summary

- **MTP speculative decoding** uses a lightweight auxiliary model to draft 1–2 tokens ahead during greedy generation.
- **Verification** occurs in a single layer-major pass through the target model ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), line 34432).
- **Acceptance** requires matching argmax tokens with a configurable logit margin (`DS4_MTP_MIN_MARGIN`, default 3).
- **Configuration** happens via `--mtp` and related flags in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c), or through the `ds4_cfg` API structure.
- **Fallback** to standard decoding occurs when drafts fail verification or when tensor parallelism is active.

## Frequently Asked Questions

### What hardware requirements does MTP speculative decoding impose?

MTP requires loading an additional GGUF file into GPU memory, increasing VRAM usage by approximately 10–15% compared to the base model alone. The verifier runs the full target model, so compute requirements remain identical to standard decoding, though memory bandwidth utilization improves due to reduced total forward passes when drafts are accepted.

### Why does MTP only work with greedy decoding (temperature = 0)?

The acceptance logic in `ds4_session_eval_speculative_argmax()` assumes deterministic sampling where the target model's argmax represents the canonical next token. Stochastic sampling would require probability distribution matching rather than argmax comparison, complicating the margin-based verification implemented at lines 64924–64944.

### How should I tune the `--mtp-margin` parameter?

Start with the default value of 3.0. Decrease the margin (e.g., to 2.0) to increase acceptance rates at the risk of occasional quality degradation, or increase it (e.g., to 5.0) for stricter fidelity to the target model. Monitor acceptance rates via the `DS4_MTP_CONF_LOG=1` environment variable to empirically determine the optimal threshold for your specific model and dataset.

### Can MTP speculative decoding combine with other acceleration techniques?

MTP operates independently of quantization and tensor parallelism, though the source code in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) indicates that tensor-parallel mode disables speculative verification, falling back to per-token generation. Standard quantization schemes (Q4_K_M, Q8_0) apply normally to both the target model and MTP head GGUF files.