# What Is Speculative Decoding in MTPLX? A Deep Dive Into the Draft-Verify Architecture

> Learn about speculative decoding in MTPLX. This technique accelerates token generation using a draft-verify architecture for faster language model inference.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-08

---

**Speculative decoding in MTPLX is a latency-optimization technique that speeds up token generation by running a fast draft model in parallel with the target (full-size) model, verifying draft tokens in a single batched forward pass, and committing them only when they match the target distribution.**

Speculative decoding enables significant speedups in large language model inference without sacrificing output quality. In MTPLX, this mechanism is implemented through a dedicated draft-verify-rollback loop that integrates directly with the generation engine. The implementation centers on the DeepSeek-V4 architecture and exposes both high-level CLI flags and low-level programmatic APIs for research use.

## How Speculative Decoding Works in MTPLX

MTPLX implements speculative decoding through three coordinated phases: draft generation, target verification, and conditional commit or rollback.

### The Draft-Verify-Rollback Loop

The core architecture resides in [`mtplx/models/deepseek_v4.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/models/deepseek_v4.py). The `Model` class exposes a draft surface through methods like `__call__(return_hidden=…)`, `mtp_forward`, `mtp_update_cache`, and `make_mtp_cache`. The `inject_deepseek_v4_mtp_support` function wires this surface into the generation engine, which drives the complete loop:

- **Draft**: The fast draft model emits `K` tokens speculatively (where `K` is the **speculative depth**)
- **Verify**: The target model validates all `K` tokens in one batched forward pass
- **Accept/Reject**: Matching tokens are committed; mismatches trigger rollback
- **Rollback**: The cache trims rejected draft rows via `DeepseekV4Cache.trim`

This mechanism is sometimes called **greedy speculative decode** because the draft model greedily emits tokens that are later either accepted or rejected.

### Speculative Depth Configuration

The **speculative depth** parameter controls how many draft tokens are generated before verification. According to [`tests/test_tail_gemma4_stream_holdback.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_tail_gemma4_stream_holdback.py) (lines 245–268), users can set depths such as 2 or 3:

```python
from mtplx import Model, SamplerConfig

model = Model.from_pretrained("deepseek_v4")

# Request speculative depth of 2

sampler_cfg = SamplerConfig(speculative_depth=2)

output = model.generate("Explain speculative decoding.", sampler_cfg=sampler_cfg)

```

Higher depths increase potential speedups but also raise the probability of rejection, creating a trade-off explored in benchmarking scripts.

## Exactness Guarantee and Correctness

A critical property of MTPLX's speculative decoding is its **bit-exact guarantee**. As validated in [`tests/test_deepseek_v4_spec.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_deepseek_v4_spec.py) (lines 111–116), greedy speculative decode produces output identical to plain autoregressive decoding when draft tokens are accepted. This holds for every depth because verification uses the same target forward that would have occurred in standard decoding.

The rollback mechanism ensures correctness when verification fails. In [`mtplx/models/deepseek_v4.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/models/deepseek_v4.py) (lines 82–89), `DeepseekV4Cache.trim` removes rejected draft rows, allowing the speculative lane to rewind safely and resume from the last valid state:

```python

# Direct low-level control for research

draft_cache = model.make_mtp_cache()
draft_logits = model.mtp_forward(draft_cache, input_ids, depth=3)
verified = model.mtp_update_cache(draft_cache, input_ids, draft_logits)

if verified.accepted:
    model.commit(draft_cache)
else:
    draft_cache.trim()  # Roll back rejected rows

```

## Performance Characteristics and CLI Integration

### Latency Reduction

By emitting `K` draft tokens at once, per-token latency drops from approximately 1× AR time to roughly `1/(K+1)` of AR time while preserving final output quality. This relationship is documented in [`scripts/deepseek_v4_mtpk_bench.py`](https://github.com/youssofal/MTPLX/blob/main/scripts/deepseek_v4_mtpk_bench.py) commentary around line 293.

### Command-Line Interface

MTPLX exposes speculative decoding through the `--speculative-depth` flag in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py). The `auto` mode (default) lets the engine select optimal depth for the loaded model:

```bash

# Enable speculative decoding with depth 3

$ mtplx serve --model deepseek_v4 --speculative-depth 3

# Use auto-selected depth

$ mtplx serve --model deepseek_v4 --speculative-depth auto

```

### Concurrent Speculative Decoding

Version 2.6.0 introduced **concurrent speculative decoding** for batched requests. The scheduler mode `--scheduler-mode mtp_batch` gives each request its own independent draft state, as described in [`docs/releases/v2.6.0.md`](https://github.com/youssofal/MTPLX/blob/main/docs/releases/v2.6.0.md) (lines 3–15). This enables throughput scaling without cross-request interference.

## Key Implementation Files

| File | Role |
|------|------|
| [`mtplx/models/deepseek_v4.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/models/deepseek_v4.py) | Core implementation: speculative lane, cache rollback, draft-model wiring |
| [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) | Parses `--speculative-depth` and forwards to generation engine |
| [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py) | Registers models as supporting speculative decode |
| [`tests/test_deepseek_v4_spec.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_deepseek_v4_spec.py) | Validates exactness against pure AR decoding |
| [`docs/releases/v2.6.0.md`](https://github.com/youssofal/MTPLX/blob/main/docs/releases/v2.6.0.md) | Documents concurrent speculative decoding |

## Summary

- **Speculative decoding in MTPLX** combines a fast draft model with target verification to reduce per-token latency
- The draft-verify-rollback loop in [`mtplx/models/deepseek_v4.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/models/deepseek_v4.py) enables greedy speculative decode with bit-exact guarantees
- **Speculative depth** (`--speculative-depth`) controls the trade-off between speedup and rejection probability
- **Rollback mechanics** via `DeepseekV4Cache.trim` ensure correctness when draft tokens fail verification
- **Concurrent mode** supports independent draft states across batched requests

## Frequently Asked Questions

### Is speculative decoding in MTPLX lossless?

Yes. Greedy speculative decoding in MTPLX is **bit-identical** to standard autoregressive decoding when draft tokens are accepted. The verification step uses the same target model forward that would occur in non-speculative generation, and the rejection-repair path ensures correct outputs on mismatch.

### How do I choose the right speculative depth?

Start with `auto` mode, which lets MTPLX select depth based on model architecture. For manual tuning, depths of 2–4 typically balance speedup and acceptance rate. Higher depths increase potential parallelism but raise rejection probability, especially for complex prompts. Benchmark with [`scripts/deepseek_v4_mtpk_bench.py`](https://github.com/youssofal/MTPLX/blob/main/scripts/deepseek_v4_mtpk_bench.py) to find optimal values for your workload.

### Can multiple requests use speculative decoding simultaneously?

Yes. The `mtp_batch` scheduler mode (`--scheduler-mode mtp_batch`) provides **concurrent speculative decoding** where each request maintains independent draft state. This prevents cross-request cache pollution and enables speculative decoding at batch scale.

### What models support speculative decoding in MTPLX?

Currently, DeepSeek-V4 variants implement full speculative decoding support through the `inject_deepseek_v4_mtp_support` function. Model support is registered in [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py), which controls artifact building and capability exposure. Check the registry or documentation for the latest supported architectures.