# MTP (Multi-Token Prediction) Support in ds4: Configuration and Usage Guide

> Learn about MTP Multi-Token Prediction in ds4. Discover how this speculative decoding technique speeds up greedy generation by drafting and verifying multiple tokens in parallel. Configure and use MTP effectively.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**Multi-Token Prediction (MTP) in ds4 is a speculative decoding mechanism that uses a lightweight support GGUF model to draft multiple tokens per step, allowing the main model to verify them in parallel for faster greedy generation.**

The `antirez/ds4` inference engine implements Multi-Token Prediction (MTP) as an optional acceleration path for greedy decoding scenarios. By loading a secondary "draft" model alongside your primary GGUF, ds4 can propose and validate short token sequences in a single forward pass, reducing per-token latency while maintaining output quality. This implementation leverages a two-model architecture where the primary model retains full authority over final token selection.

## How MTP Works in ds4

### The Two-Model Architecture

MTP requires exactly two GGUF files: your primary target model and a lightweight **support** model containing draft weights. According to the source code in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), the engine loads the support GGUF as a second memory mapping (see the comment *"MTP loads a second GGUF mapping …"*). The support model is typically orders of magnitude smaller than the main model, allowing it to generate candidate tokens with minimal computational overhead.

The primary model retains the full KV cache and authoritative logits, while the support model provides speculative continuations. This separation ensures that draft generation does not pollute the main model's context state, as the draft logits are computed transiently and **not** persisted in the KV cache.

### Draft Generation and Verification

During each decode step, the MTP model proposes up to *N* draft tokens specified by `--mtp-draft N`. The main model then performs a single verification forward pass that recomputes logits for the prefix and compares them against the draft predictions. If the draft matches the main model's distribution within a configurable confidence margin, ds4 commits all draft tokens at once.

The verification logic in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 56459-56507) implements two modes:

- **Approximate verification**: Uses the `--mtp-margin` threshold (or `DS4_MTP_MIN_MARGIN` environment variable) to determine acceptance based on probability ratios
- **Strict verification**: Activated by setting `DS4_MTP_STRICT=1`, forcing an exact match path that bypasses the confidence margin heuristic

When verification fails, the engine falls back to standard autoregressive decoding for that step, ensuring no quality degradation occurs.

## Configuring MTP in ds4

### Command-Line Options

The CLI parser in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) (lines 63015-63019) registers the following MTP-specific flags:

- `--mtp FILE`: Path to the MTP support GGUF (required to enable the feature)
- `--mtp-draft N`: Number of draft tokens to generate per step (default: 1)
- `--mtp-margin F`: Confidence margin for approximate verification; higher values increase conservatism
- `--mtp-timing`: Print acceptance counters and timing diagnostics to stderr
- `--mtp-spec-log`: Enable verbose speculative decoding logging
- `--mtp-conf-log`: Log verification confidence decisions for debugging

### Environment Variables

For finer runtime control, ds4 respects several environment variables:

- `DS4_MTP_MIN_MARGIN`: Overrides the `--mtp-margin` value
- `DS4_MTP_STRICT`: Forces exact verification when set to `1`
- `DS4_MTP_PROBE`: Enables probe mode for testing draft generation
- `DS4_MTP_FULL_LOGITS`: Computes full logits during verification rather than partial
- `DS4_MTP_KEEP_ACCEPTED`: Retains accepted draft tokens in specific edge cases

Note that when native session batching is active, MTP is automatically disabled (as documented in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) around line 23223).

## Practical Configuration Examples

Enable basic MTP with default settings:

```bash
./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
      --mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
      --temp 0

```

Increase draft depth to 2 tokens, which often provides the optimal latency reduction without excessive rejection rates:

```bash
./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
      --mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
      --mtp-draft 2 \
      --temp 0

```

Add a confidence margin to avoid slow partial accepts on uncertain drafts:

```bash
./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
      --mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
      --mtp-draft 2 \
      --mtp-margin 0.2 \
      --temp 0

```

Profile acceptance rates and timing overhead:

```bash
./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
      --mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
      --mtp-draft 2 \
      --mtp-timing \
      --temp 0

```

## Debugging and Testing

The logging instrumentation in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (around lines 55730-55743) provides visibility into the speculative decoding loop. For continuous integration testing, the test harness in [`tests/ds4_test.c`](https://github.com/antirez/ds4/blob/main/tests/ds4_test.c) supports the `--mtp-verify-depth` flag and the `DS4_TEST_MTP` environment variable to exercise verification logic without requiring full model weights.

When diagnosing low acceptance rates, enable `--mtp-conf-log` to inspect which draft tokens fail verification and why. The acceptance counters printed by `--mtp-timing` reveal the ratio of speculative successes to total steps, helping you tune `--mtp-draft` and `--mtp-margin` for your specific model combination.

## Summary

- **MTP requires two models**: A primary GGUF and a smaller support GGUF loaded via `--mtp`
- **Draft depth controls speculation**: Use `--mtp-draft N` to generate 1-N tokens per step, with 2 often being the practical sweet spot
- **Verification ensures quality**: The main model validates drafts using either strict or approximate matching controlled by `--mtp-margin` and `DS4_MTP_STRICT`
- **Instrumentation aids tuning**: Flags like `--mtp-timing` and `--mtp-conf-log` expose internal acceptance statistics
- **TP incompatibility**: MTP automatically disables when tensor parallelism batching is active in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c)

## Frequently Asked Questions

### What is the optimal --mtp-draft value for most use cases?

A draft depth of **2 tokens** typically provides the best latency improvement without excessive verification failures. Values above 3 often see diminishing returns due to cascading rejection rates, while `--mtp-draft 1` minimizes risk but leaves performance gains on the table. Profile your specific model pair using `--mtp-timing` to identify the inflection point.

### How does the --mtp-margin parameter affect token acceptance?

The `--mtp-margin` flag sets a probability ratio threshold for approximate verification. Higher values (e.g., 0.3-0.5) make the engine more conservative, accepting only drafts that closely match the main model's distribution. Lower values increase acceptance rates but risk quality degradation on edge cases. Set `DS4_MTP_STRICT=1` to bypass this heuristic entirely and require exact logit matches.

### Can I use MTP with batched inference or tensor parallelism?

No. According to the implementation in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) (line 23223), MTP is automatically disabled when native session batching or tensor parallelism is active. The speculative decoding loop assumes single-sequence generation contexts where draft and verification passes can be interleaved without batch synchronization overhead.

### How do I debug MTP acceptance rates in ds4?

Enable `--mtp-timing` to print acceptance counters and microseconds spent in draft vs. verification phases. For detailed per-token analysis, add `--mtp-conf-log` to see which specific draft positions failed verification and their confidence scores. The test suite in [`tests/ds4_test.c`](https://github.com/antirez/ds4/blob/main/tests/ds4_test.c) provides a `--mtp-verify-depth` option for isolated verification testing without full inference overhead.