MTP (Multi-Token Prediction) Support in ds4: Configuration and Usage Guide

Multi-Token Prediction (MTP) in ds4 is a speculative decoding mechanism that uses a lightweight support GGUF model to draft multiple tokens per step, allowing the main model to verify them in parallel for faster greedy generation.

The antirez/ds4 inference engine implements Multi-Token Prediction (MTP) as an optional acceleration path for greedy decoding scenarios. By loading a secondary "draft" model alongside your primary GGUF, ds4 can propose and validate short token sequences in a single forward pass, reducing per-token latency while maintaining output quality. This implementation leverages a two-model architecture where the primary model retains full authority over final token selection.

How MTP Works in ds4

The Two-Model Architecture

MTP requires exactly two GGUF files: your primary target model and a lightweight support model containing draft weights. According to the source code in ds4.c, the engine loads the support GGUF as a second memory mapping (see the comment "MTP loads a second GGUF mapping …"). The support model is typically orders of magnitude smaller than the main model, allowing it to generate candidate tokens with minimal computational overhead.

The primary model retains the full KV cache and authoritative logits, while the support model provides speculative continuations. This separation ensures that draft generation does not pollute the main model's context state, as the draft logits are computed transiently and not persisted in the KV cache.

Draft Generation and Verification

During each decode step, the MTP model proposes up to N draft tokens specified by --mtp-draft N. The main model then performs a single verification forward pass that recomputes logits for the prefix and compares them against the draft predictions. If the draft matches the main model's distribution within a configurable confidence margin, ds4 commits all draft tokens at once.

The verification logic in ds4.c (lines 56459-56507) implements two modes:

  • Approximate verification: Uses the --mtp-margin threshold (or DS4_MTP_MIN_MARGIN environment variable) to determine acceptance based on probability ratios
  • Strict verification: Activated by setting DS4_MTP_STRICT=1, forcing an exact match path that bypasses the confidence margin heuristic

When verification fails, the engine falls back to standard autoregressive decoding for that step, ensuring no quality degradation occurs.

Configuring MTP in ds4

Command-Line Options

The CLI parser in ds4_cli.c (lines 63015-63019) registers the following MTP-specific flags:

  • --mtp FILE: Path to the MTP support GGUF (required to enable the feature)
  • --mtp-draft N: Number of draft tokens to generate per step (default: 1)
  • --mtp-margin F: Confidence margin for approximate verification; higher values increase conservatism
  • --mtp-timing: Print acceptance counters and timing diagnostics to stderr
  • --mtp-spec-log: Enable verbose speculative decoding logging
  • --mtp-conf-log: Log verification confidence decisions for debugging

Environment Variables

For finer runtime control, ds4 respects several environment variables:

  • DS4_MTP_MIN_MARGIN: Overrides the --mtp-margin value
  • DS4_MTP_STRICT: Forces exact verification when set to 1
  • DS4_MTP_PROBE: Enables probe mode for testing draft generation
  • DS4_MTP_FULL_LOGITS: Computes full logits during verification rather than partial
  • DS4_MTP_KEEP_ACCEPTED: Retains accepted draft tokens in specific edge cases

Note that when native session batching is active, MTP is automatically disabled (as documented in ds4_tp.c around line 23223).

Practical Configuration Examples

Enable basic MTP with default settings:

./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
      --mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
      --temp 0

Increase draft depth to 2 tokens, which often provides the optimal latency reduction without excessive rejection rates:

./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
      --mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
      --mtp-draft 2 \
      --temp 0

Add a confidence margin to avoid slow partial accepts on uncertain drafts:

./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
      --mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
      --mtp-draft 2 \
      --mtp-margin 0.2 \
      --temp 0

Profile acceptance rates and timing overhead:

./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
      --mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
      --mtp-draft 2 \
      --mtp-timing \
      --temp 0

Debugging and Testing

The logging instrumentation in ds4.c (around lines 55730-55743) provides visibility into the speculative decoding loop. For continuous integration testing, the test harness in tests/ds4_test.c supports the --mtp-verify-depth flag and the DS4_TEST_MTP environment variable to exercise verification logic without requiring full model weights.

When diagnosing low acceptance rates, enable --mtp-conf-log to inspect which draft tokens fail verification and why. The acceptance counters printed by --mtp-timing reveal the ratio of speculative successes to total steps, helping you tune --mtp-draft and --mtp-margin for your specific model combination.

Summary

  • MTP requires two models: A primary GGUF and a smaller support GGUF loaded via --mtp
  • Draft depth controls speculation: Use --mtp-draft N to generate 1-N tokens per step, with 2 often being the practical sweet spot
  • Verification ensures quality: The main model validates drafts using either strict or approximate matching controlled by --mtp-margin and DS4_MTP_STRICT
  • Instrumentation aids tuning: Flags like --mtp-timing and --mtp-conf-log expose internal acceptance statistics
  • TP incompatibility: MTP automatically disables when tensor parallelism batching is active in ds4_tp.c

Frequently Asked Questions

What is the optimal --mtp-draft value for most use cases?

A draft depth of 2 tokens typically provides the best latency improvement without excessive verification failures. Values above 3 often see diminishing returns due to cascading rejection rates, while --mtp-draft 1 minimizes risk but leaves performance gains on the table. Profile your specific model pair using --mtp-timing to identify the inflection point.

How does the --mtp-margin parameter affect token acceptance?

The --mtp-margin flag sets a probability ratio threshold for approximate verification. Higher values (e.g., 0.3-0.5) make the engine more conservative, accepting only drafts that closely match the main model's distribution. Lower values increase acceptance rates but risk quality degradation on edge cases. Set DS4_MTP_STRICT=1 to bypass this heuristic entirely and require exact logit matches.

Can I use MTP with batched inference or tensor parallelism?

No. According to the implementation in ds4_tp.c (line 23223), MTP is automatically disabled when native session batching or tensor parallelism is active. The speculative decoding loop assumes single-sequence generation contexts where draft and verification passes can be interleaved without batch synchronization overhead.

How do I debug MTP acceptance rates in ds4?

Enable --mtp-timing to print acceptance counters and microseconds spent in draft vs. verification phases. For detailed per-token analysis, add --mtp-conf-log to see which specific draft positions failed verification and their confidence scores. The test suite in tests/ds4_test.c provides a --mtp-verify-depth option for isolated verification testing without full inference overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →