MTP (Multi-Token Prediction) Support in ds4: Configuration and Usage Guide
Multi-Token Prediction (MTP) in ds4 is a speculative decoding mechanism that uses a lightweight support GGUF model to draft multiple tokens per step, allowing the main model to verify them in parallel for faster greedy generation.
The antirez/ds4 inference engine implements Multi-Token Prediction (MTP) as an optional acceleration path for greedy decoding scenarios. By loading a secondary "draft" model alongside your primary GGUF, ds4 can propose and validate short token sequences in a single forward pass, reducing per-token latency while maintaining output quality. This implementation leverages a two-model architecture where the primary model retains full authority over final token selection.
How MTP Works in ds4
The Two-Model Architecture
MTP requires exactly two GGUF files: your primary target model and a lightweight support model containing draft weights. According to the source code in ds4.c, the engine loads the support GGUF as a second memory mapping (see the comment "MTP loads a second GGUF mapping …"). The support model is typically orders of magnitude smaller than the main model, allowing it to generate candidate tokens with minimal computational overhead.
The primary model retains the full KV cache and authoritative logits, while the support model provides speculative continuations. This separation ensures that draft generation does not pollute the main model's context state, as the draft logits are computed transiently and not persisted in the KV cache.
Draft Generation and Verification
During each decode step, the MTP model proposes up to N draft tokens specified by --mtp-draft N. The main model then performs a single verification forward pass that recomputes logits for the prefix and compares them against the draft predictions. If the draft matches the main model's distribution within a configurable confidence margin, ds4 commits all draft tokens at once.
The verification logic in ds4.c (lines 56459-56507) implements two modes:
- Approximate verification: Uses the
--mtp-marginthreshold (orDS4_MTP_MIN_MARGINenvironment variable) to determine acceptance based on probability ratios - Strict verification: Activated by setting
DS4_MTP_STRICT=1, forcing an exact match path that bypasses the confidence margin heuristic
When verification fails, the engine falls back to standard autoregressive decoding for that step, ensuring no quality degradation occurs.
Configuring MTP in ds4
Command-Line Options
The CLI parser in ds4_cli.c (lines 63015-63019) registers the following MTP-specific flags:
--mtp FILE: Path to the MTP support GGUF (required to enable the feature)--mtp-draft N: Number of draft tokens to generate per step (default: 1)--mtp-margin F: Confidence margin for approximate verification; higher values increase conservatism--mtp-timing: Print acceptance counters and timing diagnostics to stderr--mtp-spec-log: Enable verbose speculative decoding logging--mtp-conf-log: Log verification confidence decisions for debugging
Environment Variables
For finer runtime control, ds4 respects several environment variables:
DS4_MTP_MIN_MARGIN: Overrides the--mtp-marginvalueDS4_MTP_STRICT: Forces exact verification when set to1DS4_MTP_PROBE: Enables probe mode for testing draft generationDS4_MTP_FULL_LOGITS: Computes full logits during verification rather than partialDS4_MTP_KEEP_ACCEPTED: Retains accepted draft tokens in specific edge cases
Note that when native session batching is active, MTP is automatically disabled (as documented in ds4_tp.c around line 23223).
Practical Configuration Examples
Enable basic MTP with default settings:
./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
--mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
--temp 0
Increase draft depth to 2 tokens, which often provides the optimal latency reduction without excessive rejection rates:
./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
--mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
--mtp-draft 2 \
--temp 0
Add a confidence margin to avoid slow partial accepts on uncertain drafts:
./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
--mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
--mtp-draft 2 \
--mtp-margin 0.2 \
--temp 0
Profile acceptance rates and timing overhead:
./ds4 -m gguf/DeepSeek-V4-Flash.gguf \
--mtp gguf/DeepSeek-V4-Flash-MTP-support.gguf \
--mtp-draft 2 \
--mtp-timing \
--temp 0
Debugging and Testing
The logging instrumentation in ds4.c (around lines 55730-55743) provides visibility into the speculative decoding loop. For continuous integration testing, the test harness in tests/ds4_test.c supports the --mtp-verify-depth flag and the DS4_TEST_MTP environment variable to exercise verification logic without requiring full model weights.
When diagnosing low acceptance rates, enable --mtp-conf-log to inspect which draft tokens fail verification and why. The acceptance counters printed by --mtp-timing reveal the ratio of speculative successes to total steps, helping you tune --mtp-draft and --mtp-margin for your specific model combination.
Summary
- MTP requires two models: A primary GGUF and a smaller support GGUF loaded via
--mtp - Draft depth controls speculation: Use
--mtp-draft Nto generate 1-N tokens per step, with 2 often being the practical sweet spot - Verification ensures quality: The main model validates drafts using either strict or approximate matching controlled by
--mtp-marginandDS4_MTP_STRICT - Instrumentation aids tuning: Flags like
--mtp-timingand--mtp-conf-logexpose internal acceptance statistics - TP incompatibility: MTP automatically disables when tensor parallelism batching is active in
ds4_tp.c
Frequently Asked Questions
What is the optimal --mtp-draft value for most use cases?
A draft depth of 2 tokens typically provides the best latency improvement without excessive verification failures. Values above 3 often see diminishing returns due to cascading rejection rates, while --mtp-draft 1 minimizes risk but leaves performance gains on the table. Profile your specific model pair using --mtp-timing to identify the inflection point.
How does the --mtp-margin parameter affect token acceptance?
The --mtp-margin flag sets a probability ratio threshold for approximate verification. Higher values (e.g., 0.3-0.5) make the engine more conservative, accepting only drafts that closely match the main model's distribution. Lower values increase acceptance rates but risk quality degradation on edge cases. Set DS4_MTP_STRICT=1 to bypass this heuristic entirely and require exact logit matches.
Can I use MTP with batched inference or tensor parallelism?
No. According to the implementation in ds4_tp.c (line 23223), MTP is automatically disabled when native session batching or tensor parallelism is active. The speculative decoding loop assumes single-sequence generation contexts where draft and verification passes can be interleaved without batch synchronization overhead.
How do I debug MTP acceptance rates in ds4?
Enable --mtp-timing to print acceptance counters and microseconds spent in draft vs. verification phases. For detailed per-token analysis, add --mtp-conf-log to see which specific draft positions failed verification and their confidence scores. The test suite in tests/ds4_test.c provides a --mtp-verify-depth option for isolated verification testing without full inference overhead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →