How MTP Speculative Decoding Works in the DeepSeek V4 Engine (antirez/ds4)
Multi-Token Prediction (MTP) speculative decoding accelerates generation by drafting short token sequences with a lightweight auxiliary model and verifying them in a single forward pass through the target model, committing drafts only when they match the target's argmax within a configurable margin.
The antirez/ds4 engine implements MTP speculative decoding as a legacy drafter path that proposes candidate token suffixes during greedy generation. This technique reduces per-token latency by accepting multiple tokens per target model evaluation when the lightweight MTP head accurately predicts the target model's outputs.
Architecture and Source Locations
The implementation spans several components within the ds4.c core:
- MTP Draft Generation (
ds4.c, lines 57973–60496): Executes the GLM MTP drafting loop to produce short suffixes using a separate raw cache. - Speculative Driver (
ds4.c, line 64773): Theds4_session_eval_speculative_argmax()function orchestrates drafting, verification, and result merging. - Target Verifier (
ds4.c, line 34432): Performs layer-major speculative verification over the entire draft suffix. - Acceptance Logic (
ds4.c, lines 64924–64944): Implements margin-based confidence checks usingDS4_MTP_MIN_MARGIN. - CLI Interface (
ds4_help.c): Exposes--mtp,--mtp-draft,--mtp-margin, and--glm-mtpflags.
Step-by-Step Execution Flow
Draft Generation with the MTP Head
When the engine operates in greedy mode (temperature ≤ 0), the GLM MTP block (ds4.c, lines 42828–43100) receives the current token embedding and hidden state (enorm, hnorm). The MTP head—a small auxiliary network loaded from a separate GGUF file—runs a few transformer layers (the "nextn" block) to predict the next one or two tokens. These draft tokens populate a dedicated MTP cache, isolated from the main KV cache to prevent side effects (see comment at line 15129).
Target Model Verification
After drafting, the engine invokes the layer-major speculative target verifier (ds4.c, line 34432). Unlike standard per-token decoding, this verifier runs the full target model once over the entire draft suffix, computing logits for each position and storing them in s->mtp_logits. This single forward pass evaluates all drafted positions simultaneously, minimizing overhead.
Acceptance Criteria and Token Commit
The acceptance test compares the draft logits against the target logits at each position:
- Argmax Match: The draft token must equal the target model's argmax at that position.
- Margin Check: The target logit must exceed the draft logit by at least the configured margin (
DS4_MTP_MIN_MARGIN, default 3.0).
If both conditions hold for a token, the engine commits it by copying the MTP cache state into the main KV cache via ds4_session_eval_speculative_argmax(). The process repeats from the new position; if verification fails, the engine falls back to standard per-token decoding.
Configuration and Usage
Command-Line Interface
Enable MTP speculative decoding via the CLI flags defined in ds4_help.c:
./ds4 --model deepseek_v4.gguf \
--mtp deepseek_mtp.gguf \
--glm-mtp \
--mtp-draft 2 \
--mtp-margin 3.0
--mtp FILE: Loads the auxiliary MTP head GGUF.--glm-mtp: Forces greedy-only mode (required for MTP).--mtp-draft N: Sets the maximum draft tokens per step (typically 1–2).--mtp-margin F: Configures the acceptance threshold (default 3.0).
C API Integration
Programmatically configure MTP through the ds4_cfg structure:
#include "ds4.h"
int main(void) {
ds4_cfg cfg = ds4_default_cfg();
cfg.model_path = "deepseek_v4.gguf";
cfg.mtp_path = "deepseek_mtp.gguf";
cfg.gen.temperature = 0.0f; // Required for MTP
cfg.mtp_draft = 2;
cfg.mtp_margin = 3.0f;
ds4_engine *engine = ds4_engine_new(&cfg);
ds4_session *sess = ds4_engine_open_session(engine);
// Automatically uses speculative decoding when conditions are met
int token = ds4_session_eval_speculative_argmax(sess, 0, 100);
ds4_engine_free(engine);
return 0;
}
The ds4_session_eval_speculative_argmax() function automatically routes through the MTP path when mtp_path is set and temperature is zero.
Key Implementation Details
Greedy-Only Constraint: MTP speculative decoding only activates when cli_greedy_argmax_requested() returns true (ds4_cli.c). This restriction exists because the acceptance logic assumes deterministic argmax sampling from the target model.
Separate Caching Strategy: The engine maintains distinct raw caches for the MTP head and target model. This isolation prevents the drafter's intermediate tensors from contaminating the target model's KV cache during verification.
Verification Depth: Under tensor-parallel regimes (ds4_tp.c), the engine falls back to per-token decoding rather than speculative batches, as noted in the TP implementation comments.
Summary
- MTP speculative decoding uses a lightweight auxiliary model to draft 1–2 tokens ahead during greedy generation.
- Verification occurs in a single layer-major pass through the target model (
ds4.c, line 34432). - Acceptance requires matching argmax tokens with a configurable logit margin (
DS4_MTP_MIN_MARGIN, default 3). - Configuration happens via
--mtpand related flags inds4_help.c, or through theds4_cfgAPI structure. - Fallback to standard decoding occurs when drafts fail verification or when tensor parallelism is active.
Frequently Asked Questions
What hardware requirements does MTP speculative decoding impose?
MTP requires loading an additional GGUF file into GPU memory, increasing VRAM usage by approximately 10–15% compared to the base model alone. The verifier runs the full target model, so compute requirements remain identical to standard decoding, though memory bandwidth utilization improves due to reduced total forward passes when drafts are accepted.
Why does MTP only work with greedy decoding (temperature = 0)?
The acceptance logic in ds4_session_eval_speculative_argmax() assumes deterministic sampling where the target model's argmax represents the canonical next token. Stochastic sampling would require probability distribution matching rather than argmax comparison, complicating the margin-based verification implemented at lines 64924–64944.
How should I tune the --mtp-margin parameter?
Start with the default value of 3.0. Decrease the margin (e.g., to 2.0) to increase acceptance rates at the risk of occasional quality degradation, or increase it (e.g., to 5.0) for stricter fidelity to the target model. Monitor acceptance rates via the DS4_MTP_CONF_LOG=1 environment variable to empirically determine the optimal threshold for your specific model and dataset.
Can MTP speculative decoding combine with other acceleration techniques?
MTP operates independently of quantization and tensor parallelism, though the source code in ds4_tp.c indicates that tensor-parallel mode disables speculative verification, falling back to per-token generation. Standard quantization schemes (Q4_K_M, Q8_0) apply normally to both the target model and MTP head GGUF files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →