How ds4 Speculative Decoding Works: Implementation Guide and Setup

TLDR: ds4 uses a lightweight MTP draft model to propose future tokens that are verified against the main model's greedy output, accepting matches until a mismatch triggers a KV-cache rollback, guaranteeing bit-for-bit identical results to autoregressive decoding while accelerating generation on GPU hardware.

ds4, the high-performance inference engine developed by antirez, implements speculative decoding (also called MTP draft decoding) to minimize per-token latency during transformer generation. This technique relies on a smaller draft model to predict token sequences, which the main model then verifies in parallel. The implementation guarantees that final outputs remain identical to a pure greedy autoregressive pass while significantly improving throughput on CUDA, Metal, and ROCm backends.

Architecture of ds4 Speculative Decoding

The core mechanism resides in ds4.c within the ds4_session_eval_speculative_argmax function, which orchestrates a five-stage draft-verify-commit pipeline.

Draft Generation with the MTP Model

The process begins when the MTP (Multi-Token Prediction) draft model generates a short sequence of proposed tokens. As implemented at line 56530 of ds4.c, this lightweight model executes a forward pass to produce a configurable number of draft tokens controlled by the --mtp-draft N parameter. The draft suffix represents the model's prediction of the next N tokens in the sequence.

Token Verification and Commit

The main target model then evaluates the proposed suffix layer-by-layer to verify each token against its own greedy argmax selection. The verification logic at line 44115 of ds4.c compares the draft token against the target model's chosen token at each position. When tokens match, they are committed to the output stream and the KV-cache; upon the first mismatch, verification halts immediately.

KV-Cache Management and Rollback

Draft tokens are stored in a speculative KV ring buffer separate from the main cache. According to the source at line 54117 of ds4.c, when verification succeeds, these entries merge into the regular KV cache to maintain state consistency. If verification fails, the engine executes a rollback by discarding the speculative KV rows and resuming generation from the last committed token, ensuring the state remains equivalent to a non-speculative greedy run.

Batched Verification for Efficiency

When draft depth exceeds two tokens, ds4 optimizes GPU utilization through batched verification. The code at line 45625 of ds4.c implements a "Batched output head for speculative verification" that processes multiple speculative tokens simultaneously, reducing kernel launch overhead and improving memory bandwidth efficiency on the target model.

How to Enable Speculative Decoding in ds4

Enabling this feature requires the optional MTP draft model and GPU support, as the speculative code path is guarded by #ifndef DS4_NO_GPU and falls back to standard greedy decoding on CPU-only builds.

1. Download the MTP Draft Model

Fetch the lightweight draft model using the provided utility script referenced in gguf-tools/imatrix/dataset/rendered_prompts.txt at line 35823:

./download_model.sh mtp

This retrieves the MTP.gguf file containing the draft network weights.

2. Launch ds4 with the MTP Flag

Activate speculative decoding by passing both the main model and the draft model to the CLI, as parsed in ds4_cli.c:

ds4 --model DeepSeek-V4-Flash.gguf \
    --mtp MTP.gguf \
    --port 8080

The --mtp flag instructs the engine to load the draft model and initialize the speculative evaluation loop defined in ds4_tp.c.

3. Configure Draft Depth

Control the speculation window size using the --mtp-draft parameter. The default value is 1, but increasing this can improve throughput at the cost of higher verification overhead:

ds4 --model DeepSeek-V4-Flash.gguf \
    --mtp MTP.gguf \
    --mtp-draft 2

Higher values trigger the batched verification paths described in the "Batched output head" implementation.

4. Verify Correctness

Validate that speculative commits match autoregressive tokens using the test harness flag defined in tests/ds4_test.c at line 6594:

ds4 ... --mtp-verify-depth

This runs a parity check ensuring that speculative decoding produces identical results to greedy generation.

Testing and Verification

The test suite in tests/ds4_test.c includes comprehensive validation of the speculative pipeline. The --mtp-verify-depth flag triggers assertions that compare speculative token commits against reference autoregressive outputs, confirming that the KV-cache rollback mechanisms and draft verification logic maintain mathematical equivalence to the non-accelerated path.

Summary

  • ds4 speculative decoding uses an MTP draft model to predict token sequences that the main model verifies in parallel.
  • The ds4_session_eval_speculative_argmax function in ds4.c handles the core draft-verify-commit cycle, with specific logic at lines 44115 and 54117 managing verification and rollback.
  • Draft tokens reside in a speculative KV ring buffer that merges on commit or discards on mismatch, preserving greedy decoding equivalence.
  • Enable the feature by running ./download_model.sh mtp and launching with --mtp MTP.gguf --mtp-draft N.
  • This capability requires GPU support and is disabled in DS4_NO_GPU builds.

Frequently Asked Questions

What is the MTP model in ds4?

The MTP (Multi-Token Prediction) model is a lightweight draft network shipped separately from the main weights. According to the source in ds4_tp.c, this model predicts future tokens to create a "draft suffix" that the main model verifies, allowing ds4 to accept multiple tokens per forward pass when the draft agrees with the target model's greedy selection.

Does speculative decoding change the output quality?

No. The implementation guarantees bit-for-bit identical output compared to standard greedy decoding. As noted in the rollback logic at line 54117 of ds4.c, any verification failure triggers an immediate KV-cache rollback to the last committed state, ensuring the final sequence matches exactly what autoregressive generation would produce.

Can I use speculative decoding on CPU-only builds?

No. Speculative decoding requires GPU acceleration. The codebase wraps speculative logic with #ifndef DS4_NO_GPU guards, and CPU-only builds automatically fall back to single-step greedy decoding even if --mtp flags are provided.

How do I tune the draft depth for optimal performance?

Use the --mtp-draft N flag where N is typically between 1 and 4. The optimal value depends on your draft model's accuracy and GPU memory bandwidth. Higher values engage the batched verification head (line 45625 in ds4.c) but increase wasted computation if the draft model predictions diverge frequently from the target model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →