DSpark Speculative Decoding: How to Enable DeepSeek V4 Flash Draft-Verify Acceleration
DSpark speculative decoding is an experimental auxiliary draft model that generates up to five future tokens using a lightweight ~5.6 GiB checkpoint, which the main DeepSeek V4 Flash model then verifies to reduce full verification passes and increase generation throughput.
The antirez/ds4 repository implements DSpark speculative decoding to accelerate inference for the DeepSeek V4 Flash model. By employing a separate draft model to predict token sequences before verification, this mechanism aims to minimize expensive forward passes through the main model during predictable text continuations such as code generation.
What is DSpark Speculative Decoding?
DSpark speculative decoding utilizes an auxiliary draft model released by DeepSeek that operates alongside the main DeepSeek V4 Flash model. The system works through a draft-verify loop where the draft model proposes candidate tokens and the main model verifies them in parallel.
The DSpark support GGUF stores a lightweight draft checkpoint (approximately 5.6 GiB) capable of proposing up to five future tokens by reusing hidden states from the main model. The main Flash model then verifies these proposals, committing only the accepted prefix. If a draft suffix is rejected or falls below a confidence threshold, the runtime falls back to ordinary target decoding.
This approach targets reduction in full verification passes to increase throughput, particularly for predictable continuations. However, because the draft work incurs overhead, performance gains are prompt-dependent; some inputs may show no speed improvement or may even degrade performance. DSpark speculative decoding is experimental and requires explicit activation.
How DSpark Speculative Decoding Works Under the Hood
The Draft-Verify Architecture
The architecture relies on several key components working in concert:
- DSpark support GGUF: Contains the draft model weights and metadata including block size and Markov rank
--mtpflag: Supplies the support GGUF to the engine using the existing Multi-Token Prediction path mechanism--dsparkflag: Activates the speculative decoder, triggering the draft-verify loop in the runtime- Confidence threshold (
--dspark-confidence): Prunes low-probability draft blocks with a default value of 0.7; setting to 0 forces fixed-size five-token blocks for diagnostics - Strict mode (
--dspark-strict): Loads the DSpark support model but maintains target-only decoding, useful for correctness testing and baseline comparison
Configuration Flag Implementation
The command-line parser in ds4_cli.c handles DSpark activation through explicit flag detection. At lines 1851-1860, the parser sets the relevant engine configuration fields:
// ds4_cli.c – handling of DSpark flags
} else if (!strcmp(arg, "--dspark")) {
c.engine.dspark = true;
} else if (!strcmp(arg, "--dspark-confidence")) {
c.engine.dspark = true;
c.engine.dspark_confidence_threshold =
parse_float_range(need_arg(&i, argc, argv, arg), arg, 0.0f, 1.0f);
c.engine.dspark_confidence_threshold_set = true;
} else if (!strcmp(arg, "--dspark-strict")) {
c.engine.dspark = true;
c.engine.dspark_strict = true;
}
The same flag handling appears in ds4_server.c (lines 12851-12858) for the server binary, ensuring consistent behavior across deployment modes.
How to Enable DSpark Speculative Decoding
1. Download the DSpark Support Model
Download the required support checkpoint (approximately 5.6 GiB) using the provided script:
./download_model.sh ds4f-dspark
This command retrieves DeepSeek-V4-Flash-DSpark-support-0731.gguf into the gguf/ directory.
2. Run with DSpark Activation
Execute the engine with both the main model and the DSpark support GGUF:
./ds4 -m ds4flash.gguf \
--mtp gguf/DeepSeek-V4-Flash-DSpark-support-0731.gguf \
--dspark --temp 0
Required parameters:
--mtpsupplies the support checkpoint path--dsparkactivates the speculative decoder--temp 0enforces greedy decoding (mandatory for DSpark operation)
3. Tune Acceptance Behavior
Adjust the speculative decoding behavior using optional flags:
-
Confidence pruning (default 0.7): Lower values accept more aggressive drafts
./ds4 ... --dspark-confidence 0.5 -
Fixed-size blocks: Set confidence to 0 for diagnostic five-token blocks
./ds4 ... --dspark-confidence 0 -
Strict mode: Load support model without speculative execution for testing
./ds4 ... --dspark-strict
Practical Command Examples
# Basic DSpark run with greedy decoding
./ds4 -m ds4flash.gguf \
--mtp gguf/DeepSeek-V4-Flash-DSpark-support-0731.gguf \
--dspark --temp 0
# Aggressive acceptance threshold at 0.6 confidence
./ds4 -m ds4flash.gguf \
--mtp gguf/DeepSeek-V4-Flash-DSpark-support-0731.gguf \
--dspark --dspark-confidence 0.6 --temp 0
# Strict mode for baseline performance comparison
./ds4 -m ds4flash.gguf \
--mtp gguf/DeepSeek-V4-Flash-DSpark-support-0731.gguf \
--dspark-strict --temp 0
These examples assume the main Flash GGUF (ds4flash.gguf) is already present, downloadable via ./download_model.sh ds4f-q2, ds4f-q2-q4, or ds4f-q4.
Summary
- DSpark speculative decoding employs a ~5.6 GiB auxiliary draft model to predict up to five tokens ahead of the main DeepSeek V4 Flash model, reducing verification passes.
- Enable the feature by downloading the support GGUF and invoking
ds4with--mtp,--dspark, and--temp 0flags. - Tune performance using
--dspark-confidence(default 0.7) to control acceptance thresholds, or set to 0 for fixed-size diagnostic blocks. - Test correctness using
--dspark-strictmode, which loads the support model while maintaining target-only decoding for baseline comparison. - Performance varies by prompt predictability; the feature is experimental and most effective on structured content like code.
Frequently Asked Questions
What hardware requirements does DSpark speculative decoding have?
DSpark speculative decoding requires sufficient VRAM to load both the main DeepSeek V4 Flash model and the ~5.6 GiB DSpark support GGUF simultaneously. The draft model adds computational overhead during the verify phase, so optimal performance requires GPUs with spare compute capacity to handle the draft-verify loop efficiently.
Why does DSpark require greedy decoding (--temp 0)?
The draft-verify mechanism relies on deterministic token prediction to maintain consistency between the draft model's proposals and the main model's verification path. Random sampling would desynchronize the draft and target distributions, causing frequent rejections of draft tokens and eliminating the throughput benefits of speculative execution.
How do I know if DSpark is actually improving my generation speed?
Monitor tokens-per-second metrics during generation and compare against --dspark-strict mode, which loads the support model but disables speculative decoding. According to the source implementation, DSpark reduces full verification passes for predictable continuations like code, but unpredictable prompts may show no improvement or minor slowdown due to draft computation overhead.
Can I use DSpark with quantization formats other than GGUF?
The current antirez/ds4 implementation specifically requires the DSpark support model in GGUF format, as evidenced by the download_model.sh ds4f-dspark command and the --mtp flag integration. The quantization tool gguf-tools/deepseek4-quantize.c (lines 2642-2650) handles DSpark support GGUF generation through --dspark-manifest and --dspark-support options, indicating tight format coupling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →