How to Configure DSpark Speculative Decoding for Faster Token Generation

To enable DSpark speculative decoding in DS4, pass the DSpark support GGUF with --mtp, enable the feature with --dspark, and use greedy decoding (--temp 0).

DSpark is DeepSeek's auxiliary draft model that accelerates token generation in the DS4 inference engine. When properly configured, DSpark can propose up to five future tokens per verification pass, significantly increasing effective throughput on favorable prompts. This guide covers the complete configuration process based on the antirez/ds4 source code.

How DSpark Speculative Decoding Works

DSpark operates as a multi-token prediction (MTP) system alongside the main Flash model. The architecture follows a two-phase pattern:

  1. Draft phase — DSpark reads the main model's hidden-state tensors and proposes a short token block (default: 5 tokens).
  2. Verification phase — The main Flash model processes the same hidden states and validates whether the drafted tokens match its own probability distribution.
  3. Commit phase — Accepted tokens advance the KV cache by the full block; rejected tokens fall back to standard single-token decoding.

This approach yields speed gains when the draft model accurately predicts the main model's behavior. However, the draft model adds computational overhead, making proper tuning essential.

Reference: The DSpark concept is documented in the repository README at lines 93-102.

Required Components

You need three artifacts to run DSpark speculative decoding:

Component Download Command Size
Flash model (e.g., q2-imatrix) ./download_model.sh q2-imatrix ~8 GiB
DSpark support GGUF ./download_model.sh dspark-support ~5.6 GiB
Compiled binary make (Metal), make cuda-spark (CUDA), or make strix-halo (ROCm) —

The DSpark support file is not a standalone model. It contains the auxiliary network weights that attach to the main Flash model at runtime.

Reference: Download instructions appear in the README at lines 108-112.

Enabling DSpark at Runtime

Basic Activation

The minimal invocation requires three flags: the main model, the MTP support file, and the --dspark toggle:

./ds4 -m ds4flash.gguf \
  --mtp gguf/DeepSeek-V4-Flash-DSpark-support.gguf \
  --dspark \
  --temp 0
  • --mtp specifies the DSpark support GGUF path
  • --dspark activates the experimental speculative path (defined in ds4_help.c at line 186)
  • --temp 0 enforces greedy decoding; sampling randomness defeats speculative verification

Tuning the Confidence Threshold

The --dspark-confidence parameter controls how aggressively the engine accepts draft tokens. The default value is 0.7 (70% probability).


# More permissive: accept drafts with ≥50% confidence

./ds4 ... --dspark-confidence 0.5

# Diagnostic mode: force full 5-token blocks regardless of quality

./ds4 ... --dspark-confidence 0

Lowering the threshold increases acceptance rates on predictable text (code, structured documents) but raises verification costs when drafts misalign. Setting confidence to zero disables pruning entirely—useful for benchmarking raw speculative throughput, but potentially slower on real prompts.

Reference: Flag parsing appears in ds4_cli.c at line 1853 and ds4_server.c at line 12709.

Strict Mode for Testing

The --dspark-strict flag loads DSpark tensors without activating speculation:

./ds4 ... --dspark --dspark-strict

This mode validates that the support file loads correctly and that hidden-state plumbing functions, without incurring speculative overhead. Documented in ds4_help.c at line 188.

Performance Optimization Strategies

Technique Configuration When to Apply
Greedy decoding --temp 0 Always required; sampling breaks verification logic
Permissive acceptance --dspark-confidence 0.5 Predictable text (repetitive code, narratives)
Maximum block forcing --dspark-confidence 0 Synthetic benchmarks only
Resident model loading Disable --ssd-streaming Avoid SSD latency for both draft and verifier
Modest KV footprint Use q2-imatrix instead of larger quants Preserves GPU memory for DSpark buffers

Profiling Draft Quality

Use the test flag --dspark-verify-depth to observe how many tokens survive verification:

./ds4 ... --dspark-verify-depth

This emits per-generation statistics showing actual acceptance rates. The flag definition resides in tests/ds4_test.c at line 6726.

Complete Working Example

#!/bin/bash

# Step 1: Download required models

./download_model.sh q2-imatrix
./download_model.sh dspark-support

# Step 2: Run with optimized speculative decoding

./ds4 \
  -m ds4flash.gguf \
  --mtp gguf/DeepSeek-V4-Flash-DSpark-support.gguf \
  --dspark \
  --dspark-confidence 0.5 \
  --temp 0 \
  -p "Define a Python dataclass for a 2D point with distance methods."

# Step 3: Monitor stderr for DSpark statistics, e.g.:

# ds4: DSpark stats accepted=4/5 draft=5 t/s=38.2 …

The statistics line is emitted from ds4.c at approximately line 56021, reporting accepted tokens, draft length, and effective tokens-per-second.

Server Deployment

The DS4 server binary accepts identical DSpark flags:

./ds4-server \
  --dspark \
  --mtp gguf/DeepSeek-V4-Flash-DSpark-support.gguf \
  --temp 0 \
  --dspark-confidence 0.6

All CLI parsing logic is shared between ds4, ds4-server, and ds4_agent.c.

Key Source Files

Understanding these files helps with troubleshooting and advanced tuning:

File Purpose
ds4.c Core engine; validates --dspark/--mtp pairing, emits statistics
ds4_help.c Generates help text; documents all DSpark flags
ds4_cli.c / ds4_server.c Command-line and server argument parsing
gguf-tools/deepseek4-quantize.c Support GGUF generation utilities
tests/ds4_test.c Unit tests for cache handling and verification depth

Summary

  • DSpark requires three components: Flash model, support GGUF, and --dspark flag
  • Greedy decoding (--temp 0) is mandatory for speculative verification to function
  • --dspark-confidence is the primary tuning knob: lower values increase acceptance but raise verification costs
  • Confidence of 0 forces full blocks for benchmarking; values 0.5-0.7 work best for production
  • Resident model loading and modest quants prevent memory pressure that slows the draft-verifier pipeline

Frequently Asked Questions

What happens if I forget --temp 0?

Sampling introduces randomness that breaks the deterministic verification assumption. The engine may accept or reject drafts unpredictably, eliminating throughput gains and potentially reducing performance below standard decoding.

Can I use DSpark with any GGUF model?

No. DSpark is specifically designed for DeepSeek-V4-Flash models. The support GGUF contains architecture-specific projection layers that map Flash hidden states to DSpark token predictions. Attempting to load unsupported models triggers an error in ds4.c during initialization.

Why would I use --dspark-strict instead of omitting --dspark entirely?

Strict mode validates that the support file loads correctly and that tensor shapes align without the performance variables of speculation. Use it when integrating a new Flash/DSpark version pair to isolate loading bugs from runtime behavior issues.

How do I know if DSpark is helping my specific workload?

Enable --dspark-verify-depth and observe the acceptance ratio in stderr output. Ratios above 60% typically yield net speedups; below 40% often indicate the draft model misaligns with your prompt style, and disabling DSpark may be faster.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →