How the MTP Speculative Decoding Path Differs from DSpark in ds4
The MTP speculative decoding path uses a legacy external head with per-token verification and cannot run under tensor parallelism, while DSpark uses integrated draft blocks with block-level verification and supports TP on the leader.
The ds4 inference engine implements two distinct speculative drafting mechanisms: MTP (Multi-Token Prediction) and DSpark. Both accelerate decoding by predicting multiple tokens ahead of verification, but they differ fundamentally in architecture, tensor-parallel compatibility, and configuration. Understanding these differences is essential for optimizing inference with GLM-DSA models.
Model Architecture and Implementation
MTP: External Legacy Head
MTP speculative decoding relies on a separate tiny model loaded as an "MTP GGUF" file. This external head attaches to the main model solely for draft-token probes and runs a discrete draft cycle.
In ds4.c, the enable check at lines 49794-49798 confirms the strict requirements:
/* ds4_engine_glm_mtp_spec_enabled – lines 49794-49798 */
The path activates only when:
- The engine is not CPU-only
- The model family is GLM-DSA
DS4_N_NEXTN_PREDICTcompile-time constant is non-zero- The
--glm-mtpflag is set
DSpark: Integrated Draft Blocks
DSpark speculative decoding uses a DSpark support GGUF containing "draft-block" tensors. Unlike MTP, the draft generation happens inside the main model graph via dedicated blocks rather than through an external attachment.
The engine validates DSpark availability in ds4.c at line 5203, where it checks for "DSpark metadata error" if tensors are missing or malformed:
/* DSpark tensor validation during model loading */
Enablement requires:
engine->support_kind == DS4_SUPPORT_DSPARK- A DSpark support GGUF loaded via
--mtp
Speculative Workflow Comparison
MTP Draft Cycle
The MTP path executes a per-token draft cycle as documented in ds4_tp.c lines 510-512:
- Run the MTP block to propose up to
--mtp-drafttokens - Verify each draft token with the MTP verifier (tiny row-count target verifier)
- Fall back to per-token decode on verification failure
/* MTP draft cycle – ds4_tp.c comment lines 510-512 */
This granular approach trades throughput for flexibility but creates overhead from repeated verification calls.
DSpark Block-Based Drafting
DSpark operates on draft blocks configured by --dspark-block-size. The block produces an entire token chunk in a single forward pass, with verification occurring at the block level rather than per-token.
Verification logic appears in ds4.c at line 61402 with DSpark verifier logs. Block rejection triggers either:
- Full block skip
- Fallback to regular decode path
This batch-oriented approach reduces verification overhead and improves throughput for coherent token sequences.
Tensor-Parallel Compatibility
MTP Falls Back Under TP
The legacy MTP path cannot operate under tensor parallelism. In ds4_tp_validate_engine_options at lines 510-513, the engine explicitly guards against TP usage with MTP:
/* ds4_tp.c lines 510-513 – TP guard for MTP */
When TP is active, MTP speculative decoding is disabled entirely, forcing per-token decode.
DSpark Works with TP on Leader
DSpark speculative drafting supports tensor parallelism through architectural differences. Its tensors install on a single executor tier and mirror across TP workers via the normal TP frame (DS4_TP_FRAME_VERIFY). The draft generation runs leader-only, but verification distributes across workers.
This makes DSpark the only speculative path compatible with distributed inference in ds4.
Configuration and CLI Flags
Both paths share the --mtp flag for loading support GGUFs but diverge in dedicated options:
| MTP Flags | DSpark Flags |
|---|---|
--mtp FILE |
--mtp FILE (same flag, different GGUF format) |
--mtp-draft N |
--dspark |
--mtp-margin |
--dspark-confidence 0.0-1.0 |
--glm-mtp |
--dspark-strict |
--glm-mtp-timing |
--dspark-block-size |
Source: ds4_help.c lines 179-185 (MTP) and 186-188 (DSpark)
Command-Line Examples
# MTP speculative decoding (legacy, non-TP only)
ds4 --mtp path/to/mtp.gguf --mtp-draft 2 --glm-mtp
# DSpark speculative decoding (TP-compatible)
ds4 --mtp path/to/dspark.gguf --dspark --dspark-confidence 0.7
Key Source Files
Understanding the implementation requires examining these files:
ds4.c– Core engine withds4_engine_glm_mtp_spec_enabled()enable-checks and DSpark support validationds4_tp.c– Tensor-parallel validation and speculative drafting comments (lines 510-513)ds4_help.c– CLI option definitions for both pathstests/ds4_test.c– Test harness for--mtp-verify-depthand--dspark-verify-depthgguf-tools/deepseek4-quantize.c– DSpark tensor creation tooling
Summary
- MTP uses an external legacy head with per-token verification and requires non-TP, non-CPU, GLM-DSA conditions
- DSpark uses integrated draft blocks with block-level verification and supports tensor parallelism on the leader
- Both paths load via
--mtpbut require different GGUF formats (MTP head vs. DSpark support tensors) - DSpark offers superior scalability for distributed deployments while MTP remains available for specialized single-node configurations
Frequently Asked Questions
Can I use MTP and DSpark simultaneously?
No, the ds4 engine treats them as mutually exclusive speculative paths. The support_kind field determines which path activates: DS4_SUPPORT_DSPARK enables DSpark, while the legacy MTP conditions must all be met for that path. Attempting to load both types of GGUFs creates undefined behavior.
Why does DSpark use the same --mtp flag as MTP?
Historical naming convention. The --mtp flag originally served only Multi-Token Prediction, then extended to accept DSpark support GGUFs. The engine distinguishes formats by tensor inspection—MTP GGUFs contain head-specific tensors, while DSpark GGUFs contain "draft-block" metadata validated at line 5203 of ds4.c.
How do I verify which speculative path is active at runtime?
Check the engine logs for initialization messages. DSpark reports "DSpark metadata" validation and tensor installation counts. MTP logs appear only when ds4_engine_glm_mtp_spec_enabled() returns true, prefixed with GLM-MTP timing markers if --glm-mtp-timing is enabled. Under TP, MTP silently disables with no log entry.
What happens when DSpark block verification fails?
The engine implements a tiered fallback per ds4.c line 61402. Partial block acceptance truncates to verified tokens; complete rejection skips the block entirely. The --dspark-strict flag modifies this behavior to require full verification or immediate fallback to regular decode, trading speculation depth for latency consistency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →