How the MTP Speculative Decoding Path Differs from DSpark in ds4

The MTP speculative decoding path uses a legacy external head with per-token verification and cannot run under tensor parallelism, while DSpark uses integrated draft blocks with block-level verification and supports TP on the leader.

The ds4 inference engine implements two distinct speculative drafting mechanisms: MTP (Multi-Token Prediction) and DSpark. Both accelerate decoding by predicting multiple tokens ahead of verification, but they differ fundamentally in architecture, tensor-parallel compatibility, and configuration. Understanding these differences is essential for optimizing inference with GLM-DSA models.

Model Architecture and Implementation

MTP: External Legacy Head

MTP speculative decoding relies on a separate tiny model loaded as an "MTP GGUF" file. This external head attaches to the main model solely for draft-token probes and runs a discrete draft cycle.

In ds4.c, the enable check at lines 49794-49798 confirms the strict requirements:

/* ds4_engine_glm_mtp_spec_enabled – lines 49794-49798 */

The path activates only when:

  • The engine is not CPU-only
  • The model family is GLM-DSA
  • DS4_N_NEXTN_PREDICT compile-time constant is non-zero
  • The --glm-mtp flag is set

DSpark: Integrated Draft Blocks

DSpark speculative decoding uses a DSpark support GGUF containing "draft-block" tensors. Unlike MTP, the draft generation happens inside the main model graph via dedicated blocks rather than through an external attachment.

The engine validates DSpark availability in ds4.c at line 5203, where it checks for "DSpark metadata error" if tensors are missing or malformed:

/* DSpark tensor validation during model loading */

Enablement requires:

  • engine->support_kind == DS4_SUPPORT_DSPARK
  • A DSpark support GGUF loaded via --mtp

Speculative Workflow Comparison

MTP Draft Cycle

The MTP path executes a per-token draft cycle as documented in ds4_tp.c lines 510-512:

  1. Run the MTP block to propose up to --mtp-draft tokens
  2. Verify each draft token with the MTP verifier (tiny row-count target verifier)
  3. Fall back to per-token decode on verification failure
/* MTP draft cycle – ds4_tp.c comment lines 510-512 */

This granular approach trades throughput for flexibility but creates overhead from repeated verification calls.

DSpark Block-Based Drafting

DSpark operates on draft blocks configured by --dspark-block-size. The block produces an entire token chunk in a single forward pass, with verification occurring at the block level rather than per-token.

Verification logic appears in ds4.c at line 61402 with DSpark verifier logs. Block rejection triggers either:

  • Full block skip
  • Fallback to regular decode path

This batch-oriented approach reduces verification overhead and improves throughput for coherent token sequences.

Tensor-Parallel Compatibility

MTP Falls Back Under TP

The legacy MTP path cannot operate under tensor parallelism. In ds4_tp_validate_engine_options at lines 510-513, the engine explicitly guards against TP usage with MTP:

/* ds4_tp.c lines 510-513 – TP guard for MTP */

When TP is active, MTP speculative decoding is disabled entirely, forcing per-token decode.

DSpark Works with TP on Leader

DSpark speculative drafting supports tensor parallelism through architectural differences. Its tensors install on a single executor tier and mirror across TP workers via the normal TP frame (DS4_TP_FRAME_VERIFY). The draft generation runs leader-only, but verification distributes across workers.

This makes DSpark the only speculative path compatible with distributed inference in ds4.

Configuration and CLI Flags

Both paths share the --mtp flag for loading support GGUFs but diverge in dedicated options:

MTP Flags DSpark Flags
--mtp FILE --mtp FILE (same flag, different GGUF format)
--mtp-draft N --dspark
--mtp-margin --dspark-confidence 0.0-1.0
--glm-mtp --dspark-strict
--glm-mtp-timing --dspark-block-size

Source: ds4_help.c lines 179-185 (MTP) and 186-188 (DSpark)

Command-Line Examples


# MTP speculative decoding (legacy, non-TP only)

ds4 --mtp path/to/mtp.gguf --mtp-draft 2 --glm-mtp

# DSpark speculative decoding (TP-compatible)

ds4 --mtp path/to/dspark.gguf --dspark --dspark-confidence 0.7

Key Source Files

Understanding the implementation requires examining these files:

  • ds4.c – Core engine with ds4_engine_glm_mtp_spec_enabled() enable-checks and DSpark support validation
  • ds4_tp.c – Tensor-parallel validation and speculative drafting comments (lines 510-513)
  • ds4_help.c – CLI option definitions for both paths
  • tests/ds4_test.c – Test harness for --mtp-verify-depth and --dspark-verify-depth
  • gguf-tools/deepseek4-quantize.c – DSpark tensor creation tooling

Summary

  • MTP uses an external legacy head with per-token verification and requires non-TP, non-CPU, GLM-DSA conditions
  • DSpark uses integrated draft blocks with block-level verification and supports tensor parallelism on the leader
  • Both paths load via --mtp but require different GGUF formats (MTP head vs. DSpark support tensors)
  • DSpark offers superior scalability for distributed deployments while MTP remains available for specialized single-node configurations

Frequently Asked Questions

Can I use MTP and DSpark simultaneously?

No, the ds4 engine treats them as mutually exclusive speculative paths. The support_kind field determines which path activates: DS4_SUPPORT_DSPARK enables DSpark, while the legacy MTP conditions must all be met for that path. Attempting to load both types of GGUFs creates undefined behavior.

Why does DSpark use the same --mtp flag as MTP?

Historical naming convention. The --mtp flag originally served only Multi-Token Prediction, then extended to accept DSpark support GGUFs. The engine distinguishes formats by tensor inspection—MTP GGUFs contain head-specific tensors, while DSpark GGUFs contain "draft-block" metadata validated at line 5203 of ds4.c.

How do I verify which speculative path is active at runtime?

Check the engine logs for initialization messages. DSpark reports "DSpark metadata" validation and tensor installation counts. MTP logs appear only when ds4_engine_glm_mtp_spec_enabled() returns true, prefixed with GLM-MTP timing markers if --glm-mtp-timing is enabled. Under TP, MTP silently disables with no log entry.

What happens when DSpark block verification fails?

The engine implements a tiered fallback per ds4.c line 61402. Partial block acceptance truncates to verified tokens; complete rejection skips the block entirely. The --dspark-strict flag modifies this behavior to require full verification or immediate fallback to regular decode, trading speculation depth for latency consistency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →