# How the MTP Speculative Decoding Path Differs from DSpark in ds4

> Discover the key differences between MTP speculative decoding and DSpark. Understand their verification methods, tensor parallelism support, and head implementations for ds4.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-04

---

**The MTP speculative decoding path uses a legacy external head with per-token verification and cannot run under tensor parallelism, while DSpark uses integrated draft blocks with block-level verification and supports TP on the leader.**

The **ds4** inference engine implements two distinct speculative drafting mechanisms: **MTP (Multi-Token Prediction)** and **DSpark**. Both accelerate decoding by predicting multiple tokens ahead of verification, but they differ fundamentally in architecture, tensor-parallel compatibility, and configuration. Understanding these differences is essential for optimizing inference with GLM-DSA models.

## Model Architecture and Implementation

### MTP: External Legacy Head

MTP speculative decoding relies on a **separate tiny model** loaded as an "MTP GGUF" file. This external head attaches to the main model solely for draft-token probes and runs a discrete draft cycle.

In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), the enable check at lines 49794-49798 confirms the strict requirements:

```c
/* ds4_engine_glm_mtp_spec_enabled – lines 49794-49798 */

```

The path activates only when:
- The engine is **not CPU-only**
- The model family is **GLM-DSA**
- `DS4_N_NEXTN_PREDICT` compile-time constant is **non-zero**
- The `--glm-mtp` flag is **set**

### DSpark: Integrated Draft Blocks

DSpark speculative decoding uses a **DSpark support GGUF** containing "draft-block" tensors. Unlike MTP, the draft generation happens **inside the main model graph** via dedicated blocks rather than through an external attachment.

The engine validates DSpark availability in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) at line 5203, where it checks for "DSpark metadata error" if tensors are missing or malformed:

```c
/* DSpark tensor validation during model loading */

```

Enablement requires:
- `engine->support_kind == DS4_SUPPORT_DSPARK`
- A DSpark support GGUF loaded via `--mtp`

## Speculative Workflow Comparison

### MTP Draft Cycle

The MTP path executes a **per-token draft cycle** as documented in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) lines 510-512:

1. Run the MTP block to propose up to `--mtp-draft` tokens
2. Verify each draft token with the **MTP verifier** (tiny row-count target verifier)
3. Fall back to per-token decode on verification failure

```c
/* MTP draft cycle – ds4_tp.c comment lines 510-512 */

```

This granular approach trades throughput for flexibility but creates overhead from repeated verification calls.

### DSpark Block-Based Drafting

DSpark operates on **draft blocks** configured by `--dspark-block-size`. The block produces an entire token chunk in a single forward pass, with verification occurring at the **block level** rather than per-token.

Verification logic appears in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) at line 61402 with DSpark verifier logs. Block rejection triggers either:
- Full block skip
- Fallback to regular decode path

This batch-oriented approach reduces verification overhead and improves throughput for coherent token sequences.

## Tensor-Parallel Compatibility

### MTP Falls Back Under TP

The legacy MTP path **cannot operate under tensor parallelism**. In `ds4_tp_validate_engine_options` at lines 510-513, the engine explicitly guards against TP usage with MTP:

```c
/* ds4_tp.c lines 510-513 – TP guard for MTP */

```

When TP is active, MTP speculative decoding is disabled entirely, forcing per-token decode.

### DSpark Works with TP on Leader

DSpark speculative drafting **supports tensor parallelism** through architectural differences. Its tensors install on a single executor tier and mirror across TP workers via the normal TP frame (`DS4_TP_FRAME_VERIFY`). The draft generation runs **leader-only**, but verification distributes across workers.

This makes DSpark the **only speculative path compatible with distributed inference** in ds4.

## Configuration and CLI Flags

Both paths share the `--mtp` flag for loading support GGUFs but diverge in dedicated options:

| MTP Flags | DSpark Flags |
|-----------|--------------|
| `--mtp FILE` | `--mtp FILE` (same flag, different GGUF format) |
| `--mtp-draft N` | `--dspark` |
| `--mtp-margin` | `--dspark-confidence 0.0-1.0` |
| `--glm-mtp` | `--dspark-strict` |
| `--glm-mtp-timing` | `--dspark-block-size` |

Source: [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) lines 179-185 (MTP) and 186-188 (DSpark)

### Command-Line Examples

```bash

# MTP speculative decoding (legacy, non-TP only)

ds4 --mtp path/to/mtp.gguf --mtp-draft 2 --glm-mtp

# DSpark speculative decoding (TP-compatible)

ds4 --mtp path/to/dspark.gguf --dspark --dspark-confidence 0.7

```

## Key Source Files

Understanding the implementation requires examining these files:

- **[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)** – Core engine with `ds4_engine_glm_mtp_spec_enabled()` enable-checks and DSpark support validation
- **[`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c)** – Tensor-parallel validation and speculative drafting comments (lines 510-513)
- **[`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c)** – CLI option definitions for both paths
- **[`tests/ds4_test.c`](https://github.com/antirez/ds4/blob/main/tests/ds4_test.c)** – Test harness for `--mtp-verify-depth` and `--dspark-verify-depth`
- **[`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c)** – DSpark tensor creation tooling

## Summary

- **MTP** uses an external legacy head with per-token verification and **requires non-TP, non-CPU, GLM-DSA conditions**
- **DSpark** uses integrated draft blocks with block-level verification and **supports tensor parallelism on the leader**
- Both paths load via `--mtp` but require **different GGUF formats** (MTP head vs. DSpark support tensors)
- **DSpark offers superior scalability** for distributed deployments while MTP remains available for specialized single-node configurations

## Frequently Asked Questions

### Can I use MTP and DSpark simultaneously?

No, the ds4 engine treats them as mutually exclusive speculative paths. The `support_kind` field determines which path activates: `DS4_SUPPORT_DSPARK` enables DSpark, while the legacy MTP conditions must all be met for that path. Attempting to load both types of GGUFs creates undefined behavior.

### Why does DSpark use the same `--mtp` flag as MTP?

Historical naming convention. The `--mtp` flag originally served only Multi-Token Prediction, then extended to accept DSpark support GGUFs. The engine distinguishes formats by tensor inspection—MTP GGUFs contain head-specific tensors, while DSpark GGUFs contain "draft-block" metadata validated at line 5203 of [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c).

### How do I verify which speculative path is active at runtime?

Check the engine logs for initialization messages. DSpark reports "DSpark metadata" validation and tensor installation counts. MTP logs appear only when `ds4_engine_glm_mtp_spec_enabled()` returns true, prefixed with GLM-MTP timing markers if `--glm-mtp-timing` is enabled. Under TP, MTP silently disables with no log entry.

### What happens when DSpark block verification fails?

The engine implements a tiered fallback per [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) line 61402. Partial block acceptance truncates to verified tokens; complete rejection skips the block entirely. The `--dspark-strict` flag modifies this behavior to require full verification or immediate fallback to regular decode, trading speculation depth for latency consistency.