# Performance Trade-offs Between MTP Speculative Decoding and DSpark in ds4

> Explore MTP speculative decoding vs DSpark in ds4. Discover how MTP saves VRAM with single-token drafting for speed, while DSpark boosts throughput with batched multi-token drafting, requiring more memory.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: performance
- Published: 2026-08-09

---

**MTP speculative decoding minimizes VRAM usage with single-token drafting for modest speed-ups, while DSpark maximizes throughput via batched multi-token drafting at the cost of loading a full support model and increased memory consumption.**

The `ds4` inference engine by antirez implements two distinct speculative decoding strategies to accelerate transformer generation. Understanding the performance trade-offs between MTP speculative decoding and DSpark is critical for optimizing latency and throughput based on available hardware constraints and deployment scenarios.

## Architectural Design Goals

The fundamental difference between these approaches lies in their drafting philosophy and parallelization strategy.

### MTP Legacy Mode

MTP (Multi-Token Prediction) operates as a legacy "one-token-ahead" drafting mechanism. In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), the engine defines this path under `DS4_SUPPORT_MTP_LEGACY`, where each draft step executes the MTP sub-graph before falling back to per-token decoding. This design targets minimal memory overhead, requiring only a small MTP GGUF file containing legacy tensors. As implemented in the source, MTP defaults to a draft depth of 1 (configurable via `--mtp-draft`), making it suitable for environments where VRAM is scarce.

### DSpark Batch Drafting

DSpark (DeepSeek V4 Flash) introduces a batch-drafting architecture that uses a separate **DSpark support model**. According to [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c), this support model shares the base model architecture but employs a reduced hidden size. The system drafts multiple tokens simultaneously using configurable block sizes (default 5 via `--dspark-block-size`). In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 56700-56856, DSpark is defined under `DS4_SUPPORT_DSPARK` and is explicitly designed for tensor-parallel (TP) setups, allowing tensors to be placed on dedicated executor tiers.

## Memory and Resource Utilization

The memory footprint represents the most significant operational trade-off between these methods.

**MTP speculative decoding** allocates only a few extra tensors for the MTP sub-graph. The engine falls back to regular decoding automatically if the MTP model cannot be loaded ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) line 57565), ensuring robustness in memory-constrained environments.

**DSpark** requires loading a full DSpark support GGUF, adding a complete set of expert tensors to VRAM. As noted in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 57582-57596, this substantially increases memory consumption compared to the MTP path. The implementation allows DSpark to be disabled (`--dspark-strict`) if verification fails or VRAM is insufficient ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 56855-56856).

## GPU Utilization and Verification Overhead

Kernel execution patterns differ dramatically between the two approaches.

MTP executes a tiny sub-graph on each draft step, resulting in frequent small kernel launches that often leave the GPU under-utilized. Verification occurs per individual token, creating repeated cheap but frequent verification passes.

DSpark batches draft tokens into blocks, allowing the GPU to process larger work-groups per kernel launch. The verifier operates on entire blocks rather than individual tokens ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 62610-63003), reducing the number of verification passes. This batching yields higher GPU utilization and better throughput, particularly evident in the scheduler logs showing "spec enter/skip" and "partial drafted/verified" states.

## Latency and Scalability Characteristics

Latency profiles favor different use cases for each method.

**MTP** exhibits low per-step latency, making it effective for very short drafts. However, the pipeline stalls after each token, limiting overall speed-up. It scales poorly across multiple GPUs or larger batch sizes because each draft step requires a separate kernel launch. Tensor-parallel support exists but is restricted; [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) line 510 notes that speculative drafting is allowed only on the leader rank.

**DSpark** incurs higher per-step latency due to block preparation overhead, but the amortized latency per token drops significantly when block size exceeds 1. The architecture is explicitly designed for tensor-parallel deployments, enabling scaling across GPUs with proper workload distribution.

## Implementation and Configuration

Enable MTP speculative decoding with minimal configuration:

```c
/* Enable MTP speculative decoding (default draft depth = 1) */
int main(int argc, char **argv) {
    ds4_options opt = {0};
    opt.mtp_file = getenv("DS4_TEST_MTP");   // legacy MTP GGUF
    opt.mtp_draft = 1;                       // max draft tokens
    ds4_engine *engine = ds4_new(&opt);
    ds4_generate(engine, prompt);
}

```

Configure DSpark for high-throughput batch drafting:

```c
/* Enable DSpark with a 5-token draft block */
int main(int argc, char **argv) {
    ds4_options opt = {0};
    opt.mtp_file = getenv("DS4_TEST_DSPARK");   // DSpark support GGUF
    opt.dspark = true;                          // turn on DSpark
    opt.dspark_block_size = 5;                  // draft 5 tokens per block
    ds4_engine *engine = ds4_new(&opt);
    ds4_generate(engine, prompt);
}

```

CLI options defined in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) provide runtime control through `--mtp`, `--mtp-draft`, `--dspark`, and `--dspark-block-size` flags.

## Summary

- **MTP speculative decoding** consumes minimal VRAM using legacy single-token drafting, making it the optimal fallback for memory-constrained environments despite limited GPU utilization.
- **DSpark** requires loading a full support model (generated via [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c)) and increases VRAM usage significantly, but delivers superior throughput via batched token drafting.
- MTP verification processes tokens individually with low latency per step, while DSpark amortizes overhead across configurable blocks (default 5 tokens).
- DSpark scales efficiently across tensor-parallel GPU configurations, whereas MTP is restricted to leader-only execution in distributed setups.
- Choose MTP when VRAM is limited and draft sequences are short; select DSpark when maximizing throughput and hardware resources permit the additional memory overhead.

## Frequently Asked Questions

### How do I choose between MTP and DSpark for my ds4 deployment?

Select **MTP speculative decoding** when running on hardware with limited VRAM or when serving models with strict memory constraints, as it requires only a small GGUF file and minimal tensor allocation. Choose **DSpark** when you have sufficient GPU memory available and need maximum generation throughput, particularly for long sequences or high batch sizes where batched drafting provides amortized latency benefits.

### What are the default configuration values for each drafting method?

According to [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) and the engine initialization in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), MTP defaults to a draft depth of 1 token (`--mtp-draft 1`), while DSpark defaults to a block size of 5 tokens (`--dspark-block-size 5`). These defaults reflect their respective design philosophies: MTP targets minimal overhead per step, while DSpark optimizes for batch efficiency.

### Can I use DSpark on multi-GPU setups?

Yes, DSpark is explicitly designed for tensor-parallel (TP) deployments. The source code in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 56700-56856 implements support for placing DSpark tensors on dedicated executor tiers, enabling scalable performance across multiple GPUs. In contrast, [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) line 510 indicates that MTP speculative drafting is restricted to the leader rank only.

### What happens if the DSpark support model fails to load?

The ds4 engine includes fallback logic at [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) lines 56855-56856 that allows DSpark to be disabled via the `--dspark-strict` flag if verification fails or if VRAM is insufficient. If DSpark initialization fails and strict mode is not enabled, the engine can fall back to standard decoding or MTP if available, ensuring service continuity even when resources are constrained.