# DFlash Speculative Decoding: Why Block Diffusion Outperforms Traditional Draft-LM Methods

> Discover how DFlash speculative decoding achieves 3-6x speedup over traditional methods by using a block-diffusion head for parallel token generation with minimal memory. Learn more.

- Repository: [Z Lab/dflash](https://github.com/z-lab/dflash)
- Tags: deep-dive
- Published: 2026-04-17

---

**DFlash replaces the separate draft language model with a lightweight block-diffusion head that generates 16-32 tokens in parallel, achieving 3-6× speedup with minimal memory overhead compared to conventional speculative decoding.**

DFlash is an open-source speculative decoding implementation from the `z-lab/dflash` repository that introduces block-diffusion-based drafting. Unlike traditional approaches that rely on a smaller language model to generate candidate tokens sequentially, DFlash uses a diffusion model to predict entire token blocks in parallel, fundamentally changing the speed-memory trade-off in large language model inference.

## DFlash vs. Traditional Speculative Decoding

Traditional speculative decoding methods—such as those implemented in vLLM or DeepSpeed—use a discrete **draft language model** (typically 2-3B parameters) to generate candidate tokens one at a time. DFlash fundamentally differs by treating the draft as a **block-diffusion process** that denoises masked token blocks conditioned on the target model's hidden states.

| Feature | Traditional Draft-LM | DFlash Block Diffusion |
|---------|---------------------|------------------------|
| **Draft model size** | 2-3B parameters (~10-15% of target) | Few MB diffusion head; orders of magnitude smaller |
| **Generation granularity** | 1 token per draft step | 16-32 tokens per block in parallel |
| **Memory overhead** | 2× KV cache (target + draft) | Shared KV cache; draft adds minimal memory |
| **Speed gains** | ~2-3× on GPU | 4-6× on Apple Silicon (MLX), 3-5× on GPU/CPU |
| **Training cost** | Full LM training/fine-tuning required | Small diffusion head trained on target hidden states |

## Implementation Architecture

### Block Diffusion Draft Model

The core innovation resides in `DFlashDraftModel.forward` within [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) (lines 33-47). The draft model receives **noise embeddings** representing a masked token block and **target hidden states** from selected layers of the main model:

```python
hidden_states = noise_embedding
target_hidden = self.hidden_norm(self.fc(target_hidden))
position_embeddings = self.rotary_emb(hidden_states, position_ids)
for layer in self.layers:
    hidden_states = layer(
        hidden_states=hidden_states,
        target_hidden=target_hidden,
        ...
    )

```

This architecture allows the draft to generate `block_size` tokens (typically 16-32) in a single forward pass, rather than autoregressively.

### Block-Wise Acceptance Logic

DFlash verifies entire blocks against the target model's posterior distribution. The acceptance logic in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) (line 35) computes the longest matching prefix:

```python
acceptance_length = (block_output_ids[:, 1:] == posterior[:, :-1]).cumprod(dim=1).sum(dim=1)[0].item()

```

Only the accepted prefix is committed to the sequence, and the KV cache is cropped to the new position. This block-level verification reduces the overhead of token-by-token rejection.

### Sliding-Window KV Management

For long-context generation, DFlash implements optional sliding-window KV caching. When configured without a sliding window, the system can trim the draft's prompt cache via `trim_prompt_cache` to prevent unbounded growth. This is particularly important in the MLX backend ([`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py)), where memory constraints are stricter than on server GPUs.

## Code Examples

### Transformers Backend (PyTorch)

The primary API for GPU/CPU inference uses `spec_generate` in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) (lines 49-66):

```python
from transformers import AutoModelForCausalLM, AutoModel, AutoTokenizer

# Load target LLM

target = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3.5-27B", torch_dtype="auto", device_map="auto"
).eval()

# Load DFlash draft model (tiny diffusion checkpoint)

draft = AutoModel.from_pretrained(
    "z-lab/Qwen3.5-27B-DFlash", trust_remote_code=True, dtype="auto", device_map="auto"
).eval()

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-27B")
messages = [{"role": "user", "content": "Explain quantum tunneling in simple terms."}]
input_ids = tokenizer.apply_chat_template(
    messages, return_tensors="pt", add_generation_prompt=True, enable_thinking=False
).to(draft.device)

# Speculative generation

output_ids = draft.spec_generate(
    target=target,
    input_ids=input_ids,
    max_new_tokens=256,
    stop_token_ids=[tokenizer.eos_token_id],
    temperature=0.0,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=False))

```

### SGLang Server Integration

DFlash integrates with SGLang's scheduler for production deployment:

```bash
export SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1
python -m sglang.launch_server \
    --model-path Qwen/Qwen3.5-35B-A3B \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path z-lab/Qwen3.5-35B-A3B-DFlash \
    --speculative-num-draft-tokens 16 \
    --tp-size 1 \
    --attention-backend trtllm_mha \
    --speculative-draft-attention-backend fa4 \
    --mem-fraction-static 0.75 \
    --mamba-scheduler-strategy extra_buffer \
    --trust-remote-code

```

### MLX Backend (Apple Silicon)

For Apple Silicon devices, use the pure-Python diffusion implementation in [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py) (lines 54-68):

```python
from dflash.model_mlx import load, load_draft, stream_generate

model, tokenizer = load("Qwen/Qwen3.5-4B")
draft = load_draft("z-lab/Qwen3.5-4B-DFlash", sliding_window_size=None)

prompt = "Write a haiku about sunrise."
for chunk in stream_generate(
    model, draft, tokenizer, prompt,
    block_size=16, max_tokens=128, temperature=0.6
):
    print(chunk.text, end="", flush=True)

```

## Summary

- **DFlash** replaces the traditional draft language model with a **block-diffusion head** that generates 16-32 tokens in parallel, rather than one token at a time.
- The draft model is **orders of magnitude smaller** (few MB vs. 2-3B parameters) and shares the target model's KV cache, eliminating the 2× memory overhead of conventional speculative decoding.
- **Block-wise acceptance** in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) verifies entire token blocks against the target posterior, accepting the longest matching prefix to maximize throughput.
- DFlash achieves **4-6× speedup** on Apple Silicon (MLX) and **3-5×** on GPU/CPU, particularly excelling in long-context scenarios where memory efficiency is critical.
- The implementation provides unified APIs across **Transformers**, **SGLang**, and **MLX** backends, requiring only a lightweight draft checkpoint (`z-lab/<model>-DFlash`) to enable speculative generation.

## Frequently Asked Questions

### How does DFlash differ from standard vLLM speculative decoding?

Standard vLLM speculative decoding relies on a separate **draft language model** (typically 2-3B parameters) that generates tokens autoregressively. DFlash replaces this with a **block-diffusion model** that predicts entire blocks of 16-32 tokens simultaneously using the target model's hidden states. This reduces draft model size from billions to millions of parameters and eliminates the need for separate KV cache management, cutting memory overhead by roughly half.

### What hardware platforms support DFlash?

DFlash ships with three official backends: **PyTorch/Transformers** for NVIDIA and AMD GPUs, **SGLang** for production server deployment with tensor parallelism, and **MLX** for Apple Silicon (M1/M2/M3). The MLX backend is particularly optimized for DFlash's block-diffusion approach, achieving 4-6× throughput gains on resource-constrained mobile devices where traditional draft language models would not fit in memory.

### How does block-wise acceptance work in DFlash?

Instead of verifying tokens one-by-one, DFlash compares the entire draft block against the target model's posterior distribution in parallel. The acceptance logic in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) computes the longest prefix where draft tokens match target predictions using a cumulative product of equality checks. Only the accepted prefix is committed, and the KV cache is cropped to that position, minimizing wasted computation while maximizing the number of tokens generated per forward pass.

### Can I use DFlash with existing model checkpoints?

Yes. DFlash is designed as a **plug-and-play** augmentation. You load your existing target model (e.g., Qwen, Llama, or Mistral) from Hugging Face or a local path, then load the corresponding DFlash draft checkpoint from `z-lab/<model-name>-DFlash`. The draft model reuses the target's tokenizer and positional embeddings, requiring no architectural changes to the base model. Integration is available via the `spec_generate` API for Transformers or the `--speculative-algorithm DFLASH` flag for SGLang.