DFlash Speculative Decoding: Why Block Diffusion Outperforms Traditional Draft-LM Methods
DFlash replaces the separate draft language model with a lightweight block-diffusion head that generates 16-32 tokens in parallel, achieving 3-6× speedup with minimal memory overhead compared to conventional speculative decoding.
DFlash is an open-source speculative decoding implementation from the z-lab/dflash repository that introduces block-diffusion-based drafting. Unlike traditional approaches that rely on a smaller language model to generate candidate tokens sequentially, DFlash uses a diffusion model to predict entire token blocks in parallel, fundamentally changing the speed-memory trade-off in large language model inference.
DFlash vs. Traditional Speculative Decoding
Traditional speculative decoding methods—such as those implemented in vLLM or DeepSpeed—use a discrete draft language model (typically 2-3B parameters) to generate candidate tokens one at a time. DFlash fundamentally differs by treating the draft as a block-diffusion process that denoises masked token blocks conditioned on the target model's hidden states.
| Feature | Traditional Draft-LM | DFlash Block Diffusion |
|---|---|---|
| Draft model size | 2-3B parameters (~10-15% of target) | Few MB diffusion head; orders of magnitude smaller |
| Generation granularity | 1 token per draft step | 16-32 tokens per block in parallel |
| Memory overhead | 2× KV cache (target + draft) | Shared KV cache; draft adds minimal memory |
| Speed gains | ~2-3× on GPU | 4-6× on Apple Silicon (MLX), 3-5× on GPU/CPU |
| Training cost | Full LM training/fine-tuning required | Small diffusion head trained on target hidden states |
Implementation Architecture
Block Diffusion Draft Model
The core innovation resides in DFlashDraftModel.forward within dflash/model.py (lines 33-47). The draft model receives noise embeddings representing a masked token block and target hidden states from selected layers of the main model:
hidden_states = noise_embedding
target_hidden = self.hidden_norm(self.fc(target_hidden))
position_embeddings = self.rotary_emb(hidden_states, position_ids)
for layer in self.layers:
hidden_states = layer(
hidden_states=hidden_states,
target_hidden=target_hidden,
...
)
This architecture allows the draft to generate block_size tokens (typically 16-32) in a single forward pass, rather than autoregressively.
Block-Wise Acceptance Logic
DFlash verifies entire blocks against the target model's posterior distribution. The acceptance logic in dflash/model.py (line 35) computes the longest matching prefix:
acceptance_length = (block_output_ids[:, 1:] == posterior[:, :-1]).cumprod(dim=1).sum(dim=1)[0].item()
Only the accepted prefix is committed to the sequence, and the KV cache is cropped to the new position. This block-level verification reduces the overhead of token-by-token rejection.
Sliding-Window KV Management
For long-context generation, DFlash implements optional sliding-window KV caching. When configured without a sliding window, the system can trim the draft's prompt cache via trim_prompt_cache to prevent unbounded growth. This is particularly important in the MLX backend (dflash/model_mlx.py), where memory constraints are stricter than on server GPUs.
Code Examples
Transformers Backend (PyTorch)
The primary API for GPU/CPU inference uses spec_generate in dflash/model.py (lines 49-66):
from transformers import AutoModelForCausalLM, AutoModel, AutoTokenizer
# Load target LLM
target = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen3.5-27B", torch_dtype="auto", device_map="auto"
).eval()
# Load DFlash draft model (tiny diffusion checkpoint)
draft = AutoModel.from_pretrained(
"z-lab/Qwen3.5-27B-DFlash", trust_remote_code=True, dtype="auto", device_map="auto"
).eval()
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-27B")
messages = [{"role": "user", "content": "Explain quantum tunneling in simple terms."}]
input_ids = tokenizer.apply_chat_template(
messages, return_tensors="pt", add_generation_prompt=True, enable_thinking=False
).to(draft.device)
# Speculative generation
output_ids = draft.spec_generate(
target=target,
input_ids=input_ids,
max_new_tokens=256,
stop_token_ids=[tokenizer.eos_token_id],
temperature=0.0,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=False))
SGLang Server Integration
DFlash integrates with SGLang's scheduler for production deployment:
export SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1
python -m sglang.launch_server \
--model-path Qwen/Qwen3.5-35B-A3B \
--speculative-algorithm DFLASH \
--speculative-draft-model-path z-lab/Qwen3.5-35B-A3B-DFlash \
--speculative-num-draft-tokens 16 \
--tp-size 1 \
--attention-backend trtllm_mha \
--speculative-draft-attention-backend fa4 \
--mem-fraction-static 0.75 \
--mamba-scheduler-strategy extra_buffer \
--trust-remote-code
MLX Backend (Apple Silicon)
For Apple Silicon devices, use the pure-Python diffusion implementation in dflash/model_mlx.py (lines 54-68):
from dflash.model_mlx import load, load_draft, stream_generate
model, tokenizer = load("Qwen/Qwen3.5-4B")
draft = load_draft("z-lab/Qwen3.5-4B-DFlash", sliding_window_size=None)
prompt = "Write a haiku about sunrise."
for chunk in stream_generate(
model, draft, tokenizer, prompt,
block_size=16, max_tokens=128, temperature=0.6
):
print(chunk.text, end="", flush=True)
Summary
- DFlash replaces the traditional draft language model with a block-diffusion head that generates 16-32 tokens in parallel, rather than one token at a time.
- The draft model is orders of magnitude smaller (few MB vs. 2-3B parameters) and shares the target model's KV cache, eliminating the 2× memory overhead of conventional speculative decoding.
- Block-wise acceptance in
dflash/model.pyverifies entire token blocks against the target posterior, accepting the longest matching prefix to maximize throughput. - DFlash achieves 4-6× speedup on Apple Silicon (MLX) and 3-5× on GPU/CPU, particularly excelling in long-context scenarios where memory efficiency is critical.
- The implementation provides unified APIs across Transformers, SGLang, and MLX backends, requiring only a lightweight draft checkpoint (
z-lab/<model>-DFlash) to enable speculative generation.
Frequently Asked Questions
How does DFlash differ from standard vLLM speculative decoding?
Standard vLLM speculative decoding relies on a separate draft language model (typically 2-3B parameters) that generates tokens autoregressively. DFlash replaces this with a block-diffusion model that predicts entire blocks of 16-32 tokens simultaneously using the target model's hidden states. This reduces draft model size from billions to millions of parameters and eliminates the need for separate KV cache management, cutting memory overhead by roughly half.
What hardware platforms support DFlash?
DFlash ships with three official backends: PyTorch/Transformers for NVIDIA and AMD GPUs, SGLang for production server deployment with tensor parallelism, and MLX for Apple Silicon (M1/M2/M3). The MLX backend is particularly optimized for DFlash's block-diffusion approach, achieving 4-6× throughput gains on resource-constrained mobile devices where traditional draft language models would not fit in memory.
How does block-wise acceptance work in DFlash?
Instead of verifying tokens one-by-one, DFlash compares the entire draft block against the target model's posterior distribution in parallel. The acceptance logic in dflash/model.py computes the longest prefix where draft tokens match target predictions using a cumulative product of equality checks. Only the accepted prefix is committed, and the KV cache is cropped to that position, minimizing wasted computation while maximizing the number of tokens generated per forward pass.
Can I use DFlash with existing model checkpoints?
Yes. DFlash is designed as a plug-and-play augmentation. You load your existing target model (e.g., Qwen, Llama, or Mistral) from Hugging Face or a local path, then load the corresponding DFlash draft checkpoint from z-lab/<model-name>-DFlash. The draft model reuses the target's tokenizer and positional embeddings, requiring no architectural changes to the base model. Integration is available via the spec_generate API for Transformers or the --speculative-algorithm DFLASH flag for SGLang.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →