How to Configure Speculative Decoding Parameters for Maximum Throughput in DFlash

Set block_size to the maximum value supported by your draft model configuration, use the default target_layer_ids for optimal cache reuse, and enable sliding_window_size only when processing prompts exceeding the model's maximum position embeddings to prevent memory overflow while maintaining peak token generation speed.

DFlash is an open-source speculative decoding framework that accelerates large language model inference by running a smaller draft model in parallel with the target model. The z-lab/dflash repository implements a draft-then-verify pipeline where configuring the speculative decoding parameters correctly determines whether you achieve peak throughput or waste GPU cycles on rejected token blocks.

Understanding DFlash's Speculative Decoding Pipeline

In DFlash, the DFlashDraftModel generates candidate token blocks while the target model verifies them in parallel. The core speculative loop in stream_generate (MLX backend) or dflash_generate (PyTorch backend) assembles draft-only token blocks using block_size and mask_token_id, then verifies them with the target model. Throughput optimization requires balancing the cost of target model verification against the acceptance rate of draft-generated tokens.

Key Parameters for Throughput Optimization

block_size

The block_size parameter controls how many tokens the draft model generates before the target model verifies them. Larger blocks amortize the cost of target model verification across more tokens, directly increasing throughput when the draft model maintains high accuracy. This parameter is defined in DFlashDraftModel.__init__ for both the MLX and PyTorch backends.

Set block_size to the maximum value allowed by your model's configuration (config.block_size). Override only if memory constraints dictate smaller blocks or if your specific workload exhibits low draft acceptance rates.

target_layer_ids

The target_layer_ids parameter determines which hidden-state layers of the target model are exposed to the draft model as context. Proper spacing reduces the amount of data the draft must copy and improves cache reuse between the models. This is computed in DFlashDraftModel.__init__ from config.dflash_config["target_layer_ids"] or generated via build_target_layer_ids.

Use the evenly-spaced default shipped with the draft checkpoint, or generate a custom list with build_target_layer_ids(num_target_layers, num_draft_layers) to optimize for your specific architecture.

sliding_window_size

The optional sliding_window_size enables a rotating KV-cache for very long prompts, preventing unbounded memory growth while still allowing large blocks. This is implemented in load_draft for the MLX backend and DFlashDraftModel.__init__ for the PyTorch backend, creating a RotatingKVCache when specified.

Set this to a sensible power-of-two size (e.g., 4096 or 8192) when working with prompts exceeding max_position_embeddings. Leave as None for shorter inputs to avoid unnecessary cache rotation overhead.

mask_token_id

The mask_token_id specifies the token used to pad speculative blocks and must match the target model's mask or padding token. This is stored in DFlashConfig.mask_token_id and propagated to the draft model. Use the value from the draft's config.json (usually 0) unless your target model uses a different padding scheme.

Implementation in the MLX Backend

In dflash/model_mlx.py, the load_draft function (lines 60-90) constructs the DFlashDraftModel with the configuration parameters. The stream_generate function (lines 108-158) implements the core speculative loop, reading block_size and mask_token_id to assemble draft blocks.

import mlx_lm
from dflash.model_mlx import load_draft, stream_generate

# Load the target model (any supported MLX model)

target = mlx_lm.load("Qwen/Qwen2-7B-Instruct")

# Load a draft model – request a large sliding window only if needed

draft = load_draft(
    draft_id="z-lab/dflash-qwen2-7b-draft",   # example draft checkpoint

    sliding_window_size=8192,                 # optional; None for short prompts

)

# Use the draft's native block size (the highest throughput setting)

block_size = draft.config.block_size

# Run speculative generation

for response in stream_generate(
    model=target,
    draft=draft,
    tokenizer=target.tokenizer,
    prompt="Explain the theory of relativity in simple terms.",
    block_size=block_size,        # largest possible block

    max_tokens=1024,
    temperature=0.0,
):
    print(response.text, end="")

Implementation in the PyTorch Backend

In dflash/model.py, the DFlashDraftModel class (lines 12-23) initializes the draft model with configuration parameters. The dflash_generate function (lines 71-119) drives the speculative generation loop, using block_size to determine draft block dimensions.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from dflash.model import DFlashDraftModel, dflash_generate

# Load the target model

target = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2-7B-Instruct", torch_dtype=torch.float16).cuda()
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2-7B-Instruct")

# Load a draft checkpoint (the config carries block_size & target_layer_ids)

draft = DFlashDraftModel.from_pretrained("z-lab/dflash-qwen2-7b-draft")
draft.block_size = draft.config.block_size          # ensures maximal block size

draft.sliding_window_size = 8192                    # optional for very long prompts

# Encode prompt

input_ids = tokenizer.encode("Summarize the plot of *Moby‑Dick*.", return_tensors="pt").cuda()

# Run speculative generation

output_ids = dflash_generate(
    model=draft,
    target=target,
    input_ids=input_ids,
    max_new_tokens=1024,
    stop_token_ids=tokenizer.eos_token_id,
    temperature=0.0,
    block_size=draft.block_size,   # largest block supported by the draft

)

print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

Summary

To maximize throughput in DFlash's speculative decoding pipeline:

  • Set block_size to the maximum value supported by your draft model's configuration to amortize verification costs across the largest possible token blocks.
  • Use the default target_layer_ids shipped with your draft checkpoint, or generate optimized spacing with build_target_layer_ids to minimize data copying between draft and target models.
  • Enable sliding_window_size (set to 4096 or 8192) only when processing prompts exceeding the model's maximum position embeddings to prevent KV-cache overflow.
  • Ensure mask_token_id matches your target model's padding token to avoid mismatched speculative blocks.

Frequently Asked Questions

What is the optimal block_size for maximum throughput?

The optimal block_size is the maximum value defined in your draft model's configuration (config.block_size). Larger blocks amortize the cost of target model verification across more tokens, directly increasing throughput. Only reduce this value if you encounter memory constraints or if your draft model exhibits low acceptance rates on your specific workload.

How does target_layer_ids affect performance?

The target_layer_ids parameter controls which hidden-state layers from the target model are exposed to the draft model. Properly spaced layer IDs reduce the volume of data copied between models and improve cache reuse, minimizing synchronization overhead. Use the evenly-spaced defaults provided with your draft checkpoint, or generate custom spacing using build_target_layer_ids(num_target_layers, num_draft_layers).

When should I use sliding_window_size?

Enable sliding_window_size when processing prompts that exceed the model's max_position_embeddings or when you need to prevent unbounded KV-cache growth during long generation sessions. Set this to a power-of-two value such as 4096 or 8192 to enable the RotatingKVCache mechanism. For prompts well within the model's context limits, leave this as None to avoid unnecessary cache rotation overhead.

Can I mix MLX and PyTorch backends?

While DFlash supports both MLX and PyTorch backends, you cannot mix them within a single speculative decoding pipeline. The draft and target models must use the same backend framework because they share tensor formats and memory management strategies. Choose the MLX backend for Apple Silicon optimizations or the PyTorch backend for CUDA-accelerated inference on NVIDIA GPUs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →