# How to Configure Speculative Decoding Parameters for Maximum Throughput in DFlash

> Maximize DFlash throughput by configuring speculative decoding. Optimize block size, cache reuse with default target layers, and use sliding window for large prompts to maintain peak token generation speed and prevent overflow.

- Repository: [Z Lab/dflash](https://github.com/z-lab/dflash)
- Tags: how-to-guide
- Published: 2026-04-17

---

**Set `block_size` to the maximum value supported by your draft model configuration, use the default `target_layer_ids` for optimal cache reuse, and enable `sliding_window_size` only when processing prompts exceeding the model's maximum position embeddings to prevent memory overflow while maintaining peak token generation speed.**

DFlash is an open-source speculative decoding framework that accelerates large language model inference by running a smaller draft model in parallel with the target model. The `z-lab/dflash` repository implements a draft-then-verify pipeline where configuring the speculative decoding parameters correctly determines whether you achieve peak throughput or waste GPU cycles on rejected token blocks.

## Understanding DFlash's Speculative Decoding Pipeline

In DFlash, the `DFlashDraftModel` generates candidate token blocks while the target model verifies them in parallel. The core speculative loop in `stream_generate` (MLX backend) or `dflash_generate` (PyTorch backend) assembles draft-only token blocks using `block_size` and `mask_token_id`, then verifies them with the target model. Throughput optimization requires balancing the cost of target model verification against the acceptance rate of draft-generated tokens.

## Key Parameters for Throughput Optimization

### block_size

The `block_size` parameter controls how many tokens the draft model generates before the target model verifies them. Larger blocks amortize the cost of target model verification across more tokens, directly increasing throughput when the draft model maintains high accuracy. This parameter is defined in `DFlashDraftModel.__init__` for both the MLX and PyTorch backends.

Set `block_size` to the maximum value allowed by your model's configuration (`config.block_size`). Override only if memory constraints dictate smaller blocks or if your specific workload exhibits low draft acceptance rates.

### target_layer_ids

The `target_layer_ids` parameter determines which hidden-state layers of the target model are exposed to the draft model as context. Proper spacing reduces the amount of data the draft must copy and improves cache reuse between the models. This is computed in `DFlashDraftModel.__init__` from `config.dflash_config["target_layer_ids"]` or generated via `build_target_layer_ids`.

Use the evenly-spaced default shipped with the draft checkpoint, or generate a custom list with `build_target_layer_ids(num_target_layers, num_draft_layers)` to optimize for your specific architecture.

### sliding_window_size

The optional `sliding_window_size` enables a rotating KV-cache for very long prompts, preventing unbounded memory growth while still allowing large blocks. This is implemented in `load_draft` for the MLX backend and `DFlashDraftModel.__init__` for the PyTorch backend, creating a `RotatingKVCache` when specified.

Set this to a sensible power-of-two size (e.g., 4096 or 8192) when working with prompts exceeding `max_position_embeddings`. Leave as `None` for shorter inputs to avoid unnecessary cache rotation overhead.

### mask_token_id

The `mask_token_id` specifies the token used to pad speculative blocks and must match the target model's mask or padding token. This is stored in `DFlashConfig.mask_token_id` and propagated to the draft model. Use the value from the draft's [`config.json`](https://github.com/z-lab/dflash/blob/main/config.json) (usually 0) unless your target model uses a different padding scheme.

## Implementation in the MLX Backend

In [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py), the `load_draft` function (lines 60-90) constructs the `DFlashDraftModel` with the configuration parameters. The `stream_generate` function (lines 108-158) implements the core speculative loop, reading `block_size` and `mask_token_id` to assemble draft blocks.

```python
import mlx_lm
from dflash.model_mlx import load_draft, stream_generate

# Load the target model (any supported MLX model)

target = mlx_lm.load("Qwen/Qwen2-7B-Instruct")

# Load a draft model – request a large sliding window only if needed

draft = load_draft(
    draft_id="z-lab/dflash-qwen2-7b-draft",   # example draft checkpoint

    sliding_window_size=8192,                 # optional; None for short prompts

)

# Use the draft's native block size (the highest throughput setting)

block_size = draft.config.block_size

# Run speculative generation

for response in stream_generate(
    model=target,
    draft=draft,
    tokenizer=target.tokenizer,
    prompt="Explain the theory of relativity in simple terms.",
    block_size=block_size,        # largest possible block

    max_tokens=1024,
    temperature=0.0,
):
    print(response.text, end="")

```

## Implementation in the PyTorch Backend

In [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py), the `DFlashDraftModel` class (lines 12-23) initializes the draft model with configuration parameters. The `dflash_generate` function (lines 71-119) drives the speculative generation loop, using `block_size` to determine draft block dimensions.

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from dflash.model import DFlashDraftModel, dflash_generate

# Load the target model

target = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2-7B-Instruct", torch_dtype=torch.float16).cuda()
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2-7B-Instruct")

# Load a draft checkpoint (the config carries block_size & target_layer_ids)

draft = DFlashDraftModel.from_pretrained("z-lab/dflash-qwen2-7b-draft")
draft.block_size = draft.config.block_size          # ensures maximal block size

draft.sliding_window_size = 8192                    # optional for very long prompts

# Encode prompt

input_ids = tokenizer.encode("Summarize the plot of *Moby‑Dick*.", return_tensors="pt").cuda()

# Run speculative generation

output_ids = dflash_generate(
    model=draft,
    target=target,
    input_ids=input_ids,
    max_new_tokens=1024,
    stop_token_ids=tokenizer.eos_token_id,
    temperature=0.0,
    block_size=draft.block_size,   # largest block supported by the draft

)

print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

```

## Summary

To maximize throughput in DFlash's speculative decoding pipeline:

- Set `block_size` to the maximum value supported by your draft model's configuration to amortize verification costs across the largest possible token blocks.
- Use the default `target_layer_ids` shipped with your draft checkpoint, or generate optimized spacing with `build_target_layer_ids` to minimize data copying between draft and target models.
- Enable `sliding_window_size` (set to 4096 or 8192) only when processing prompts exceeding the model's maximum position embeddings to prevent KV-cache overflow.
- Ensure `mask_token_id` matches your target model's padding token to avoid mismatched speculative blocks.

## Frequently Asked Questions

### What is the optimal block_size for maximum throughput?

The optimal `block_size` is the maximum value defined in your draft model's configuration (`config.block_size`). Larger blocks amortize the cost of target model verification across more tokens, directly increasing throughput. Only reduce this value if you encounter memory constraints or if your draft model exhibits low acceptance rates on your specific workload.

### How does target_layer_ids affect performance?

The `target_layer_ids` parameter controls which hidden-state layers from the target model are exposed to the draft model. Properly spaced layer IDs reduce the volume of data copied between models and improve cache reuse, minimizing synchronization overhead. Use the evenly-spaced defaults provided with your draft checkpoint, or generate custom spacing using `build_target_layer_ids(num_target_layers, num_draft_layers)`.

### When should I use sliding_window_size?

Enable `sliding_window_size` when processing prompts that exceed the model's `max_position_embeddings` or when you need to prevent unbounded KV-cache growth during long generation sessions. Set this to a power-of-two value such as 4096 or 8192 to enable the `RotatingKVCache` mechanism. For prompts well within the model's context limits, leave this as `None` to avoid unnecessary cache rotation overhead.

### Can I mix MLX and PyTorch backends?

While DFlash supports both MLX and PyTorch backends, you cannot mix them within a single speculative decoding pipeline. The draft and target models must use the same backend framework because they share tensor formats and memory management strategies. Choose the MLX backend for Apple Silicon optimizations or the PyTorch backend for CUDA-accelerated inference on NVIDIA GPUs.