How to Configure the Block Size in DFlash for Optimal Performance

Block size in DFlash controls how many tokens the draft model predicts in each speculative step, typically ranging from 4 to 16 tokens, where values around 8–16 provide the best balance between throughput and acceptance rate on most hardware.

Configuring the block size is essential when optimizing speculative decoding performance in DFlash, an open-source inference acceleration library. The block size determines how many tokens the draft model generates before the target model verifies them, directly impacting the trade-off between raw speed and draft acceptance rates. In the z-lab/dflash repository, this parameter is exposed through model configuration files, Python APIs, and CLI tools.

What Is Block Size in DFlash?

Block size defines the number of tokens the draft model predicts in a single speculative "block" before the target model validates them. It governs the fundamental trade-off between throughput (larger blocks mean more tokens per draft) and acceptance rate (larger blocks are less likely to be fully accepted, potentially causing costly rollbacks).

According to the source code in dflash/model_mlx.py (lines 40-44), the block_size is read from the draft model's config.json and stored in the DFlashConfig class. Similarly, in dflash/model.py (lines 19-20), the PyTorch DFlashDraftModel initializes self.block_size from the configuration during model load.

Locating Block Size in the Source Code

The block size parameter flows through three critical files in the repository:

  • dflash/model_mlx.py (lines 40-44 and 59-60): Defines DFlashConfig.block_size loaded from JSON and implements the fallback logic in stream_generate() where block_size = draft.config.block_size if not explicitly provided.

  • dflash/model.py (lines 19-20 and 76-77): Stores the attribute in DFlashDraftModel.block_size and handles the fallback in dflash_generate() with the logic block_size = model.block_size if block_size is None else block_size.

  • dflash/benchmark.py (line 227): Exposes the --block-size CLI argument for performance testing.

Finding the Optimal Block Size

Empirical testing across the z-lab/dflash codebase suggests that most models achieve peak effective throughput with block sizes between 8 and 16 tokens. The exact optimal value depends on model architecture, hardware capabilities, and prompt length.

  • Small blocks (1–4 tokens) deliver the highest acceptance rates (nearly 100%) but provide minimal speculative acceleration since the draft model only predicts a few tokens per step.

  • Medium blocks (8–16 tokens) strike the practical balance, typically yielding 2–4× speedups on GPU or Apple Silicon MLX hardware while maintaining respectable acceptance rates above 60-80%.

  • Large blocks (32+ tokens) maximize parallel speculative computation but often suffer acceptance rates below 30%, causing frequent rollbacks that negate the performance gains.

How to Configure the Block Size

You can set the block size through four methods, depending on whether you need a permanent configuration or a temporary override.

Via the Draft Model Configuration

Edit the draft model's config.json to permanently set the default for all inference runs:

{
  "block_size": 16
}

This value is automatically loaded into DFlashConfig.block_size when the model initializes, as implemented in dflash/model_mlx.py lines 40-44.

Via the Generation API

Pass an explicit block_size parameter to override the config default without modifying files.

MLX Backend using stream_generate:

from dflash.model_mlx import stream_generate

for result in stream_generate(
    model, draft, tokenizer, prompt,
    block_size=12,  # Override config default

    max_tokens=512
):
    print(result.text, end="", flush=True)

The fallback logic at lines 59-60 of dflash/model_mlx.py uses your provided value instead of draft.config.block_size.

PyTorch Backend using spec_generate:

from dflash.model import DFlashDraftModel

draft = DFlashDraftModel.from_pretrained("z-lab/Qwen3.5-4B-DFlash")
output_ids = draft.spec_generate(
    target_model,
    input_ids,
    block_size=10,  # Custom block size

    max_new_tokens=128
)

This passes through to dflash_generate where lines 76-77 of dflash/model.py apply the override.

Via Command-Line Benchmarks

Use the --block-size flag when running the benchmark utility to test different values:

python -m dflash.benchmark \
    --backend mlx \
    --model Qwen/Qwen3.5-4B \
    --draft-model z-lab/Qwen3.5-4B-DFlash \
    --block-size 14 \
    --dataset gsm8k \
    --max-samples 64

The argument parser at line 227 of dflash/benchmark.py handles this input and reports per-block throughput and acceptance statistics.

Via Environment Variable

For containerized deployments, set DFLASH_BLOCK_SIZE before launching the inference server. The server code reads this variable and forwards it to the underlying stream_generate function.

Practical Workflow for Optimization

Follow this workflow to determine the ideal block size for your specific hardware and model combination:

  1. Run a benchmark sweep across 4, 8, 12, and 16 tokens:
for bs in 4 8 12 16; do
  python -m dflash.benchmark \
      --backend mlx \
      --draft-model z-lab/Qwen3.5-4B-DFlash \
      --block-size $bs \
      --max-samples 32
done
  1. Analyze the output for throughput (tok/s) and mean acceptance length. Throughput will rise with block size until acceptance rate drops precipitously.

  2. Lock the configuration by updating config.json with the winning value or hardcoding it in your production inference script.

Production Considerations

When deploying DFlash in production environments, keep these constraints in mind:

  • Do not exceed the native block size: The code automatically caps values at int(draft.config.block_size), so passing larger values silently reduces them to the model's native limit.

  • Enable sliding windows for long sequences: When generating very long text, combine larger block sizes with --draft-sliding-window-size to prevent KV cache memory blow-up.

  • Monitor cache pressure: The MLX backend calls mx.clear_cache() every 256 tokens in stream_generate. If you encounter out-of-memory errors, reduce the block size or enable sliding windows.

Summary

  • Block size in DFlash determines how many tokens the draft model generates speculatively before target verification, sourced from config.json via DFlashConfig in dflash/model_mlx.py.
  • The optimal range is typically 8–16 tokens, balancing the speed of speculative generation against the acceptance rate of the target model.
  • Override the default via the block_size parameter in stream_generate() or spec_generate(), or use the --block-size CLI flag in dflash/benchmark.py.
  • Always validate configuration changes using the built-in benchmark utility to measure actual throughput and acceptance rates on your specific hardware.

Frequently Asked Questions

What happens if I set the block size too high?

If you configure a block size larger than the model's native capability, the DFlash code automatically clamps the value to draft.config.block_size as seen in dflash/model_mlx.py line 59. However, even with automatic capping, excessively large values (e.g., 32+) often result in acceptance rates below 30%, causing frequent verification failures that slow down overall generation compared to smaller blocks.

Can I change the block size without modifying the model files?

Yes. You can pass the block_size parameter directly to the generation functions without touching config.json. In the MLX backend, stream_generate() accepts this argument (lines 59-60), while the PyTorch backend supports it in spec_generate() which passes through to dflash_generate() (lines 76-77). This allows runtime experimentation without permanent configuration changes.

How do I know if my block size is optimal?

Use the benchmark utility provided in dflash/benchmark.py with the --block-size flag to measure throughput and acceptance statistics. The optimal configuration shows peak tok/s with acceptance rates above 60%. If acceptance drops below 30% or throughput plateaus, reduce the block size to the previous value that maintained high acceptance.

Does block size affect memory usage?

Yes, larger block sizes increase memory pressure because the draft model must maintain KV cache states for more speculative tokens simultaneously. For very long generation sequences, combine your chosen block size with --draft-sliding-window-size to limit cache growth, or reduce the block size if you encounter out-of-memory errors during inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →