# How to Configure the Block Size in DFlash for Optimal Performance

> Learn how to configure DFlash block size for optimal performance. Adjust token prediction to balance throughput and acceptance rate on your hardware.

- Repository: [Z Lab/dflash](https://github.com/z-lab/dflash)
- Tags: how-to-guide
- Published: 2026-04-17

---

**Block size in DFlash controls how many tokens the draft model predicts in each speculative step, typically ranging from 4 to 16 tokens, where values around 8–16 provide the best balance between throughput and acceptance rate on most hardware.**

Configuring the block size is essential when optimizing speculative decoding performance in DFlash, an open-source inference acceleration library. The block size determines how many tokens the draft model generates before the target model verifies them, directly impacting the trade-off between raw speed and draft acceptance rates. In the z-lab/dflash repository, this parameter is exposed through model configuration files, Python APIs, and CLI tools.

## What Is Block Size in DFlash?

Block size defines the number of tokens the draft model predicts in a single speculative "block" before the target model validates them. It governs the fundamental trade-off between **throughput** (larger blocks mean more tokens per draft) and **acceptance rate** (larger blocks are less likely to be fully accepted, potentially causing costly rollbacks).

According to the source code in [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py) (lines 40-44), the `block_size` is read from the draft model's [`config.json`](https://github.com/z-lab/dflash/blob/main/config.json) and stored in the `DFlashConfig` class. Similarly, in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) (lines 19-20), the PyTorch `DFlashDraftModel` initializes `self.block_size` from the configuration during model load.

## Locating Block Size in the Source Code

The block size parameter flows through three critical files in the repository:

- **[`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py) (lines 40-44 and 59-60)**: Defines `DFlashConfig.block_size` loaded from JSON and implements the fallback logic in `stream_generate()` where `block_size = draft.config.block_size` if not explicitly provided.

- **[`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) (lines 19-20 and 76-77)**: Stores the attribute in `DFlashDraftModel.block_size` and handles the fallback in `dflash_generate()` with the logic `block_size = model.block_size if block_size is None else block_size`.

- **[`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) (line 227)**: Exposes the `--block-size` CLI argument for performance testing.

## Finding the Optimal Block Size

Empirical testing across the z-lab/dflash codebase suggests that **most models achieve peak effective throughput with block sizes between 8 and 16 tokens**. The exact optimal value depends on model architecture, hardware capabilities, and prompt length.

- **Small blocks (1–4 tokens)** deliver the highest acceptance rates (nearly 100%) but provide minimal speculative acceleration since the draft model only predicts a few tokens per step.

- **Medium blocks (8–16 tokens)** strike the practical balance, typically yielding 2–4× speedups on GPU or Apple Silicon MLX hardware while maintaining respectable acceptance rates above 60-80%.

- **Large blocks (32+ tokens)** maximize parallel speculative computation but often suffer acceptance rates below 30%, causing frequent rollbacks that negate the performance gains.

## How to Configure the Block Size

You can set the block size through four methods, depending on whether you need a permanent configuration or a temporary override.

### Via the Draft Model Configuration

Edit the draft model's [`config.json`](https://github.com/z-lab/dflash/blob/main/config.json) to permanently set the default for all inference runs:

```json
{
  "block_size": 16
}

```

This value is automatically loaded into `DFlashConfig.block_size` when the model initializes, as implemented in [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py) lines 40-44.

### Via the Generation API

Pass an explicit `block_size` parameter to override the config default without modifying files.

**MLX Backend using `stream_generate`:**

```python
from dflash.model_mlx import stream_generate

for result in stream_generate(
    model, draft, tokenizer, prompt,
    block_size=12,  # Override config default

    max_tokens=512
):
    print(result.text, end="", flush=True)

```

The fallback logic at lines 59-60 of [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py) uses your provided value instead of `draft.config.block_size`.

**PyTorch Backend using `spec_generate`:**

```python
from dflash.model import DFlashDraftModel

draft = DFlashDraftModel.from_pretrained("z-lab/Qwen3.5-4B-DFlash")
output_ids = draft.spec_generate(
    target_model,
    input_ids,
    block_size=10,  # Custom block size

    max_new_tokens=128
)

```

This passes through to `dflash_generate` where lines 76-77 of [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) apply the override.

### Via Command-Line Benchmarks

Use the `--block-size` flag when running the benchmark utility to test different values:

```bash
python -m dflash.benchmark \
    --backend mlx \
    --model Qwen/Qwen3.5-4B \
    --draft-model z-lab/Qwen3.5-4B-DFlash \
    --block-size 14 \
    --dataset gsm8k \
    --max-samples 64

```

The argument parser at line 227 of [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) handles this input and reports per-block throughput and acceptance statistics.

### Via Environment Variable

For containerized deployments, set `DFLASH_BLOCK_SIZE` before launching the inference server. The server code reads this variable and forwards it to the underlying `stream_generate` function.

## Practical Workflow for Optimization

Follow this workflow to determine the ideal block size for your specific hardware and model combination:

1. **Run a benchmark sweep** across 4, 8, 12, and 16 tokens:

```bash
for bs in 4 8 12 16; do
  python -m dflash.benchmark \
      --backend mlx \
      --draft-model z-lab/Qwen3.5-4B-DFlash \
      --block-size $bs \
      --max-samples 32
done

```

2. **Analyze the output** for throughput (tok/s) and mean acceptance length. Throughput will rise with block size until acceptance rate drops precipitously.

3. **Lock the configuration** by updating [`config.json`](https://github.com/z-lab/dflash/blob/main/config.json) with the winning value or hardcoding it in your production inference script.

## Production Considerations

When deploying DFlash in production environments, keep these constraints in mind:

- **Do not exceed the native block size**: The code automatically caps values at `int(draft.config.block_size)`, so passing larger values silently reduces them to the model's native limit.

- **Enable sliding windows for long sequences**: When generating very long text, combine larger block sizes with `--draft-sliding-window-size` to prevent KV cache memory blow-up.

- **Monitor cache pressure**: The MLX backend calls `mx.clear_cache()` every 256 tokens in `stream_generate`. If you encounter out-of-memory errors, reduce the block size or enable sliding windows.

## Summary

- Block size in DFlash determines how many tokens the draft model generates speculatively before target verification, sourced from [`config.json`](https://github.com/z-lab/dflash/blob/main/config.json) via `DFlashConfig` in [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py).
- The optimal range is typically **8–16 tokens**, balancing the speed of speculative generation against the acceptance rate of the target model.
- Override the default via the `block_size` parameter in `stream_generate()` or `spec_generate()`, or use the `--block-size` CLI flag in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py).
- Always validate configuration changes using the built-in benchmark utility to measure actual throughput and acceptance rates on your specific hardware.

## Frequently Asked Questions

### What happens if I set the block size too high?

If you configure a block size larger than the model's native capability, the DFlash code automatically clamps the value to `draft.config.block_size` as seen in [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py) line 59. However, even with automatic capping, excessively large values (e.g., 32+) often result in acceptance rates below 30%, causing frequent verification failures that slow down overall generation compared to smaller blocks.

### Can I change the block size without modifying the model files?

Yes. You can pass the `block_size` parameter directly to the generation functions without touching [`config.json`](https://github.com/z-lab/dflash/blob/main/config.json). In the MLX backend, `stream_generate()` accepts this argument (lines 59-60), while the PyTorch backend supports it in `spec_generate()` which passes through to `dflash_generate()` (lines 76-77). This allows runtime experimentation without permanent configuration changes.

### How do I know if my block size is optimal?

Use the benchmark utility provided in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) with the `--block-size` flag to measure throughput and acceptance statistics. The optimal configuration shows peak tok/s with acceptance rates above 60%. If acceptance drops below 30% or throughput plateaus, reduce the block size to the previous value that maintained high acceptance.

### Does block size affect memory usage?

Yes, larger block sizes increase memory pressure because the draft model must maintain KV cache states for more speculative tokens simultaneously. For very long generation sequences, combine your chosen block size with `--draft-sliding-window-size` to limit cache growth, or reduce the block size if you encounter out-of-memory errors during inference.