# How DFlash Handles Thinking Tokens in Qwen3 Models

> Discover how DFlash manages thinking tokens in Qwen3 models. Learn about its tokenizer API integration and guard-rails for model compatibility.

- Repository: [Z Lab/dflash](https://github.com/z-lab/dflash)
- Tags: how-to-guide
- Published: 2026-04-17

---

**DFlash delegates thinking token handling entirely to the tokenizer's `apply_chat_template` API and enforces guard-rails to block unsupported models like Qwen3-4B and Qwen3-8B.**

DFlash is a speculative decoding framework that accelerates inference for large language models. When working with Qwen3 models that support "thinking" capabilities—where the model generates intermediate reasoning tokens before the final answer—DFlash does not implement custom logic to process these tokens. Instead, it relies on the underlying tokenizer and model-specific training configurations to manage thinking tokens correctly.

## Tokenizer Delegation: The Core Mechanism

DFlash treats thinking tokens as a tokenizer-level concern rather than a framework-level feature. The framework passes the `enable_thinking` flag directly to the tokenizer's chat template application.

### How `apply_chat_template` Handles Thinking Tokens

When the `--enable-thinking` flag is set via command line, DFlash invokes the tokenizer in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) (lines 103-108):

```python
input_ids = torch.tensor(
    tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True,
        enable_thinking=True,  # Tokenizer inserts special thinking tokens

    ),
).to(device)

```

The tokenizer inserts special thinking tokens (such as `<|thinking|>`) into the prompt based on the model's chat template definition. These tokens are only meaningful if the draft model was trained with thinking traces; DFlash itself does not interpret or modify these tokens.

## Model-Specific Guard-Rails for Qwen3

Not all Qwen3 models support thinking tokens. DFlash implements explicit assertions to prevent users from enabling thinking mode on incompatible draft models, which would produce suboptimal results.

### Why Qwen3-4B and Qwen3-8B Block Thinking Mode

The Qwen3-4B and Qwen3-8B draft checkpoints were not trained with thinking traces. In [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) (lines 503-506), DFlash aborts execution if thinking is enabled for these specific models:

```python
assert not (args.enable_thinking and any(
    x in args.model.lower() for x in ["qwen3-4b", "qwen3-8b"]
)), (
    "DFlash draft models for Qwen3-4B and Qwen3-8B were not trained with thinking traces. "
    "Using --enable-thinking will lead to suboptimal performance."
)

```

This guard-rail ensures that users cannot accidentally enable thinking mode on models that lack the necessary training data to generate meaningful reasoning tokens.

## Generation Loop: No Special Processing

Once tokenization is complete, DFlash's speculative decoding engine treats thinking tokens identically to standard tokens. The framework does not implement special handling for reasoning tokens during the generation phase.

### How `dflash_generate` Treats All Tokens Equally

The core speculative decoding routine, `dflash_generate` in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py), processes token IDs as opaque integers. Whether a token represents ordinary text, a thinking delimiter, or a special control token is irrelevant to the block-diffusion algorithm. The model simply treats each token as the next candidate to be drafted or verified by the target model.

This architecture keeps DFlash agnostic to the semantic meaning of tokens, allowing it to support new token types—including future thinking formats—without code changes.

## Supported Configurations and Usage Examples

Thinking tokens are supported when using Qwen3.5-based draft models, which were trained with thinking traces. DFlash provides different usage patterns depending on the backend (MLX or Transformers).

### Enabling Thinking with Qwen3.5 Models

For MLX-based inference with Qwen3.5-4B, enable thinking through the tokenizer in [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py):

```python
from dflash.model_mlx import load, load_draft, stream_generate

model, tokenizer = load("Qwen/Qwen3.5-4B")
draft = load_draft("z-lab/Qwen3.5-4B-DFlash")

messages = [{"role": "user", "content": "Explain the significance of the Riemann hypothesis."}]
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True,  # Inserts thinking tokens for Qwen3.5

)

for response in stream_generate(model, draft, tokenizer, prompt,
                                block_size=16, max_tokens=1024, temperature=0.6):
    print(response.text, end="", flush=True)

```

### Disabling Thinking for Legacy Qwen3 Drafts

When using Qwen3-8B (which lacks thinking training data), explicitly disable thinking in your Transformers-based script:

```python
from transformers import AutoModel, AutoModelForCausalLM, AutoTokenizer
import torch

draft = AutoModel.from_pretrained(
    "z-lab/Qwen3-8B-DFlash-b16", trust_remote_code=True, device_map="cuda:0"
)
target = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3-8B", device_map="cuda:0"
)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")

messages = [{"role": "user", "content": "How many primes are less than 100?"}]
input_ids = tokenizer.apply_chat_template(
    messages,
    return_tensors="pt",
    add_generation_prompt=True,
    enable_thinking=False,  # Required for Qwen3-8B compatibility

).to(draft.device)

output = draft.spec_generate(
    input_ids=input_ids,
    max_new_tokens=512,
    temperature=0.0,
    target=target,
    stop_token_ids=[tokenizer.eos_token_id],
)
print(tokenizer.decode(output[0], skip_special_tokens=False))

```

## Summary

- **DFlash delegates thinking token handling to the tokenizer's `apply_chat_template` API**, passing the `enable_thinking` flag directly to the model's chat template.
- **Guard-rails prevent misuse on incompatible models**: DFlash explicitly blocks thinking mode for Qwen3-4B and Qwen3-8B drafts in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) because these models lack thinking trace training data.
- **No special generation logic**: The `dflash_generate` function in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) treats thinking tokens as standard token IDs during speculative decoding, maintaining framework agnosticism.
- **Supported configurations**: Only Qwen3.5-based draft models (e.g., Qwen3.5-4B) support thinking tokens, as demonstrated in [`dflash/model_mlx.py`](https://github.com/z-lab/dflash/blob/main/dflash/model_mlx.py) and the project README.

## Frequently Asked Questions

### Does DFlash modify thinking tokens during speculative decoding?

No. DFlash does not implement any special logic to detect, filter, or modify thinking tokens during the generation loop. The `dflash_generate` function in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) processes all token IDs—including thinking tokens—as opaque integers in its block-diffusion algorithm.

### Why can't I use thinking mode with Qwen3-4B?

The Qwen3-4B and Qwen3-8B draft checkpoints were not trained with thinking traces. DFlash enforces this limitation through an assertion in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) (lines 503-506) that aborts execution if `--enable-thinking` is used with these specific model variants, preventing suboptimal performance.

### Which DFlash draft models support thinking tokens?

Draft models based on Qwen3.5 support thinking tokens. Specifically, the `z-lab/Qwen3.5-4B-DFlash` and similar Qwen3.5 variants were trained with thinking traces and can safely use the `enable_thinking=True` flag in the tokenizer's `apply_chat_template` method.

### How do I enable thinking mode in my DFlash inference script?

Pass `enable_thinking=True` to the tokenizer's `apply_chat_template` method when preparing your prompt. For example: `tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=True)`. Ensure you are using a Qwen3.5-based draft model, as Qwen3-4B/8B drafts do not support this feature.