How DFlash Handles Thinking Tokens in Qwen3 Models

DFlash delegates thinking token handling entirely to the tokenizer's apply_chat_template API and enforces guard-rails to block unsupported models like Qwen3-4B and Qwen3-8B.

DFlash is a speculative decoding framework that accelerates inference for large language models. When working with Qwen3 models that support "thinking" capabilities—where the model generates intermediate reasoning tokens before the final answer—DFlash does not implement custom logic to process these tokens. Instead, it relies on the underlying tokenizer and model-specific training configurations to manage thinking tokens correctly.

Tokenizer Delegation: The Core Mechanism

DFlash treats thinking tokens as a tokenizer-level concern rather than a framework-level feature. The framework passes the enable_thinking flag directly to the tokenizer's chat template application.

How apply_chat_template Handles Thinking Tokens

When the --enable-thinking flag is set via command line, DFlash invokes the tokenizer in dflash/benchmark.py (lines 103-108):

input_ids = torch.tensor(
    tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True,
        enable_thinking=True,  # Tokenizer inserts special thinking tokens

    ),
).to(device)

The tokenizer inserts special thinking tokens (such as <|thinking|>) into the prompt based on the model's chat template definition. These tokens are only meaningful if the draft model was trained with thinking traces; DFlash itself does not interpret or modify these tokens.

Model-Specific Guard-Rails for Qwen3

Not all Qwen3 models support thinking tokens. DFlash implements explicit assertions to prevent users from enabling thinking mode on incompatible draft models, which would produce suboptimal results.

Why Qwen3-4B and Qwen3-8B Block Thinking Mode

The Qwen3-4B and Qwen3-8B draft checkpoints were not trained with thinking traces. In dflash/benchmark.py (lines 503-506), DFlash aborts execution if thinking is enabled for these specific models:

assert not (args.enable_thinking and any(
    x in args.model.lower() for x in ["qwen3-4b", "qwen3-8b"]
)), (
    "DFlash draft models for Qwen3-4B and Qwen3-8B were not trained with thinking traces. "
    "Using --enable-thinking will lead to suboptimal performance."
)

This guard-rail ensures that users cannot accidentally enable thinking mode on models that lack the necessary training data to generate meaningful reasoning tokens.

Generation Loop: No Special Processing

Once tokenization is complete, DFlash's speculative decoding engine treats thinking tokens identically to standard tokens. The framework does not implement special handling for reasoning tokens during the generation phase.

How dflash_generate Treats All Tokens Equally

The core speculative decoding routine, dflash_generate in dflash/model.py, processes token IDs as opaque integers. Whether a token represents ordinary text, a thinking delimiter, or a special control token is irrelevant to the block-diffusion algorithm. The model simply treats each token as the next candidate to be drafted or verified by the target model.

This architecture keeps DFlash agnostic to the semantic meaning of tokens, allowing it to support new token types—including future thinking formats—without code changes.

Supported Configurations and Usage Examples

Thinking tokens are supported when using Qwen3.5-based draft models, which were trained with thinking traces. DFlash provides different usage patterns depending on the backend (MLX or Transformers).

Enabling Thinking with Qwen3.5 Models

For MLX-based inference with Qwen3.5-4B, enable thinking through the tokenizer in dflash/model_mlx.py:

from dflash.model_mlx import load, load_draft, stream_generate

model, tokenizer = load("Qwen/Qwen3.5-4B")
draft = load_draft("z-lab/Qwen3.5-4B-DFlash")

messages = [{"role": "user", "content": "Explain the significance of the Riemann hypothesis."}]
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True,  # Inserts thinking tokens for Qwen3.5

)

for response in stream_generate(model, draft, tokenizer, prompt,
                                block_size=16, max_tokens=1024, temperature=0.6):
    print(response.text, end="", flush=True)

Disabling Thinking for Legacy Qwen3 Drafts

When using Qwen3-8B (which lacks thinking training data), explicitly disable thinking in your Transformers-based script:

from transformers import AutoModel, AutoModelForCausalLM, AutoTokenizer
import torch

draft = AutoModel.from_pretrained(
    "z-lab/Qwen3-8B-DFlash-b16", trust_remote_code=True, device_map="cuda:0"
)
target = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3-8B", device_map="cuda:0"
)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")

messages = [{"role": "user", "content": "How many primes are less than 100?"}]
input_ids = tokenizer.apply_chat_template(
    messages,
    return_tensors="pt",
    add_generation_prompt=True,
    enable_thinking=False,  # Required for Qwen3-8B compatibility

).to(draft.device)

output = draft.spec_generate(
    input_ids=input_ids,
    max_new_tokens=512,
    temperature=0.0,
    target=target,
    stop_token_ids=[tokenizer.eos_token_id],
)
print(tokenizer.decode(output[0], skip_special_tokens=False))

Summary

  • DFlash delegates thinking token handling to the tokenizer's apply_chat_template API, passing the enable_thinking flag directly to the model's chat template.
  • Guard-rails prevent misuse on incompatible models: DFlash explicitly blocks thinking mode for Qwen3-4B and Qwen3-8B drafts in dflash/benchmark.py because these models lack thinking trace training data.
  • No special generation logic: The dflash_generate function in dflash/model.py treats thinking tokens as standard token IDs during speculative decoding, maintaining framework agnosticism.
  • Supported configurations: Only Qwen3.5-based draft models (e.g., Qwen3.5-4B) support thinking tokens, as demonstrated in dflash/model_mlx.py and the project README.

Frequently Asked Questions

Does DFlash modify thinking tokens during speculative decoding?

No. DFlash does not implement any special logic to detect, filter, or modify thinking tokens during the generation loop. The dflash_generate function in dflash/model.py processes all token IDs—including thinking tokens—as opaque integers in its block-diffusion algorithm.

Why can't I use thinking mode with Qwen3-4B?

The Qwen3-4B and Qwen3-8B draft checkpoints were not trained with thinking traces. DFlash enforces this limitation through an assertion in dflash/benchmark.py (lines 503-506) that aborts execution if --enable-thinking is used with these specific model variants, preventing suboptimal performance.

Which DFlash draft models support thinking tokens?

Draft models based on Qwen3.5 support thinking tokens. Specifically, the z-lab/Qwen3.5-4B-DFlash and similar Qwen3.5 variants were trained with thinking traces and can safely use the enable_thinking=True flag in the tokenizer's apply_chat_template method.

How do I enable thinking mode in my DFlash inference script?

Pass enable_thinking=True to the tokenizer's apply_chat_template method when preparing your prompt. For example: tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=True). Ensure you are using a Qwen3.5-based draft model, as Qwen3-4B/8B drafts do not support this feature.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →