# How to Debug DFlash Integration with vLLM: A Complete Troubleshooting Guide

> Debug DFlash integration with vLLM using LOGURU_LEVEL DEBUG and inspect the _send_vllm payload in dflash/benchmark.py. Resolve connection and compatibility errors effectively.

- Repository: [Z Lab/dflash](https://github.com/z-lab/dflash)
- Tags: how-to-guide
- Published: 2026-04-17

---

**Enable `LOGURU_LEVEL=DEBUG` and inspect the `_send_vllm` payload in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) to surface connection, compatibility, or token-mismatch errors when integrating DFlash with vLLM.**

Debugging DFlash integration with vLLM requires tracing data flow from the draft model through HTTP requests to the target server. The `z-lab/dflash` repository provides a dedicated vLLM server mode in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) that handles speculative decoding orchestration, but mismatches in tokenization, configuration, or network parameters often cause silent failures or zero acceptance rates.

## Verify the vLLM Server Setup

Before invoking DFlash, confirm the target model is reachable and responding with valid JSON. The integration expects a standard vLLM OpenAI-compatible endpoint.

Launch the server with the specific model and port:

```bash

# Example: launch Qwen3.5-27B with vLLM nightly

uv pip install -e ".[vllm]"
uv pip install -U vllm --torch-backend=auto --extra-index-url https://wheels.vllm.ai/nightly
vllm serve Qwen/Qwen3.5-27B --port 8000

```

Verify connectivity by querying the models endpoint:

```bash
curl http://127.0.0.1:8000/v1/models

```

If this returns an HTML error page or connection timeout, resolve the vLLM server configuration (CUDA version, torch build, or model compatibility) before proceeding with DFlash debugging.

## Run DFlash with the vLLM Backend

Once the server is stable, execute the benchmark with the `--backend vllm` flag to route requests through the HTTP integration path.

```bash
python -m dflash.benchmark \
  --backend vllm \
  --model Qwen/Qwen3.5-27B \
  --draft-model z-lab/Qwen3.5-27B-DFlash \
  --dataset gsm8k \
  --base-url http://127.0.0.1:8000 \
  --num-prompts 16 \
  --concurrency 4 \
  --max-new-tokens 256 \
  --temperature 0.0

```

### Required Command-Line Arguments

When debugging DFlash integration with vLLM, these arguments control the critical connection parameters:

- **`--backend vllm`**: Activates the `_run_server` code path in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) instead of local transformers inference.
- **`--base-url`**: Must match the vLLM server URL (default in benchmark is `http://127.0.0.1:30000`, but vLLM defaults to `8000`).
- **`--draft-model`**: Path to the DFlash checkpoint; required for any non-transformers backend.
- **`--block-size`**: Optional override for the draft model's native block size; omit to use the value from `config.dflash_config`.

If the benchmark crashes or returns empty results, proceed to inspect the HTTP layer and configuration compatibility.

## Inspect HTTP Requests and Responses

The `_send_vllm` function in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) (lines 99-120) constructs the JSON payload and posts to the vLLM Chat Completions endpoint. Debugging here isolates network and API contract issues.

### Debugging the _send_vllm Function

Insert diagnostic prints inside [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) to expose the exact data exchange:

```python
def _send_vllm(
    base_url: str,
    text: str,
    *,
    model: str,
    max_new_tokens: int,
    temperature: float,
    top_p: float,
    top_k: int,
    timeout_s: int,
    enable_thinking: bool = False,
) -> dict:
    body: dict = {
        "model": model,
        "messages": [{"role": "user", "content": text}],
        "max_tokens": max_new_tokens,
        "temperature": temperature,
        "top_p": top_p,
        "top_k": top_k,
        "chat_template_kwargs": {"enable_thinking": enable_thinking},
    }
    
    # Debug: print the request payload

    import json
    print("VLLM request:", json.dumps(body, indent=2))
    
    resp = requests.post(
        base_url + "/v1/chat/completions",
        json=body,
        timeout=timeout_s,
    )
    resp.raise_for_status()
    
    # Debug: print raw response before parsing

    raw = resp.text
    print("VLLM response:", raw)
    
    return resp.json()

```

Validate these three critical fields in the response JSON:

1. **`choices[0].message.content`**: Contains the generated text.
2. **`usage.completion_tokens`**: Required for throughput calculations.
3. **`usage.prompt_tokens`**: Validates the draft-to-target tokenization alignment.

If `choices` is missing or contains HTML, the vLLM server rejected the request (check model name compatibility or server logs).

## Validate Draft-Target Compatibility

DFlash requires architectural alignment between the draft and target models. The `DFlashDraftModel` class in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) reads configuration values that must match the target's expectations.

### Check block_size and mask_token_id

In [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) (lines 19-22), the draft model initializes critical parameters from its configuration:

```python
self.block_size = config.block_size
self.mask_token_id = self.config.dflash_config.get("mask_token_id", None)

```

Debug checklist for configuration mismatches:

| Symptom | Root Cause | Resolution |
|---------|------------|------------|
| `RuntimeError: CUDA out of memory` | `block_size` too large for GPU memory (e.g., 64 on a 12GB card). | Override with `--block-size 16` or upgrade GPU. |
| `IndexError: mask_token_id is None` | Draft config missing `mask_token_id` required for masked padding. | Use a draft defining `mask_token_id` or manually specify via future CLI flag. |
| `dflash_generate` returns only the input prompt | `target_layer_ids` mismatch causes `target_hidden` extraction failure. | Ensure both models share the same architecture (both Qwen3 or both LLaMA-3.1). |

### Resolve Acceptance Length Issues

The acceptance length calculation in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) (lines 134-138) determines how many draft tokens survive verification:

```python
acceptance_length = (block_output_ids[:, 1:] == posterior[:, :-1]).cumprod(dim=1).sum(dim=1)[0].item()

```

If acceptance length is consistently **1** (only the first token is kept), investigate these three common causes:

1. **Vocabulary mismatch**: Compare `tokenizer.get_vocab()` between draft and target. Identical tokenizers are required for token alignment.
2. **Thinking mode incompatibility**: Lines 303-306 in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) validate that `enable_thinking` is disabled for certain draft models. If the draft was trained without thinking traces but `--enable-thinking` is set to true, token distributions diverge.
3. **Mask token contamination**: If `mask_token_id` is incorrectly set, padding tokens may appear in `block_output_ids`, causing premature rejection of valid draft tokens.

Manual verification of token streams:

```python
print("draft tokens:", block_output_ids[0, 1:acceptance_length+1].tolist())
print("target tokens:", posterior[0, :acceptance_length].tolist())

```

## Enable Detailed Logging for Deep Debugging

Activate `loguru` debug output to trace the benchmark execution flow:

```bash
export LOGURU_LEVEL=DEBUG
python -m dflash.benchmark --backend vllm ...

```

This exposes per-prompt timing, acceptance lengths, and exception traces inside `_run_server`.

For PyTorch-level diagnostics (CUDA kernel failures, memory access errors):

```bash
export TORCH_CPP_LOG_LEVEL=INFO
export CUDA_LAUNCH_BLOCKING=1

```

`CUDA_LAUNCH_BLOCKING=1` forces synchronous CUDA execution, ensuring errors surface at the exact line of failure rather than asynchronously.

## Common Pitfalls and Quick Fixes

| Symptom | Quick Fix |
|---------|-----------|
| `requests.exceptions.ConnectTimeout` | Verify `--base-url` matches the vLLM server port (vLLM defaults to `8000`, DFlash benchmark defaults to `30000`). |
| `KeyError: 'choices'` in response | The vLLM server returned an HTML error page. Check vLLM logs for model loading failures or incompatible API arguments. |
| `CUDA kernel error: device-side assert triggered` | Set `CUDA_LAUNCH_BLOCKING=1` and inspect tensor dimensions; usually indicates `block_size` mismatch or out-of-bounds token IDs. |
| Acceptance length always 1 | Disable `--enable-thinking` for drafts not trained with thinking traces (validation at lines 303-306 of [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py)). |
| `Model not found` loading draft | Use exact HuggingFace identifier (`z-lab/Qwen3.5-27B-DFlash`). Run `huggingface-cli lfs pull` if files are incomplete. |

## End-to-End Spot-Check Example

Isolate a single prompt to debug without benchmark overhead:

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
from dflash.model import DFlashDraftModel, dflash_generate

# 1. Load target and draft models

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-27B")
target = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3.5-27B",
    torch_dtype="auto",
    device_map="auto",
).eval()

draft = DFlashDraftModel.from_pretrained(
    "z-lab/Qwen3.5-27B-DFlash",
    torch_dtype="auto",
    device_map="auto",
).eval()

# 2. Prepare prompt

prompt = "Explain why the sky is blue."
input_ids = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    add_generation_prompt=True,
    tokenize=True,
    return_tensors="pt",
).to(target.device)

# 3. Execute speculative decode

result = dflash_generate(
    draft,
    target=target,
    input_ids=input_ids,
    max_new_tokens=128,
    stop_token_ids=[tokenizer.eos_token_id],
    temperature=0.0,
    block_size=16,
    return_stats=True,
)

print("Generated:", tokenizer.decode(result.output_ids[0], skip_special_tokens=True))
print("Acceptance lengths per block:", result.acceptance_lengths)

```

This script replicates the benchmark's core logic for a single prompt, enabling interactive debugging with `pdb` or VS Code breakpoints to inspect `block_output_ids` and `posterior` tensors directly.

## Summary

- **Verify connectivity** to the vLLM server at the correct `--base-url` before running DFlash.
- **Enable debug logging** via `LOGURU_LEVEL=DEBUG` and instrument `_send_vllm` in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) to inspect raw HTTP payloads.
- **Align configurations** between draft and target models, specifically `block_size` and `mask_token_id` in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py).
- **Disable thinking mode** (`--enable-thinking=false`) for drafts not trained with thinking traces to avoid token distribution mismatches.
- **Use `CUDA_LAUNCH_BLOCKING=1`** to capture synchronous CUDA errors when encountering device-side asserts.

## Frequently Asked Questions

### Why does DFlash return only the original prompt when using vLLM mode?

This occurs when `target_layer_ids` extraction fails due to architectural mismatches between the draft and target models. Ensure both models share the same architecture family (e.g., both Qwen3 or both LLaMA-3.1) so that `target_hidden` tensors can be correctly extracted in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py).

### How do I fix "CUDA out of memory" errors during speculative decoding?

The draft model's `block_size` likely exceeds your GPU memory capacity. In [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py), the `block_size` is loaded from `config.block_size`. Override this with the `--block-size` CLI argument (e.g., `--block-size 16`) to reduce memory consumption, or upgrade to a GPU with larger VRAM.

### Why is the acceptance length always 1, indicating all draft tokens are rejected?

Consistent rejection of draft tokens stems from three primary causes: vocabulary mismatches between draft and target tokenizers, enabling thinking mode (`--enable-thinking`) on drafts not trained with thinking traces (validated at lines 303-306 of [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py)), or incorrect `mask_token_id` values causing padding tokens to contaminate the comparison in [`dflash/model.py`](https://github.com/z-lab/dflash/blob/main/dflash/model.py) lines 134-138.

### What should I check when receiving `requests.exceptions.ConnectTimeout`?

Verify that the `--base-url` argument matches the actual vLLM server port. The vLLM default is `http://127.0.0.1:8000`, while the DFlash benchmark defaults to `30000`. Ensure the server is running and the port is not blocked by firewall rules before launching the benchmark.