# DFlash Performance Across Transformers, SGLang, and vLLM Backends: A Complete Benchmark Guide

> Benchmark DFlash performance: Discover how Transformers, SGLang, and vLLM backends compare in latency and throughput for single-GPU and concurrent workloads to optimize your inference.

- Repository: [Z Lab/dflash](https://github.com/z-lab/dflash)
- Tags: performance
- Published: 2026-04-17

---

**DFlash achieves the lowest latency with the Transformers backend for single-GPU inference, while vLLM and SGLang deliver higher throughput for concurrent workloads by leveraging remote server architectures.**

The `z-lab/dflash` repository implements a speculative decoding framework that supports three distinct execution backends. Each backend handles the **DFlash generation loop** (`dflash_generate`) differently, creating unique performance profiles for latency-sensitive versus throughput-intensive workloads. The benchmark driver in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) provides a unified interface to compare these backends using the `--backend` flag.

## Execution Architecture of Each Backend

### Transformers (Local PyTorch)

The **Transformers** backend runs as a pure-Python/PyTorch process that loads both the target and draft models directly into local GPU memory. In [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) lines 198-207, the `run_transformers` function invokes the core **DFlash generation loop** (`dflash_generate`) inside the same process.

Model loading occurs in lines 219-226, where `AutoModelForCausalLM` loads the target model and `DFlashDraftModel` initializes the draft model. Performance depends heavily on the attention implementation selected in lines 185-194, where the code configures `attn_implementation` to use Flash Attention 2 or SDPA. No network latency is involved, but the entire model must fit within the local GPU memory.

### SGLang (Remote HTTP)

The **SGLang** backend delegates inference to a remote HTTP server running the SGLang inference engine. In [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) lines 271-285, the `_send_sglang` function transmits each prompt via a POST request to `base_url/generate`.

The server performs the *entire* DFlash pipeline—including draft generation and token verification—internally, returning a JSON payload with the completed sequence. Performance is constrained by the server's inference engine optimization and network round-trip time, typically adding 1 ms to tens of milliseconds of overhead per request. The local process only handles request orchestration, minimizing CPU overhead.

### vLLM (Remote HTTP)

The **vLLM** backend similarly uses a remote HTTP architecture, but targets the vLLM inference engine. In [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) lines 299-317, the `_send_vllm` function sends requests to `base_url/v1/chat/completions` with the model name, prompt, and sampling parameters.

The vLLM server runs the DFlash algorithm internally and returns the completion. Like SGLang, performance depends on server-side optimization and network latency. However, vLLM is specifically optimized for high-throughput token generation through its PagedAttention mechanism, often achieving superior throughput for large batch sizes compared to other backends.

## Performance Characteristics and Trade-offs

### Latency Analysis

For single-request latency, the **Transformers** backend provides the lowest possible latency because it eliminates network round-trips and HTTP serialization overhead. The generation loop runs directly on the local GPU, with latency determined solely by CUDA kernel execution time and memory bandwidth.

**SGLang** and **vLLM** introduce additional latency from HTTP request handling and network transmission. For very short prompts, this network overhead can dominate the total execution time, making these backends less suitable for latency-critical applications with low concurrency.

### Throughput Under Concurrency

When running with high concurrency (`--concurrency > 1`), the performance hierarchy shifts. **vLLM** typically achieves the highest throughput due to its continuous batching and PagedAttention implementation, which efficiently handles multiple concurrent requests with minimal memory overhead.

**SGLang** delivers comparable throughput when the server is configured for batch inference, though exact performance depends on the specific SGLang version and configuration. Both remote backends amortize GPU costs across multiple requests, whereas the local **Transformers** backend is limited by single-GPU batch size constraints and the overhead of the Python generation loop.

### Scalability Considerations

**SGLang** and **vLLM** offer superior horizontal scalability. You can scale across multiple GPUs or machines without modifying client code by simply updating the `--base-url` parameter to point to a larger deployment. The benchmark driver handles this through the `_run_server` function (lines 381-387), which selects the appropriate request function based on the backend flag.

The **Transformers** backend requires explicit multi-GPU coordination using `torchrun` or similar distributed launchers. The benchmark code supports this through distributed helpers (`_dist_*`), but configuration is more complex compared to the remote backends.

## Running Benchmarks

The benchmark driver in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) provides a unified CLI for comparing all three backends. The `--backend` flag (parsed in lines 81-84) selects the execution mode.

To benchmark the **Transformers** backend with local GPU execution:

```bash
python -m dflash.benchmark \
  --backend transformers \
  --model Qwen/Qwen3-8B-Instruct \
  --draft-model Qwen/Qwen3-8B-Draft \
  --dataset gsm8k \
  --max-new-tokens 512 \
  --block-size 16 \
  --num-prompts 200 \
  --concurrency 4

```

To benchmark against an **SGLang** server:

```bash
python -m dflash.benchmark \
  --backend sglang \
  --model Qwen/Qwen3-8B-Instruct \
  --draft-model Qwen/Qwen3-8B-Draft \
  --base-url http://sglang-host:30000 \
  --dataset gsm8k \
  --max-new-tokens 512 \
  --num-prompts 200 \
  --concurrency 4

```

To benchmark against a **vLLM** server:

```bash
python -m dflash.benchmark \
  --backend vllm \
  --model Qwen/Qwen3-8B-Instruct \
  --base-url http://vllm-host:8000 \
  --dataset gsm8k \
  --max-new-tokens 512 \
  --num-prompts 200 \
  --concurrency 4

```

After completion, the benchmark prints a decode summary via `print_decode_summary` (lines 20-27), showing metrics including acceptance length and decoding speedup for the Transformers backend.

## Summary

- **DFlash performance** varies significantly by backend selection, with each option optimized for different deployment scenarios.
- The **Transformers** backend delivers the lowest latency for single-GPU inference by running the `dflash_generate` loop locally without network overhead.
- **SGLang** and **vLLM** backends trade latency for scalability, using HTTP servers to handle the DFlash algorithm remotely.
- **vLLM** typically achieves the highest throughput under high concurrency due to its optimized batching and memory management.
- The benchmark script in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) provides unified testing across all backends using the `--backend` flag and shared metrics collection.

## Frequently Asked Questions

### Which DFlash backend offers the lowest latency for single requests?

The **Transformers** backend provides the lowest latency for single requests because it runs the **DFlash generation loop** (`dflash_generate`) directly in the local process without network round-trips. In [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) lines 198-207, the `run_transformers` function loads the model locally and executes generation immediately, eliminating the HTTP overhead that adds 1 ms to tens of milliseconds in the server-based backends.

### How does vLLM achieve higher throughput than the Transformers backend?

**vLLM** achieves higher throughput through its **PagedAttention** mechanism and continuous batching implementation, which efficiently handles multiple concurrent requests with minimal memory overhead. While the **Transformers** backend in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) is limited by single-GPU batch size constraints and Python loop overhead, vLLM's server architecture amortizes GPU computation across many requests. This makes vLLM superior for high-concurrency scenarios specified with `--concurrency` values greater than 1.

### Can I switch between SGLang and vLLM without changing my benchmark configuration?

Yes, you can switch between **SGLang** and **vLLM** by changing only the `--backend` and `--base-url` parameters. The benchmark driver in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) lines 381-387 uses the `_run_server` function to select the appropriate request handler (`_send_sglang` or `_send_vllm`) based on the backend flag. Both server backends use HTTP POST requests and return JSON payloads, allowing the benchmark to collect consistent metrics (latency, tokens per second) regardless of which server is handling the DFlash algorithm execution.

### What hardware requirements are needed for the Transformers backend compared to the server backends?

The **Transformers** backend requires the entire target model and draft model to fit within the local GPU memory, as evidenced by the model loading code in [`dflash/benchmark.py`](https://github.com/z-lab/dflash/blob/main/dflash/benchmark.py) lines 219-226. This backend uses `AutoModelForCausalLM` and `DFlashDraftModel` to load both models locally, requiring high-memory GPUs like A100 or H100 for large models. In contrast, **SGLang** and **vLLM** backends only require network connectivity to the server and minimal local CPU resources, as the heavy GPU computation happens remotely on the server infrastructure, which can be scaled independently across multiple GPUs or nodes.