DFlash Performance Across Transformers, SGLang, and vLLM Backends: A Complete Benchmark Guide

DFlash achieves the lowest latency with the Transformers backend for single-GPU inference, while vLLM and SGLang deliver higher throughput for concurrent workloads by leveraging remote server architectures.

The z-lab/dflash repository implements a speculative decoding framework that supports three distinct execution backends. Each backend handles the DFlash generation loop (dflash_generate) differently, creating unique performance profiles for latency-sensitive versus throughput-intensive workloads. The benchmark driver in dflash/benchmark.py provides a unified interface to compare these backends using the --backend flag.

Execution Architecture of Each Backend

Transformers (Local PyTorch)

The Transformers backend runs as a pure-Python/PyTorch process that loads both the target and draft models directly into local GPU memory. In dflash/benchmark.py lines 198-207, the run_transformers function invokes the core DFlash generation loop (dflash_generate) inside the same process.

Model loading occurs in lines 219-226, where AutoModelForCausalLM loads the target model and DFlashDraftModel initializes the draft model. Performance depends heavily on the attention implementation selected in lines 185-194, where the code configures attn_implementation to use Flash Attention 2 or SDPA. No network latency is involved, but the entire model must fit within the local GPU memory.

SGLang (Remote HTTP)

The SGLang backend delegates inference to a remote HTTP server running the SGLang inference engine. In dflash/benchmark.py lines 271-285, the _send_sglang function transmits each prompt via a POST request to base_url/generate.

The server performs the entire DFlash pipeline—including draft generation and token verification—internally, returning a JSON payload with the completed sequence. Performance is constrained by the server's inference engine optimization and network round-trip time, typically adding 1 ms to tens of milliseconds of overhead per request. The local process only handles request orchestration, minimizing CPU overhead.

vLLM (Remote HTTP)

The vLLM backend similarly uses a remote HTTP architecture, but targets the vLLM inference engine. In dflash/benchmark.py lines 299-317, the _send_vllm function sends requests to base_url/v1/chat/completions with the model name, prompt, and sampling parameters.

The vLLM server runs the DFlash algorithm internally and returns the completion. Like SGLang, performance depends on server-side optimization and network latency. However, vLLM is specifically optimized for high-throughput token generation through its PagedAttention mechanism, often achieving superior throughput for large batch sizes compared to other backends.

Performance Characteristics and Trade-offs

Latency Analysis

For single-request latency, the Transformers backend provides the lowest possible latency because it eliminates network round-trips and HTTP serialization overhead. The generation loop runs directly on the local GPU, with latency determined solely by CUDA kernel execution time and memory bandwidth.

SGLang and vLLM introduce additional latency from HTTP request handling and network transmission. For very short prompts, this network overhead can dominate the total execution time, making these backends less suitable for latency-critical applications with low concurrency.

Throughput Under Concurrency

When running with high concurrency (--concurrency > 1), the performance hierarchy shifts. vLLM typically achieves the highest throughput due to its continuous batching and PagedAttention implementation, which efficiently handles multiple concurrent requests with minimal memory overhead.

SGLang delivers comparable throughput when the server is configured for batch inference, though exact performance depends on the specific SGLang version and configuration. Both remote backends amortize GPU costs across multiple requests, whereas the local Transformers backend is limited by single-GPU batch size constraints and the overhead of the Python generation loop.

Scalability Considerations

SGLang and vLLM offer superior horizontal scalability. You can scale across multiple GPUs or machines without modifying client code by simply updating the --base-url parameter to point to a larger deployment. The benchmark driver handles this through the _run_server function (lines 381-387), which selects the appropriate request function based on the backend flag.

The Transformers backend requires explicit multi-GPU coordination using torchrun or similar distributed launchers. The benchmark code supports this through distributed helpers (_dist_*), but configuration is more complex compared to the remote backends.

Running Benchmarks

The benchmark driver in dflash/benchmark.py provides a unified CLI for comparing all three backends. The --backend flag (parsed in lines 81-84) selects the execution mode.

To benchmark the Transformers backend with local GPU execution:

python -m dflash.benchmark \
  --backend transformers \
  --model Qwen/Qwen3-8B-Instruct \
  --draft-model Qwen/Qwen3-8B-Draft \
  --dataset gsm8k \
  --max-new-tokens 512 \
  --block-size 16 \
  --num-prompts 200 \
  --concurrency 4

To benchmark against an SGLang server:

python -m dflash.benchmark \
  --backend sglang \
  --model Qwen/Qwen3-8B-Instruct \
  --draft-model Qwen/Qwen3-8B-Draft \
  --base-url http://sglang-host:30000 \
  --dataset gsm8k \
  --max-new-tokens 512 \
  --num-prompts 200 \
  --concurrency 4

To benchmark against a vLLM server:

python -m dflash.benchmark \
  --backend vllm \
  --model Qwen/Qwen3-8B-Instruct \
  --base-url http://vllm-host:8000 \
  --dataset gsm8k \
  --max-new-tokens 512 \
  --num-prompts 200 \
  --concurrency 4

After completion, the benchmark prints a decode summary via print_decode_summary (lines 20-27), showing metrics including acceptance length and decoding speedup for the Transformers backend.

Summary

  • DFlash performance varies significantly by backend selection, with each option optimized for different deployment scenarios.
  • The Transformers backend delivers the lowest latency for single-GPU inference by running the dflash_generate loop locally without network overhead.
  • SGLang and vLLM backends trade latency for scalability, using HTTP servers to handle the DFlash algorithm remotely.
  • vLLM typically achieves the highest throughput under high concurrency due to its optimized batching and memory management.
  • The benchmark script in dflash/benchmark.py provides unified testing across all backends using the --backend flag and shared metrics collection.

Frequently Asked Questions

Which DFlash backend offers the lowest latency for single requests?

The Transformers backend provides the lowest latency for single requests because it runs the DFlash generation loop (dflash_generate) directly in the local process without network round-trips. In dflash/benchmark.py lines 198-207, the run_transformers function loads the model locally and executes generation immediately, eliminating the HTTP overhead that adds 1 ms to tens of milliseconds in the server-based backends.

How does vLLM achieve higher throughput than the Transformers backend?

vLLM achieves higher throughput through its PagedAttention mechanism and continuous batching implementation, which efficiently handles multiple concurrent requests with minimal memory overhead. While the Transformers backend in dflash/benchmark.py is limited by single-GPU batch size constraints and Python loop overhead, vLLM's server architecture amortizes GPU computation across many requests. This makes vLLM superior for high-concurrency scenarios specified with --concurrency values greater than 1.

Can I switch between SGLang and vLLM without changing my benchmark configuration?

Yes, you can switch between SGLang and vLLM by changing only the --backend and --base-url parameters. The benchmark driver in dflash/benchmark.py lines 381-387 uses the _run_server function to select the appropriate request handler (_send_sglang or _send_vllm) based on the backend flag. Both server backends use HTTP POST requests and return JSON payloads, allowing the benchmark to collect consistent metrics (latency, tokens per second) regardless of which server is handling the DFlash algorithm execution.

What hardware requirements are needed for the Transformers backend compared to the server backends?

The Transformers backend requires the entire target model and draft model to fit within the local GPU memory, as evidenced by the model loading code in dflash/benchmark.py lines 219-226. This backend uses AutoModelForCausalLM and DFlashDraftModel to load both models locally, requiring high-memory GPUs like A100 or H100 for large models. In contrast, SGLang and vLLM backends only require network connectivity to the server and minimal local CPU resources, as the heavy GPU computation happens remotely on the server infrastructure, which can be scaled independently across multiple GPUs or nodes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →