How Speculative Decoding in vLLM Accelerates Token Generation

Speculative decoding in vLLM uses a lightweight draft model to predict multiple tokens ahead, validates them with a single forward pass of the full target model, and accepts valid tokens via rejection sampling to achieve up to N× speedup while maintaining the exact output distribution.

Speculative decoding is a look-ahead optimization technique that dramatically reduces latency in large language model inference. In the vllm-project/vllm repository, this feature is implemented as a configurable pipeline that generates candidate token sequences using cheap draft models or algorithms, then verifies them efficiently using the expensive target model with minimal overhead.

How Speculative Decoding Works in vLLM

The implementation follows a four-step iterative process that minimizes calls to the expensive target model:

1. Draft Token Generation

A lightweight drafter generates N speculative tokens in a single forward pass. This drafter can be a small transformer model, an N-gram lookup, or specialized heads like Eagle or Medusa. Because the draft model is significantly smaller than the target model, this step is computationally cheap.

2. Target Model Verification

The full-size target model runs one forward pass on the current prompt concatenated with the draft tokens. Crucially, this single pass produces logits for all positions, but only the prediction for the first draft token is examined initially.

3. Rejection Sampling

The RejectionSampler compares the target model's output distribution with the draft tokens. If the target model's top prediction matches the draft token, that token is accepted; otherwise, the target model's token is used and the remaining draft tokens are discarded. This ensures the final output distribution exactly matches the target model's distribution.

4. State Update and Loop

The KV cache is updated to retain only the accepted tokens' hidden states, and the process repeats from the last accepted position. Because the target model's hidden state from the verification pass can be reused for subsequent generation, no computation is wasted.

Core Components and Source File Locations

The speculative decoding pipeline is implemented across several key modules in the vLLM codebase:

Component Source File Description
SpeculativeConfig vllm/config/speculative.py Pydantic configuration object that parses CLI flags (--speculative-mode, --num-speculative-tokens) and validates draft model compatibility.
Speculator Factory vllm/v1/worker/gpu/spec_decode/__init__.py init_speculator() function that instantiates the appropriate drafter (Eagle, N-gram, Medusa, etc.) based on configuration.
GPU Model Runner vllm/v1/worker/gpu_model_runner.py Core generation loop that creates the drafter, manages the RejectionSampler, and integrates speculative decoding with KV-cache handling.
Eagle Proposer vllm/v1/spec_decode/eagle.py Implementation of Eagle/Eagle-3 draft algorithm with auxiliary hidden-state extraction for efficient verification.
N-gram Proposer vllm/v1/spec_decode/ngram_proposer.py Lightweight n-gram lookup drafter requiring no additional model.
Draft Model Proposer vllm/v1/spec_decode/draft_model.py Wrapper for running a separate lightweight transformer as the drafter.
Medusa Proposer vllm/v1/spec_decode/medusa.py Multi-head draft proposal implementation.
Rejection Sampler vllm/v1/sample/rejection_sampler.py Implements the acceptance/rejection logic and KV-cache rollback for rejected tokens.
SpecDecode Metadata vllm/v1/spec_decode/metadata.py Tracks per-request speculative metadata including draft token counts and KV-cache offsets.

Configuration and Usage Examples

Enabling Speculative Decoding via Python API

from vllm import LLM, SamplingParams

# Initialize vLLM with Eagle speculative decoding

llm = LLM(
    model="meta-llama/Meta-Llama-3.1-70B",
    gpu_memory_utilization=0.9,
    tensor_parallel_size=2,
    speculative_mode="eagle",               # Options: "ngram", "medusa", "draft_model", "eagle3"

    num_speculative_tokens=4,               # Number of draft tokens per step

    draft_model="yuhuili/EAGLE-LLaMA3-Instruct-8B",  # Optional explicit draft model

)

sampling_params = SamplingParams(
    temperature=0.7,
    top_p=0.9,
    max_tokens=128,
)

outputs = llm.generate(
    prompts=["Explain quantum computing in simple terms."],
    sampling_params=sampling_params
)

for output in outputs:
    print(output.text)

The constructor forwards speculative_mode and num_speculative_tokens to SpeculativeConfig in vllm/config/speculative.py.

Using the Command Line Interface


# Start the OpenAI-compatible server with Eagle-3 speculative decoding

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3.1-70B \
    --tensor-parallel-size 2 \
    --speculative-mode eagle3 \
    --num-speculative-tokens 8

Any client using the OpenAI SDK automatically benefits from speculative decoding without code changes.

Inspecting the Active Drafter


# After LLM initialization

print("Active drafter:", llm.llm_engine.model_runner.drafter.__class__.__name__)

# Output: "EagleProposer" or "NgramProposer", etc.

Custom Draft Model Configuration

llm = LLM(
    model="meta-llama/Meta-Llama-3.1-70B",
    speculative_mode="draft_model",
    num_speculative_tokens=3,
    draft_model="facebook/opt-125m",  # Lightweight model for fast drafting

)

The DraftModelProposer in vllm/v1/spec_decode/draft_model.py handles the draft model execution.

Performance Optimizations in vLLM

The vLLM implementation includes several architectural decisions that maximize speculative decoding efficiency:

  • GPU-Resident Draft Models: The draft model resides on the last pipeline-parallel rank to eliminate inter-node communication overhead during draft generation.

  • KV-Cache Reuse: The same KV cache serves both draft and target passes. SpecDecodeMetadata tracks which slots survive acceptance, allowing the system to retain only valid hidden states and discard rejected tokens without full recomputation.

  • Parallel Drafting: When parallel_drafting is enabled (Eagle-3), multiple draft tokens are generated concurrently in a single kernel launch rather than sequentially, reducing draft latency.

  • Auxiliary Hidden-State Sharing: Eagle-3 and extract_hidden_states modes allow the target model to emit intermediate hidden states that the drafter reuses, eliminating duplicated computation for shared layers.

  • Efficient Rejection Sampling: The RejectionSampler in vllm/v1/sample/rejection_sampler.py implements the acceptance logic with minimal overhead, ensuring the output distribution matches the target model while maximizing token acceptance rates.

Summary

  • Speculative decoding in vLLM accelerates inference by generating candidate tokens with a lightweight draft model and validating them with a single forward pass of the expensive target model.
  • The system supports multiple draft methods including Eagle, Eagle-3, Medusa, N-gram lookup, and custom draft models, configurable via SpeculativeConfig in vllm/config/speculative.py.
  • The core execution loop resides in GPUModelRunner (vllm/v1/worker/gpu_model_runner.py), which coordinates the drafter, rejection sampler, and KV-cache management.
  • Rejection sampling ensures the final output distribution exactly matches the target model's distribution, with accepted tokens retained and rejected tokens rolled back via SpecDecodeMetadata.
  • Users enable the feature through simple API flags (speculative_mode, num_speculative_tokens) or CLI arguments without modifying generation code.

Frequently Asked Questions

What is the optimal number of speculative tokens to use?

The optimal value depends on the draft model's accuracy and the target model's size. Typically, 3 to 8 speculative tokens (num_speculative_tokens=4) provides the best latency-throughput trade-off for Eagle-based drafting. Higher values increase potential speedup but also raise the probability of rejection, which wastes computation on discarded draft tokens.

Does speculative decoding change the output distribution of the model?

No. The rejection sampling mechanism in vllm/v1/sample/rejection_sampler.py ensures that the final token distribution exactly matches what the target model would have produced without speculative decoding. Tokens are only accepted if the draft matches the target's distribution; otherwise, the target's token is used and the distribution is corrected.

Can I use speculative decoding with tensor or pipeline parallelism?

Yes. vLLM's implementation places the draft model on the last pipeline-parallel rank to minimize communication overhead. When using tensor parallelism, both draft and target models are sharded across the same GPUs. The GPUModelRunner in vllm/v1/worker/gpu_model_runner.py handles the coordination of speculative decoding across distributed workers.

Which speculative method should I choose: Eagle, Medusa, or N-gram?

Eagle and Eagle-3 generally provide the best speedup for general text generation by leveraging auxiliary hidden states from the target model, achieving high acceptance rates with 4-8 draft tokens. Medusa is effective for specific model architectures that support multiple prediction heads but requires model-specific training. N-gram drafting requires no additional model and works well for repetitive or structured outputs, making it the most memory-efficient option. Configure your choice via the speculative_mode parameter in SpeculativeConfig.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →