# How Speculative Decoding in vLLM Accelerates Token Generation

> Discover how speculative decoding in vLLM accelerates token generation. Learn how a draft model and rejection sampling boost speed N times with exact output distribution.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: deep-dive
- Published: 2026-03-03

---

**Speculative decoding in vLLM uses a lightweight draft model to predict multiple tokens ahead, validates them with a single forward pass of the full target model, and accepts valid tokens via rejection sampling to achieve up to N× speedup while maintaining the exact output distribution.**

Speculative decoding is a look-ahead optimization technique that dramatically reduces latency in large language model inference. In the vllm-project/vllm repository, this feature is implemented as a configurable pipeline that generates candidate token sequences using cheap draft models or algorithms, then verifies them efficiently using the expensive target model with minimal overhead.

## How Speculative Decoding Works in vLLM

The implementation follows a four-step iterative process that minimizes calls to the expensive target model:

### 1. Draft Token Generation

A lightweight **drafter** generates *N* speculative tokens in a single forward pass. This drafter can be a small transformer model, an N-gram lookup, or specialized heads like Eagle or Medusa. Because the draft model is significantly smaller than the target model, this step is computationally cheap.

### 2. Target Model Verification

The full-size **target model** runs one forward pass on the current prompt concatenated with the draft tokens. Crucially, this single pass produces logits for all positions, but only the prediction for the *first* draft token is examined initially.

### 3. Rejection Sampling

The `RejectionSampler` compares the target model's output distribution with the draft tokens. If the target model's top prediction matches the draft token, that token is **accepted**; otherwise, the target model's token is used and the remaining draft tokens are discarded. This ensures the final output distribution exactly matches the target model's distribution.

### 4. State Update and Loop

The KV cache is updated to retain only the accepted tokens' hidden states, and the process repeats from the last accepted position. Because the target model's hidden state from the verification pass can be reused for subsequent generation, no computation is wasted.

## Core Components and Source File Locations

The speculative decoding pipeline is implemented across several key modules in the vLLM codebase:

| Component | Source File | Description |
|-----------|-------------|-------------|
| **SpeculativeConfig** | [`vllm/config/speculative.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/speculative.py) | Pydantic configuration object that parses CLI flags (`--speculative-mode`, `--num-speculative-tokens`) and validates draft model compatibility. |
| **Speculator Factory** | [`vllm/v1/worker/gpu/spec_decode/__init__.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu/spec_decode/__init__.py) | `init_speculator()` function that instantiates the appropriate drafter (Eagle, N-gram, Medusa, etc.) based on configuration. |
| **GPU Model Runner** | [`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py) | Core generation loop that creates the drafter, manages the `RejectionSampler`, and integrates speculative decoding with KV-cache handling. |
| **Eagle Proposer** | [`vllm/v1/spec_decode/eagle.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/spec_decode/eagle.py) | Implementation of Eagle/Eagle-3 draft algorithm with auxiliary hidden-state extraction for efficient verification. |
| **N-gram Proposer** | [`vllm/v1/spec_decode/ngram_proposer.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/spec_decode/ngram_proposer.py) | Lightweight n-gram lookup drafter requiring no additional model. |
| **Draft Model Proposer** | [`vllm/v1/spec_decode/draft_model.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/spec_decode/draft_model.py) | Wrapper for running a separate lightweight transformer as the drafter. |
| **Medusa Proposer** | [`vllm/v1/spec_decode/medusa.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/spec_decode/medusa.py) | Multi-head draft proposal implementation. |
| **Rejection Sampler** | [`vllm/v1/sample/rejection_sampler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/sample/rejection_sampler.py) | Implements the acceptance/rejection logic and KV-cache rollback for rejected tokens. |
| **SpecDecode Metadata** | [`vllm/v1/spec_decode/metadata.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/spec_decode/metadata.py) | Tracks per-request speculative metadata including draft token counts and KV-cache offsets. |

## Configuration and Usage Examples

### Enabling Speculative Decoding via Python API

```python
from vllm import LLM, SamplingParams

# Initialize vLLM with Eagle speculative decoding

llm = LLM(
    model="meta-llama/Meta-Llama-3.1-70B",
    gpu_memory_utilization=0.9,
    tensor_parallel_size=2,
    speculative_mode="eagle",               # Options: "ngram", "medusa", "draft_model", "eagle3"

    num_speculative_tokens=4,               # Number of draft tokens per step

    draft_model="yuhuili/EAGLE-LLaMA3-Instruct-8B",  # Optional explicit draft model

)

sampling_params = SamplingParams(
    temperature=0.7,
    top_p=0.9,
    max_tokens=128,
)

outputs = llm.generate(
    prompts=["Explain quantum computing in simple terms."],
    sampling_params=sampling_params
)

for output in outputs:
    print(output.text)

```

The constructor forwards `speculative_mode` and `num_speculative_tokens` to `SpeculativeConfig` in [`vllm/config/speculative.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/speculative.py).

### Using the Command Line Interface

```bash

# Start the OpenAI-compatible server with Eagle-3 speculative decoding

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3.1-70B \
    --tensor-parallel-size 2 \
    --speculative-mode eagle3 \
    --num-speculative-tokens 8

```

Any client using the OpenAI SDK automatically benefits from speculative decoding without code changes.

### Inspecting the Active Drafter

```python

# After LLM initialization

print("Active drafter:", llm.llm_engine.model_runner.drafter.__class__.__name__)

# Output: "EagleProposer" or "NgramProposer", etc.

```

### Custom Draft Model Configuration

```python
llm = LLM(
    model="meta-llama/Meta-Llama-3.1-70B",
    speculative_mode="draft_model",
    num_speculative_tokens=3,
    draft_model="facebook/opt-125m",  # Lightweight model for fast drafting

)

```

The `DraftModelProposer` in [`vllm/v1/spec_decode/draft_model.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/spec_decode/draft_model.py) handles the draft model execution.

## Performance Optimizations in vLLM

The vLLM implementation includes several architectural decisions that maximize speculative decoding efficiency:

- **GPU-Resident Draft Models**: The draft model resides on the last pipeline-parallel rank to eliminate inter-node communication overhead during draft generation.

- **KV-Cache Reuse**: The same KV cache serves both draft and target passes. `SpecDecodeMetadata` tracks which slots survive acceptance, allowing the system to retain only valid hidden states and discard rejected tokens without full recomputation.

- **Parallel Drafting**: When `parallel_drafting` is enabled (Eagle-3), multiple draft tokens are generated concurrently in a single kernel launch rather than sequentially, reducing draft latency.

- **Auxiliary Hidden-State Sharing**: Eagle-3 and `extract_hidden_states` modes allow the target model to emit intermediate hidden states that the drafter reuses, eliminating duplicated computation for shared layers.

- **Efficient Rejection Sampling**: The `RejectionSampler` in [`vllm/v1/sample/rejection_sampler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/sample/rejection_sampler.py) implements the acceptance logic with minimal overhead, ensuring the output distribution matches the target model while maximizing token acceptance rates.

## Summary

- Speculative decoding in vLLM accelerates inference by generating candidate tokens with a lightweight draft model and validating them with a single forward pass of the expensive target model.
- The system supports multiple draft methods including Eagle, Eagle-3, Medusa, N-gram lookup, and custom draft models, configurable via `SpeculativeConfig` in [`vllm/config/speculative.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/speculative.py).
- The core execution loop resides in `GPUModelRunner` ([`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py)), which coordinates the drafter, rejection sampler, and KV-cache management.
- Rejection sampling ensures the final output distribution exactly matches the target model's distribution, with accepted tokens retained and rejected tokens rolled back via `SpecDecodeMetadata`.
- Users enable the feature through simple API flags (`speculative_mode`, `num_speculative_tokens`) or CLI arguments without modifying generation code.

## Frequently Asked Questions

### What is the optimal number of speculative tokens to use?

The optimal value depends on the draft model's accuracy and the target model's size. Typically, 3 to 8 speculative tokens (`num_speculative_tokens=4`) provides the best latency-throughput trade-off for Eagle-based drafting. Higher values increase potential speedup but also raise the probability of rejection, which wastes computation on discarded draft tokens.

### Does speculative decoding change the output distribution of the model?

No. The rejection sampling mechanism in [`vllm/v1/sample/rejection_sampler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/sample/rejection_sampler.py) ensures that the final token distribution exactly matches what the target model would have produced without speculative decoding. Tokens are only accepted if the draft matches the target's distribution; otherwise, the target's token is used and the distribution is corrected.

### Can I use speculative decoding with tensor or pipeline parallelism?

Yes. vLLM's implementation places the draft model on the last pipeline-parallel rank to minimize communication overhead. When using tensor parallelism, both draft and target models are sharded across the same GPUs. The `GPUModelRunner` in [`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py) handles the coordination of speculative decoding across distributed workers.

### Which speculative method should I choose: Eagle, Medusa, or N-gram?

Eagle and Eagle-3 generally provide the best speedup for general text generation by leveraging auxiliary hidden states from the target model, achieving high acceptance rates with 4-8 draft tokens. Medusa is effective for specific model architectures that support multiple prediction heads but requires model-specific training. N-gram drafting requires no additional model and works well for repetitive or structured outputs, making it the most memory-efficient option. Configure your choice via the `speculative_mode` parameter in `SpeculativeConfig`.