# vLLM v1 Engine vs Legacy Engine: Key Architectural Differences Explained

> Explore the vLLM v1 engine vs legacy engine architectural differences. Discover its unified scheduler, `EngineCore` loop, and removal of legacy features for enhanced performance.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: architecture
- Published: 2026-03-03

---

**The vLLM v1 engine is a complete architectural rewrite that replaces the legacy v0 engine's fragmented prefill/decode scheduler with a unified token-based scheduler, introduces a minimal `EngineCore` execution loop, and removes legacy features like `best_of` sampling and CPU-GPU KV-cache swapping.**

The vLLM project (`vllm-project/vllm`) introduced the v1 engine (sometimes referred to as the "vI" engine) to address technical debt and performance limitations inherent in the original v0 implementation. While the legacy engine handled prefill and decode phases through separate code paths with complex bookkeeping, the v1 engine treats all tokens uniformly, enabling chunked prefill and prefix caching by default.

## Unified Scheduler: Eliminating the Prefill/Decode Split

The most significant change in the v1 engine is the **unified scheduler**. The legacy engine maintained separate logic for prompt processing (prefill) and token generation (decode), requiring distinct bookkeeping structures for "prefill-only" cache blocks and generated tokens.

In contrast, the v1 scheduler uses a simple `{request_id: num_tokens}` dictionary to track all requests uniformly. This design enables:

- **Chunked prefill** by default, allowing long prompts to be processed across multiple forward passes without blocking the queue
- **Prefix caching** without special handling for prefill phases
- **Speculative decoding** without hard splits between prompt and generation phases

As noted in the v1 user guide, this unified approach eliminates the "separate code paths for prompt tokens and generated tokens" that complicated the legacy implementation. See [`docs/usage/v1_guide.md`](https://github.com/vllm-project/vllm/blob/main/docs/usage/v1_guide.md) lines 71-77 for the architectural rationale.

## EngineCore Architecture: Separation of Concerns

The v1 engine introduces a strict separation between orchestration and execution through the **`EngineCore`** abstraction. While the legacy engine combined model execution, KV-cache management, and request scheduling in a monolithic `Engine` class, the v1 architecture splits these responsibilities:

- **`EngineCore`** ([`vllm/v1/engine/core.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/core.py) lines 83-88): The inner execution loop that only knows how to run one batch of requests. It contains the minimal logic for forward passes and token generation.

- **`EngineCoreProc`/`EngineCoreActor`**: Process-based or Ray-actor-based wrappers that handle the lifecycle and communication with `EngineCore`. The `EngineCoreActorMixin` class (lines 1796-1802 in [`core.py`](https://github.com/vllm-project/vllm/blob/main/core.py)) enables both local and distributed execution without code duplication.

- **`EngineCoreClient`** ([`vllm/v1/engine/core_client.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/core_client.py) lines 66-69): A minimal RPC client that forwards requests to the core. This replaces the broader API surface of the legacy engine with a narrow, well-defined interface.

## Worker Model: Lazy Initialization and Decoupling

The legacy engine tightly coupled workers to the engine process, requiring each worker to implicitly create a "virtual engine" for pipeline parallelism (`v0 PP virtual engine`). This made multi-executor setups complex and initialization order-critical.

The v1 engine replaces this with **`WorkerWrapperBase`** ([`vllm/v1/worker/worker_base.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/worker_base.py) lines 79-86), a lazy-initialization wrapper that:

- Decouples the executor from worker lifecycle management
- Only spins up the real worker when first needed
- Makes multiple engines in a single process trivial to implement
- Eliminates the special-case handling for pipeline parallel "virtual engines"

## KV-Cache Management: Simplified and Optimized

The legacy engine supported complex KV-cache operations including CPU-GPU swapping and "prefill-only" cache blocks, requiring the scheduler to track cache state across heterogeneous memory hierarchies.

The v1 engine implements a **unified KV-cache** (`KVCacheConfig`) with the following characteristics:

- **No CPU-GPU swapping**: The unified scheduler never pre-emptively discards cache, eliminating the need for swap operations (see removal note in [`docs/usage/v1_guide.md`](https://github.com/vllm-project/vllm/blob/main/docs/usage/v1_guide.md) lines 85-92)
- **FP8 KV-cache support**: Optional quantization for memory efficiency
- **Simplified allocation**: Single configuration object manages cache layout across all workers

## Removed Features and Zero-Config Defaults

The v1 engine intentionally removes several legacy features to reduce technical debt:

- **`best_of` sampling**: Removed in favor of simpler sampling strategies
- **Per-request logits processors**: Replaced by **global logits processors** applied uniformly across all requests (see [`vllm/v1/engine/llm_engine.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/llm_engine.py) lines 130-138)
- **Legacy logprobs handling**: Now returns raw logprobs by default with optional post-processing flags

Configuration is now **zero-config** for most deployments:

- Chunked prefill and prefix caching are **enabled by default**
- The scheduler automatically disables chunked prefill for models without KV-cache
- Global defaults are chosen for performance rather than requiring manual tuning of flags like `--enable-chunked-prefill` or `--use-cpu-kv-cache`

## Observability and Metrics

The legacy engine scattered metrics definitions across multiple modules with inconsistent tagging. The v1 engine centralizes observability through **prometheus-compatible per-engine collectors**.

The `make_per_engine` helper function ([`vllm/v1/metrics/perf.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/metrics/perf.py) lines 1300-1309) automatically tags metrics with the engine index, providing consistent labeling for multi-engine deployments without manual configuration.

## Summary

- **Unified scheduler** eliminates the prefill/decode split, enabling chunked prefill and prefix caching by default
- **EngineCore architecture** separates the execution loop from orchestration through `EngineCore`, `EngineCoreProc`, and `EngineCoreClient`
- **Lazy worker initialization** via `WorkerWrapperBase` decouples executor lifecycle from worker creation
- **Simplified KV-cache** removes CPU-GPU swapping and unifies cache management under `KVCacheConfig`
- **Zero-config defaults** remove the need for manual tuning of performance flags while eliminating legacy features like `best_of` and per-request logits processors

## Frequently Asked Questions

### What happened to the `best_of` sampling parameter in vLLM v1?

The `best_of` parameter has been **removed** in the v1 engine. This feature, which allowed the engine to generate multiple candidate sequences and return the best one, added significant complexity to the scheduler and worker coordination. Users should implement equivalent functionality at the application layer if needed, or rely on alternative sampling strategies like beam search where supported.

### How does the unified scheduler affect performance for long-context models?

The unified scheduler **improves throughput** for long-context workloads by enabling **chunked prefill** by default. In the legacy engine, long prompts could block the decode queue because prefill and decode were scheduled separately. The v1 engine treats all tokens uniformly, allowing a long prompt to be processed in chunks interleaved with decode steps, reducing head-of-line blocking and improving GPU utilization.

### Can I still use custom logits processors with the vLLM v1 engine?

Yes, but with a key architectural change: **per-request logits processors are replaced by global logits processors**. Instead of passing a logits processor for individual requests, you configure processors at the engine level via `EngineArgs`. These processors are applied to all requests uniformly during the forward pass in [`vllm/v1/engine/llm_engine.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/llm_engine.py). This simplifies the codebase while maintaining flexibility for most use cases.

### Is the v1 engine backward compatible with existing vLLM deployments?

The v1 engine is **not fully backward compatible** due to intentional removal of legacy features. Key breaking changes include the removal of `best_of` sampling, CPU-GPU KV-cache swapping, and per-request logits processors. However, the v1 engine provides **zero-config defaults** for performance optimizations that previously required manual flags. Users migrating from v0 should review the removed features list in [`docs/usage/v1_guide.md`](https://github.com/vllm-project/vllm/blob/main/docs/usage/v1_guide.md) and update their client code accordingly.