vLLM v1 Engine vs Legacy Engine: Key Architectural Differences Explained
The vLLM v1 engine is a complete architectural rewrite that replaces the legacy v0 engine's fragmented prefill/decode scheduler with a unified token-based scheduler, introduces a minimal EngineCore execution loop, and removes legacy features like best_of sampling and CPU-GPU KV-cache swapping.
The vLLM project (vllm-project/vllm) introduced the v1 engine (sometimes referred to as the "vI" engine) to address technical debt and performance limitations inherent in the original v0 implementation. While the legacy engine handled prefill and decode phases through separate code paths with complex bookkeeping, the v1 engine treats all tokens uniformly, enabling chunked prefill and prefix caching by default.
Unified Scheduler: Eliminating the Prefill/Decode Split
The most significant change in the v1 engine is the unified scheduler. The legacy engine maintained separate logic for prompt processing (prefill) and token generation (decode), requiring distinct bookkeeping structures for "prefill-only" cache blocks and generated tokens.
In contrast, the v1 scheduler uses a simple {request_id: num_tokens} dictionary to track all requests uniformly. This design enables:
- Chunked prefill by default, allowing long prompts to be processed across multiple forward passes without blocking the queue
- Prefix caching without special handling for prefill phases
- Speculative decoding without hard splits between prompt and generation phases
As noted in the v1 user guide, this unified approach eliminates the "separate code paths for prompt tokens and generated tokens" that complicated the legacy implementation. See docs/usage/v1_guide.md lines 71-77 for the architectural rationale.
EngineCore Architecture: Separation of Concerns
The v1 engine introduces a strict separation between orchestration and execution through the EngineCore abstraction. While the legacy engine combined model execution, KV-cache management, and request scheduling in a monolithic Engine class, the v1 architecture splits these responsibilities:
-
EngineCore(vllm/v1/engine/core.pylines 83-88): The inner execution loop that only knows how to run one batch of requests. It contains the minimal logic for forward passes and token generation. -
EngineCoreProc/EngineCoreActor: Process-based or Ray-actor-based wrappers that handle the lifecycle and communication withEngineCore. TheEngineCoreActorMixinclass (lines 1796-1802 incore.py) enables both local and distributed execution without code duplication. -
EngineCoreClient(vllm/v1/engine/core_client.pylines 66-69): A minimal RPC client that forwards requests to the core. This replaces the broader API surface of the legacy engine with a narrow, well-defined interface.
Worker Model: Lazy Initialization and Decoupling
The legacy engine tightly coupled workers to the engine process, requiring each worker to implicitly create a "virtual engine" for pipeline parallelism (v0 PP virtual engine). This made multi-executor setups complex and initialization order-critical.
The v1 engine replaces this with WorkerWrapperBase (vllm/v1/worker/worker_base.py lines 79-86), a lazy-initialization wrapper that:
- Decouples the executor from worker lifecycle management
- Only spins up the real worker when first needed
- Makes multiple engines in a single process trivial to implement
- Eliminates the special-case handling for pipeline parallel "virtual engines"
KV-Cache Management: Simplified and Optimized
The legacy engine supported complex KV-cache operations including CPU-GPU swapping and "prefill-only" cache blocks, requiring the scheduler to track cache state across heterogeneous memory hierarchies.
The v1 engine implements a unified KV-cache (KVCacheConfig) with the following characteristics:
- No CPU-GPU swapping: The unified scheduler never pre-emptively discards cache, eliminating the need for swap operations (see removal note in
docs/usage/v1_guide.mdlines 85-92) - FP8 KV-cache support: Optional quantization for memory efficiency
- Simplified allocation: Single configuration object manages cache layout across all workers
Removed Features and Zero-Config Defaults
The v1 engine intentionally removes several legacy features to reduce technical debt:
best_ofsampling: Removed in favor of simpler sampling strategies- Per-request logits processors: Replaced by global logits processors applied uniformly across all requests (see
vllm/v1/engine/llm_engine.pylines 130-138) - Legacy logprobs handling: Now returns raw logprobs by default with optional post-processing flags
Configuration is now zero-config for most deployments:
- Chunked prefill and prefix caching are enabled by default
- The scheduler automatically disables chunked prefill for models without KV-cache
- Global defaults are chosen for performance rather than requiring manual tuning of flags like
--enable-chunked-prefillor--use-cpu-kv-cache
Observability and Metrics
The legacy engine scattered metrics definitions across multiple modules with inconsistent tagging. The v1 engine centralizes observability through prometheus-compatible per-engine collectors.
The make_per_engine helper function (vllm/v1/metrics/perf.py lines 1300-1309) automatically tags metrics with the engine index, providing consistent labeling for multi-engine deployments without manual configuration.
Summary
- Unified scheduler eliminates the prefill/decode split, enabling chunked prefill and prefix caching by default
- EngineCore architecture separates the execution loop from orchestration through
EngineCore,EngineCoreProc, andEngineCoreClient - Lazy worker initialization via
WorkerWrapperBasedecouples executor lifecycle from worker creation - Simplified KV-cache removes CPU-GPU swapping and unifies cache management under
KVCacheConfig - Zero-config defaults remove the need for manual tuning of performance flags while eliminating legacy features like
best_ofand per-request logits processors
Frequently Asked Questions
What happened to the best_of sampling parameter in vLLM v1?
The best_of parameter has been removed in the v1 engine. This feature, which allowed the engine to generate multiple candidate sequences and return the best one, added significant complexity to the scheduler and worker coordination. Users should implement equivalent functionality at the application layer if needed, or rely on alternative sampling strategies like beam search where supported.
How does the unified scheduler affect performance for long-context models?
The unified scheduler improves throughput for long-context workloads by enabling chunked prefill by default. In the legacy engine, long prompts could block the decode queue because prefill and decode were scheduled separately. The v1 engine treats all tokens uniformly, allowing a long prompt to be processed in chunks interleaved with decode steps, reducing head-of-line blocking and improving GPU utilization.
Can I still use custom logits processors with the vLLM v1 engine?
Yes, but with a key architectural change: per-request logits processors are replaced by global logits processors. Instead of passing a logits processor for individual requests, you configure processors at the engine level via EngineArgs. These processors are applied to all requests uniformly during the forward pass in vllm/v1/engine/llm_engine.py. This simplifies the codebase while maintaining flexibility for most use cases.
Is the v1 engine backward compatible with existing vLLM deployments?
The v1 engine is not fully backward compatible due to intentional removal of legacy features. Key breaking changes include the removal of best_of sampling, CPU-GPU KV-cache swapping, and per-request logits processors. However, the v1 engine provides zero-config defaults for performance optimizations that previously required manual flags. Users migrating from v0 should review the removed features list in docs/usage/v1_guide.md and update their client code accordingly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →