How the vLLM Scheduler Handles Request Prioritization and Preemption
The vLLM scheduler uses a configurable SchedulerConfig.policy ("fcfs" or "priority") to determine request ordering; when "priority" is selected, it employs a heap-based PriorityRequestQueue ordered by Request.priority (lower integer = higher priority) and preempts the lowest-priority running requests when KV-cache allocation fails.
The vLLM inference engine implements a sophisticated scheduling mechanism that balances throughput and latency-critical workloads. Unlike simple first-come-first-served systems, the vLLM scheduler can prioritize specific requests and preempt lower-priority work to ensure high-priority tasks obtain GPU resources. This article examines the implementation details found in the vllm-project/vllm repository, covering the configuration options, priority queue mechanics, and preemption algorithms that govern request execution.
Scheduler Configuration and Policies
The scheduling behavior is determined at engine initialization through SchedulerConfig, defined in vllm/config/scheduler.py. This configuration specifies the policy field, which accepts two values:
"fcfs"(default): First-come-first-served processing using a simpledeque"priority": Priority-based scheduling using a min-heap that respects request priority levels
During scheduler construction in vllm/v1/core/sched/scheduler.py (lines 54-62), the code reads self.scheduler_config.policy and invokes create_request_queue to instantiate the appropriate queue implementation. The factory function, located in vllm/v1/core/sched/request_queue.py (lines 201-209), returns either an FCFSRequestQueue or PriorityRequestQueue based on this policy setting.
Request Prioritization Mechanics
When priority scheduling is enabled, the system relies on two core components: the Request data class and the PriorityRequestQueue.
Request Priority Attributes
Each Request object, defined in vllm/v1/request.py (lines 70-73), carries a priority: int field defaulting to 0. The class implements the __lt__ method (lines 93-104) to enable heap ordering:
def __lt__(self, other: "Request") -> bool:
# Lower priority value = higher priority
if self.priority != other.priority:
return self.priority < other.priority
# Tie-break by arrival time (earlier wins)
if self.arrival_time != other.arrival_time:
return self.arrival_time < other.arrival_time
# Final tie-break by request_id
return self.request_id < other.request_id
This comparison logic ensures that the heap always surfaces the highest-priority request (lowest integer value), using arrival time as a secondary sort key to maintain fairness among equal-priority requests.
Priority Queue Implementation
The PriorityRequestQueue class in vllm/v1/core/sched/request_queue.py (lines 31-99) wraps Python's heapq module. When the scheduler calls get_next_request(), the queue pops the highest-priority item according to the Request.__lt__ ordering. New requests are added via heapq.heappush, maintaining the invariant that the most urgent request always resides at index 0.
Preemption Logic in the Scheduling Loop
Preemption occurs exclusively under the priority policy when the KV-cache manager cannot allocate required slots for an incoming high-priority request. The logic resides in the main scheduling loop within vllm/v1/core/sched/scheduler.py.
Allocation Failure Handling
When kv_cache_manager.allocate_slots returns None (indicating insufficient free blocks), the scheduler enters a preemption loop:
while True:
new_blocks = self.kv_cache_manager.allocate_slots(...)
if new_blocks is not None:
break # Allocation succeeded
# Preemption required
if self.policy == SchedulingPolicy.PRIORITY:
# Select lowest-priority running request
preempted_req = max(
self.running,
key=lambda r: (r.priority, r.arrival_time)
)
self._preempt_request(preempted_req, scheduled_timestamp)
self.waiting.prepend_request(preempted_req)
else:
# FCFS policy: no preemption allowed
break
The max() function selects the running request with the highest priority value (lowest actual priority) and latest arrival time. This selection ensures that latency-sensitive or admin-designated high-priority requests displace less critical work.
Preemption Execution
The _preempt_request method (lines 12-31 in vllm/v1/core/sched/scheduler.py) performs the following atomic operations:
- Resource Liberation: Frees KV-cache blocks and encoder cache allocations associated with the target request
- State Reset: Resets
num_computed_tokensto zero and increments thenum_preemptionscounter - Status Update: Sets the request status to
PREEMPTED - Re-queuing: Prepends the request to the waiting queue via
prepend_request, ensuring it retains priority relative to other waiting requests when resources become available
Under the FCFS policy, the scheduler simply breaks the allocation loop without preemption, leaving the request in the waiting state until earlier requests complete and release resources naturally.
Practical Implementation Examples
Configuring Priority-Based Scheduling
To enable request prioritization and preemption, instantiate the engine with an explicit SchedulerConfig:
from vllm import LLMEngine
from vllm.config import SchedulerConfig, EngineConfig
# Configure priority scheduling with tight resource constraints
sched_cfg = SchedulerConfig.default_factory(
policy="priority",
max_num_batched_tokens=256,
max_num_seqs=8,
)
engine_cfg = EngineConfig(
model="facebook/opt-125m",
scheduler_config=sched_cfg,
cache_config={"num_gpu_blocks": 2, "block_size": 16},
)
engine = LLMEngine(engine_config=engine_cfg)
Submitting Prioritized Requests
Use the priority parameter (lower values indicate higher urgency) when adding requests:
# Critical request with highest priority (priority = -10)
engine.add_request(
request_id="urgent-admin-query",
prompt="Analyze system health metrics...",
sampling_params={"temperature": 0.0},
priority=-10,
)
# Background batch jobs with default priority (0)
for i in range(5):
engine.add_request(
request_id=f"background-job-{i}",
prompt="Summarize document batch...",
sampling_params={"temperature": 0.7},
priority=0,
)
Monitoring Preemption Events
Inspect scheduler state during execution to observe preemption behavior:
while not engine.is_finished():
engine.step()
sched = engine.scheduler
print(f"Running: {[r.request_id for r in sched.running]}")
print(f"Waiting: {[r.request_id for r in sched.waiting]}")
# Detect preempted requests
for req in sched.waiting:
if req.num_preemptions > 0:
print(f"Request {req.request_id} was preempted "
f"({req.num_preemptions} times)")
When GPU memory pressure forces preemption, you will observe low-priority requests moving from running to waiting with incremented num_preemptions counters, while high-priority requests (negative values) retain their execution slots.
FCFS Baseline Configuration
For comparison, standard first-come-first-served behavior disables preemption entirely:
sched_cfg_fcfs = SchedulerConfig.default_factory(policy="fcfs")
engine_cfg_fcfs = EngineConfig(
model="facebook/opt-125m",
scheduler_config=sched_cfg_fcfs,
cache_config={"num_gpu_blocks": 2, "block_size": 16}
)
engine_fcfs = LLMEngine(engine_config=engine_cfg_fcfs)
Key Source Files
The vLLM scheduling system spans several core modules:
vllm/config/scheduler.py: DefinesSchedulerConfigincluding thepolicyenum and token budget constraintsvllm/v1/core/sched/scheduler.py: Implements the main scheduling loop, preemption logic, and_preempt_requestmethodvllm/v1/core/sched/request_queue.py: ContainsFCFSRequestQueue,PriorityRequestQueue, and thecreate_request_queuefactoryvllm/v1/request.py: Defines theRequestdataclass withpriority,arrival_time, and comparison operators for heap orderingvllm/v1/core/kv_cache_manager.py: Handles block allocation; allocation failures trigger the preemption cascade in the scheduler
Summary
- Policy Selection: The vLLM scheduler supports
"fcfs"(non-preemptive) and"priority"(preemptive) policies configured viaSchedulerConfig.policyinvllm/config/scheduler.py - Priority Ordering: Under the priority policy,
PriorityRequestQueueuses Python'sheapqordered byRequest.priority(lower integer = higher priority) andarrival_timeto break ties - Selective Preemption: Preemption only occurs when
policy="priority"and KV-cache allocation fails; the scheduler preempts the lowest-priority running request usingmax()selection on(priority, arrival_time) - Resource Reclamation: The
_preempt_requestmethod invllm/v1/core/sched/scheduler.pyfrees cache blocks, marks requests asPREEMPTED, and re-queues them for later execution - FCFS Protection: First-come-first-served mode prevents preemption entirely, ensuring predictable but potentially slower execution for high-priority late arrivals
Frequently Asked Questions
What is the difference between FCFS and priority scheduling in vLLM?
FCFS (First-Come-First-Served) processes requests in arrival order without preemption, meaning once a request begins running, it retains its KV-cache blocks until completion. Priority scheduling allows requests with lower integer priority values (e.g., -10) to preempt running requests with higher values (e.g., 0), forcibly freeing GPU memory for the urgent task. FCFS provides predictable latency for early requests, while priority scheduling optimizes for latency-sensitive workloads.
How do I assign priority to a request in vLLM?
Pass the priority integer parameter when calling engine.add_request() or the equivalent API. Lower values indicate higher priority; for example, priority=-10 executes before the default priority=0. The priority value is stored in the Request object defined in vllm/v1/request.py and evaluated by the PriorityRequestQueue when determining scheduling order.
Does priority scheduling impact throughput?
Priority scheduling can reduce overall throughput when frequent preemptions occur, because preempted requests must be re-computed from their last checkpoint when resumed. However, it significantly improves latency for high-priority requests by ensuring they are not blocked by long-running, low-priority batch jobs. The trade-off is acceptable when serving mixed workloads where certain queries (e.g., admin commands, real-time user interactions) require immediate attention over background processing.
What happens to a request after it is preempted?
When preempted, the request's status changes to PREEMPTED, its allocated KV-cache blocks are freed via kv_cache_manager, and it is prepended to the waiting queue. The num_preemptions counter increments to track how many times the request has been interrupted. The request will resume execution when it reaches the front of the priority queue and sufficient GPU memory becomes available, though it may need to recompute tokens depending on the checkpointing strategy configured.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →