How the vLLM Scheduler Handles Request Prioritization and Preemption

The vLLM scheduler uses a configurable SchedulerConfig.policy ("fcfs" or "priority") to determine request ordering; when "priority" is selected, it employs a heap-based PriorityRequestQueue ordered by Request.priority (lower integer = higher priority) and preempts the lowest-priority running requests when KV-cache allocation fails.

The vLLM inference engine implements a sophisticated scheduling mechanism that balances throughput and latency-critical workloads. Unlike simple first-come-first-served systems, the vLLM scheduler can prioritize specific requests and preempt lower-priority work to ensure high-priority tasks obtain GPU resources. This article examines the implementation details found in the vllm-project/vllm repository, covering the configuration options, priority queue mechanics, and preemption algorithms that govern request execution.

Scheduler Configuration and Policies

The scheduling behavior is determined at engine initialization through SchedulerConfig, defined in vllm/config/scheduler.py. This configuration specifies the policy field, which accepts two values:

  • "fcfs" (default): First-come-first-served processing using a simple deque
  • "priority": Priority-based scheduling using a min-heap that respects request priority levels

During scheduler construction in vllm/v1/core/sched/scheduler.py (lines 54-62), the code reads self.scheduler_config.policy and invokes create_request_queue to instantiate the appropriate queue implementation. The factory function, located in vllm/v1/core/sched/request_queue.py (lines 201-209), returns either an FCFSRequestQueue or PriorityRequestQueue based on this policy setting.

Request Prioritization Mechanics

When priority scheduling is enabled, the system relies on two core components: the Request data class and the PriorityRequestQueue.

Request Priority Attributes

Each Request object, defined in vllm/v1/request.py (lines 70-73), carries a priority: int field defaulting to 0. The class implements the __lt__ method (lines 93-104) to enable heap ordering:

def __lt__(self, other: "Request") -> bool:
    # Lower priority value = higher priority

    if self.priority != other.priority:
        return self.priority < other.priority
    # Tie-break by arrival time (earlier wins)

    if self.arrival_time != other.arrival_time:
        return self.arrival_time < other.arrival_time
    # Final tie-break by request_id

    return self.request_id < other.request_id

This comparison logic ensures that the heap always surfaces the highest-priority request (lowest integer value), using arrival time as a secondary sort key to maintain fairness among equal-priority requests.

Priority Queue Implementation

The PriorityRequestQueue class in vllm/v1/core/sched/request_queue.py (lines 31-99) wraps Python's heapq module. When the scheduler calls get_next_request(), the queue pops the highest-priority item according to the Request.__lt__ ordering. New requests are added via heapq.heappush, maintaining the invariant that the most urgent request always resides at index 0.

Preemption Logic in the Scheduling Loop

Preemption occurs exclusively under the priority policy when the KV-cache manager cannot allocate required slots for an incoming high-priority request. The logic resides in the main scheduling loop within vllm/v1/core/sched/scheduler.py.

Allocation Failure Handling

When kv_cache_manager.allocate_slots returns None (indicating insufficient free blocks), the scheduler enters a preemption loop:

while True:
    new_blocks = self.kv_cache_manager.allocate_slots(...)
    if new_blocks is not None:
        break  # Allocation succeeded

    
    # Preemption required

    if self.policy == SchedulingPolicy.PRIORITY:
        # Select lowest-priority running request

        preempted_req = max(
            self.running, 
            key=lambda r: (r.priority, r.arrival_time)
        )
        self._preempt_request(preempted_req, scheduled_timestamp)
        self.waiting.prepend_request(preempted_req)
    else:
        # FCFS policy: no preemption allowed

        break

The max() function selects the running request with the highest priority value (lowest actual priority) and latest arrival time. This selection ensures that latency-sensitive or admin-designated high-priority requests displace less critical work.

Preemption Execution

The _preempt_request method (lines 12-31 in vllm/v1/core/sched/scheduler.py) performs the following atomic operations:

  1. Resource Liberation: Frees KV-cache blocks and encoder cache allocations associated with the target request
  2. State Reset: Resets num_computed_tokens to zero and increments the num_preemptions counter
  3. Status Update: Sets the request status to PREEMPTED
  4. Re-queuing: Prepends the request to the waiting queue via prepend_request, ensuring it retains priority relative to other waiting requests when resources become available

Under the FCFS policy, the scheduler simply breaks the allocation loop without preemption, leaving the request in the waiting state until earlier requests complete and release resources naturally.

Practical Implementation Examples

Configuring Priority-Based Scheduling

To enable request prioritization and preemption, instantiate the engine with an explicit SchedulerConfig:

from vllm import LLMEngine
from vllm.config import SchedulerConfig, EngineConfig

# Configure priority scheduling with tight resource constraints

sched_cfg = SchedulerConfig.default_factory(
    policy="priority",
    max_num_batched_tokens=256,
    max_num_seqs=8,
)

engine_cfg = EngineConfig(
    model="facebook/opt-125m",
    scheduler_config=sched_cfg,
    cache_config={"num_gpu_blocks": 2, "block_size": 16},
)

engine = LLMEngine(engine_config=engine_cfg)

Submitting Prioritized Requests

Use the priority parameter (lower values indicate higher urgency) when adding requests:


# Critical request with highest priority (priority = -10)

engine.add_request(
    request_id="urgent-admin-query",
    prompt="Analyze system health metrics...",
    sampling_params={"temperature": 0.0},
    priority=-10,
)

# Background batch jobs with default priority (0)

for i in range(5):
    engine.add_request(
        request_id=f"background-job-{i}",
        prompt="Summarize document batch...",
        sampling_params={"temperature": 0.7},
        priority=0,
    )

Monitoring Preemption Events

Inspect scheduler state during execution to observe preemption behavior:

while not engine.is_finished():
    engine.step()
    
    sched = engine.scheduler
    print(f"Running: {[r.request_id for r in sched.running]}")
    print(f"Waiting: {[r.request_id for r in sched.waiting]}")
    
    # Detect preempted requests

    for req in sched.waiting:
        if req.num_preemptions > 0:
            print(f"Request {req.request_id} was preempted "
                  f"({req.num_preemptions} times)")

When GPU memory pressure forces preemption, you will observe low-priority requests moving from running to waiting with incremented num_preemptions counters, while high-priority requests (negative values) retain their execution slots.

FCFS Baseline Configuration

For comparison, standard first-come-first-served behavior disables preemption entirely:

sched_cfg_fcfs = SchedulerConfig.default_factory(policy="fcfs")
engine_cfg_fcfs = EngineConfig(
    model="facebook/opt-125m",
    scheduler_config=sched_cfg_fcfs,
    cache_config={"num_gpu_blocks": 2, "block_size": 16}
)
engine_fcfs = LLMEngine(engine_config=engine_cfg_fcfs)

Key Source Files

The vLLM scheduling system spans several core modules:

Summary

  • Policy Selection: The vLLM scheduler supports "fcfs" (non-preemptive) and "priority" (preemptive) policies configured via SchedulerConfig.policy in vllm/config/scheduler.py
  • Priority Ordering: Under the priority policy, PriorityRequestQueue uses Python's heapq ordered by Request.priority (lower integer = higher priority) and arrival_time to break ties
  • Selective Preemption: Preemption only occurs when policy="priority" and KV-cache allocation fails; the scheduler preempts the lowest-priority running request using max() selection on (priority, arrival_time)
  • Resource Reclamation: The _preempt_request method in vllm/v1/core/sched/scheduler.py frees cache blocks, marks requests as PREEMPTED, and re-queues them for later execution
  • FCFS Protection: First-come-first-served mode prevents preemption entirely, ensuring predictable but potentially slower execution for high-priority late arrivals

Frequently Asked Questions

What is the difference between FCFS and priority scheduling in vLLM?

FCFS (First-Come-First-Served) processes requests in arrival order without preemption, meaning once a request begins running, it retains its KV-cache blocks until completion. Priority scheduling allows requests with lower integer priority values (e.g., -10) to preempt running requests with higher values (e.g., 0), forcibly freeing GPU memory for the urgent task. FCFS provides predictable latency for early requests, while priority scheduling optimizes for latency-sensitive workloads.

How do I assign priority to a request in vLLM?

Pass the priority integer parameter when calling engine.add_request() or the equivalent API. Lower values indicate higher priority; for example, priority=-10 executes before the default priority=0. The priority value is stored in the Request object defined in vllm/v1/request.py and evaluated by the PriorityRequestQueue when determining scheduling order.

Does priority scheduling impact throughput?

Priority scheduling can reduce overall throughput when frequent preemptions occur, because preempted requests must be re-computed from their last checkpoint when resumed. However, it significantly improves latency for high-priority requests by ensuring they are not blocked by long-running, low-priority batch jobs. The trade-off is acceptable when serving mixed workloads where certain queries (e.g., admin commands, real-time user interactions) require immediate attention over background processing.

What happens to a request after it is preempted?

When preempted, the request's status changes to PREEMPTED, its allocated KV-cache blocks are freed via kv_cache_manager, and it is prepended to the waiting queue. The num_preemptions counter increments to track how many times the request has been interrupted. The request will resume execution when it reaches the front of the priority queue and sufficient GPU memory becomes available, though it may need to recompute tokens depending on the checkpointing strategy configured.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →