# How the vLLM Scheduler Handles Request Prioritization and Preemption

> Discover how the vLLM scheduler prioritizes requests and preempts lower priority tasks when KV cache allocation fails. Learn about FCFS and priority scheduling policies.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**The vLLM scheduler uses a configurable `SchedulerConfig.policy` (`"fcfs"` or `"priority"`) to determine request ordering; when `"priority"` is selected, it employs a heap-based `PriorityRequestQueue` ordered by `Request.priority` (lower integer = higher priority) and preempts the lowest-priority running requests when KV-cache allocation fails.**

The vLLM inference engine implements a sophisticated scheduling mechanism that balances throughput and latency-critical workloads. Unlike simple first-come-first-served systems, the vLLM scheduler can prioritize specific requests and preempt lower-priority work to ensure high-priority tasks obtain GPU resources. This article examines the implementation details found in the `vllm-project/vllm` repository, covering the configuration options, priority queue mechanics, and preemption algorithms that govern request execution.

## Scheduler Configuration and Policies

The scheduling behavior is determined at engine initialization through `SchedulerConfig`, defined in [`vllm/config/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/scheduler.py). This configuration specifies the `policy` field, which accepts two values:

- **`"fcfs"`** (default): First-come-first-served processing using a simple `deque`
- **`"priority"`**: Priority-based scheduling using a min-heap that respects request priority levels

During scheduler construction in [`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py) (lines 54-62), the code reads `self.scheduler_config.policy` and invokes `create_request_queue` to instantiate the appropriate queue implementation. The factory function, located in [`vllm/v1/core/sched/request_queue.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/request_queue.py) (lines 201-209), returns either an `FCFSRequestQueue` or `PriorityRequestQueue` based on this policy setting.

## Request Prioritization Mechanics

When priority scheduling is enabled, the system relies on two core components: the `Request` data class and the `PriorityRequestQueue`.

### Request Priority Attributes

Each `Request` object, defined in [`vllm/v1/request.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/request.py) (lines 70-73), carries a `priority: int` field defaulting to `0`. The class implements the `__lt__` method (lines 93-104) to enable heap ordering:

```python
def __lt__(self, other: "Request") -> bool:
    # Lower priority value = higher priority

    if self.priority != other.priority:
        return self.priority < other.priority
    # Tie-break by arrival time (earlier wins)

    if self.arrival_time != other.arrival_time:
        return self.arrival_time < other.arrival_time
    # Final tie-break by request_id

    return self.request_id < other.request_id

```

This comparison logic ensures that the heap always surfaces the highest-priority request (lowest integer value), using arrival time as a secondary sort key to maintain fairness among equal-priority requests.

### Priority Queue Implementation

The `PriorityRequestQueue` class in [`vllm/v1/core/sched/request_queue.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/request_queue.py) (lines 31-99) wraps Python's `heapq` module. When the scheduler calls `get_next_request()`, the queue pops the highest-priority item according to the `Request.__lt__` ordering. New requests are added via `heapq.heappush`, maintaining the invariant that the most urgent request always resides at index `0`.

## Preemption Logic in the Scheduling Loop

Preemption occurs exclusively under the priority policy when the KV-cache manager cannot allocate required slots for an incoming high-priority request. The logic resides in the main scheduling loop within [`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py).

### Allocation Failure Handling

When `kv_cache_manager.allocate_slots` returns `None` (indicating insufficient free blocks), the scheduler enters a preemption loop:

```python
while True:
    new_blocks = self.kv_cache_manager.allocate_slots(...)
    if new_blocks is not None:
        break  # Allocation succeeded

    
    # Preemption required

    if self.policy == SchedulingPolicy.PRIORITY:
        # Select lowest-priority running request

        preempted_req = max(
            self.running, 
            key=lambda r: (r.priority, r.arrival_time)
        )
        self._preempt_request(preempted_req, scheduled_timestamp)
        self.waiting.prepend_request(preempted_req)
    else:
        # FCFS policy: no preemption allowed

        break

```

The `max()` function selects the running request with the highest priority value (lowest actual priority) and latest arrival time. This selection ensures that latency-sensitive or admin-designated high-priority requests displace less critical work.

### Preemption Execution

The `_preempt_request` method (lines 12-31 in [`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py)) performs the following atomic operations:

1. **Resource Liberation**: Frees KV-cache blocks and encoder cache allocations associated with the target request
2. **State Reset**: Resets `num_computed_tokens` to zero and increments the `num_preemptions` counter
3. **Status Update**: Sets the request status to `PREEMPTED`
4. **Re-queuing**: Prepends the request to the waiting queue via `prepend_request`, ensuring it retains priority relative to other waiting requests when resources become available

Under the FCFS policy, the scheduler simply breaks the allocation loop without preemption, leaving the request in the waiting state until earlier requests complete and release resources naturally.

## Practical Implementation Examples

### Configuring Priority-Based Scheduling

To enable request prioritization and preemption, instantiate the engine with an explicit `SchedulerConfig`:

```python
from vllm import LLMEngine
from vllm.config import SchedulerConfig, EngineConfig

# Configure priority scheduling with tight resource constraints

sched_cfg = SchedulerConfig.default_factory(
    policy="priority",
    max_num_batched_tokens=256,
    max_num_seqs=8,
)

engine_cfg = EngineConfig(
    model="facebook/opt-125m",
    scheduler_config=sched_cfg,
    cache_config={"num_gpu_blocks": 2, "block_size": 16},
)

engine = LLMEngine(engine_config=engine_cfg)

```

### Submitting Prioritized Requests

Use the `priority` parameter (lower values indicate higher urgency) when adding requests:

```python

# Critical request with highest priority (priority = -10)

engine.add_request(
    request_id="urgent-admin-query",
    prompt="Analyze system health metrics...",
    sampling_params={"temperature": 0.0},
    priority=-10,
)

# Background batch jobs with default priority (0)

for i in range(5):
    engine.add_request(
        request_id=f"background-job-{i}",
        prompt="Summarize document batch...",
        sampling_params={"temperature": 0.7},
        priority=0,
    )

```

### Monitoring Preemption Events

Inspect scheduler state during execution to observe preemption behavior:

```python
while not engine.is_finished():
    engine.step()
    
    sched = engine.scheduler
    print(f"Running: {[r.request_id for r in sched.running]}")
    print(f"Waiting: {[r.request_id for r in sched.waiting]}")
    
    # Detect preempted requests

    for req in sched.waiting:
        if req.num_preemptions > 0:
            print(f"Request {req.request_id} was preempted "
                  f"({req.num_preemptions} times)")

```

When GPU memory pressure forces preemption, you will observe low-priority requests moving from `running` to `waiting` with incremented `num_preemptions` counters, while high-priority requests (negative values) retain their execution slots.

### FCFS Baseline Configuration

For comparison, standard first-come-first-served behavior disables preemption entirely:

```python
sched_cfg_fcfs = SchedulerConfig.default_factory(policy="fcfs")
engine_cfg_fcfs = EngineConfig(
    model="facebook/opt-125m",
    scheduler_config=sched_cfg_fcfs,
    cache_config={"num_gpu_blocks": 2, "block_size": 16}
)
engine_fcfs = LLMEngine(engine_config=engine_cfg_fcfs)

```

## Key Source Files

The vLLM scheduling system spans several core modules:

- **[`vllm/config/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/scheduler.py)**: Defines `SchedulerConfig` including the `policy` enum and token budget constraints
- **[`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py)**: Implements the main scheduling loop, preemption logic, and `_preempt_request` method
- **[`vllm/v1/core/sched/request_queue.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/request_queue.py)**: Contains `FCFSRequestQueue`, `PriorityRequestQueue`, and the `create_request_queue` factory
- **[`vllm/v1/request.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/request.py)**: Defines the `Request` dataclass with `priority`, `arrival_time`, and comparison operators for heap ordering
- **[`vllm/v1/core/kv_cache_manager.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/kv_cache_manager.py)**: Handles block allocation; allocation failures trigger the preemption cascade in the scheduler

## Summary

- **Policy Selection**: The vLLM scheduler supports `"fcfs"` (non-preemptive) and `"priority"` (preemptive) policies configured via `SchedulerConfig.policy` in [`vllm/config/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/scheduler.py)
- **Priority Ordering**: Under the priority policy, `PriorityRequestQueue` uses Python's `heapq` ordered by `Request.priority` (lower integer = higher priority) and `arrival_time` to break ties
- **Selective Preemption**: Preemption only occurs when `policy="priority"` and KV-cache allocation fails; the scheduler preempts the lowest-priority running request using `max()` selection on `(priority, arrival_time)`
- **Resource Reclamation**: The `_preempt_request` method in [`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py) frees cache blocks, marks requests as `PREEMPTED`, and re-queues them for later execution
- **FCFS Protection**: First-come-first-served mode prevents preemption entirely, ensuring predictable but potentially slower execution for high-priority late arrivals

## Frequently Asked Questions

### What is the difference between FCFS and priority scheduling in vLLM?

FCFS (First-Come-First-Served) processes requests in arrival order without preemption, meaning once a request begins running, it retains its KV-cache blocks until completion. Priority scheduling allows requests with lower integer priority values (e.g., -10) to preempt running requests with higher values (e.g., 0), forcibly freeing GPU memory for the urgent task. FCFS provides predictable latency for early requests, while priority scheduling optimizes for latency-sensitive workloads.

### How do I assign priority to a request in vLLM?

Pass the `priority` integer parameter when calling `engine.add_request()` or the equivalent API. Lower values indicate higher priority; for example, `priority=-10` executes before the default `priority=0`. The priority value is stored in the `Request` object defined in [`vllm/v1/request.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/request.py) and evaluated by the `PriorityRequestQueue` when determining scheduling order.

### Does priority scheduling impact throughput?

Priority scheduling can reduce overall throughput when frequent preemptions occur, because preempted requests must be re-computed from their last checkpoint when resumed. However, it significantly improves latency for high-priority requests by ensuring they are not blocked by long-running, low-priority batch jobs. The trade-off is acceptable when serving mixed workloads where certain queries (e.g., admin commands, real-time user interactions) require immediate attention over background processing.

### What happens to a request after it is preempted?

When preempted, the request's status changes to `PREEMPTED`, its allocated KV-cache blocks are freed via `kv_cache_manager`, and it is prepended to the waiting queue. The `num_preemptions` counter increments to track how many times the request has been interrupted. The request will resume execution when it reaches the front of the priority queue and sufficient GPU memory becomes available, though it may need to recompute tokens depending on the checkpointing strategy configured.