# How to Use LoRA Adapters and Multi-LoRA in vLLM for Efficient Fine-Tuning

> Discover how to efficiently fine-tune models with LoRA adapters and Multi-LoRA in vLLM. Load adapters at runtime without reloading weights for faster deployment.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**vLLM enables efficient serving of multiple LoRA adapters by loading them into an LRU cache at runtime, allowing you to switch between fine-tuned models without reloading the base weights or restarting the server.**

vLLM's LoRA support allows you to attach lightweight adapter weights to a base model at inference time without retraining the entire model. This guide explains how to use LoRA adapters and multi-LoRA in vLLM for efficient fine-tuning, covering both the Python SDK and server deployments. Whether you need single-adapter inference or concurrent serving of multiple adapters, vLLM's architecture minimizes per-request overhead through optimized kernel implementations and intelligent caching.

## Understanding vLLM's LoRA Architecture

### Core Components

vLLM implements LoRA support through three primary abstractions defined in the source code:

- **`LoRAConfig`** ([`vllm/config/lora.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/lora.py)): Global configuration specifying `max_loras` (cache size), `max_lora_rank` (adapter dimension), and sharding options. When `max_lora_rank` exceeds 8, vLLM automatically enables `fully_sharded_loras` for tensor parallelism.

- **`LoRARequest`** ([`vllm/lora/request.py`](https://github.com/vllm-project/vllm/blob/main/vllm/lora/request.py)): Per-request descriptor containing the adapter name, a globally unique integer ID (`lora_int_id ≥ 1`), and the filesystem path to the adapter weights. Equality and hashing are based on `lora_name`, enabling automatic deduplication across requests.

- **`LoRAModelRunnerMixin`** ([`vllm/v1/worker/lora_model_runner_mixin.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/lora_model_runner_mixin.py)): Mix-in class that equips GPU and TPU model runners with LoRA capabilities. It manages the `LRUCacheWorkerLoRAManager`, constructs `LoRAMapping` objects that map tokens to adapter IDs, and handles dummy-run warm-up for CUDA graph capture.

### The LRU Cache Manager

When LoRA is enabled (`--enable-lora`), the model runner instantiates an **`LRUCacheWorkerLoRAManager`** ([`vllm/lora/worker_manager.py`](https://github.com/vllm-project/vllm/blob/main/vllm/lora/worker_manager.py)). This manager maintains a bounded cache of adapter weights (size determined by `max_loras`) and implements the following critical functions:

- **`add_adapter`**: Loads adapter weights from disk and pins them in GPU memory.
- **`set_active_adapters`**: Builds a **`LoRAMapping`** object containing `prompt_lora_mapping` and `token_lora_mapping` arrays that tell the CUDA kernels which adapter weights to apply for each token.
- **Eviction**: When the cache exceeds `max_loras`, the least-recently-used adapter is unloaded to free memory.

The `LoRAMapping` structure is defined in [`vllm/lora/layers.py`](https://github.com/vllm-project/vllm/blob/main/vllm/lora/layers.py) and is passed directly