How to Use LoRA Adapters and Multi-LoRA in vLLM for Efficient Fine-Tuning

vLLM enables efficient serving of multiple LoRA adapters by loading them into an LRU cache at runtime, allowing you to switch between fine-tuned models without reloading the base weights or restarting the server.

vLLM's LoRA support allows you to attach lightweight adapter weights to a base model at inference time without retraining the entire model. This guide explains how to use LoRA adapters and multi-LoRA in vLLM for efficient fine-tuning, covering both the Python SDK and server deployments. Whether you need single-adapter inference or concurrent serving of multiple adapters, vLLM's architecture minimizes per-request overhead through optimized kernel implementations and intelligent caching.

Understanding vLLM's LoRA Architecture

Core Components

vLLM implements LoRA support through three primary abstractions defined in the source code:

  • LoRAConfig (vllm/config/lora.py): Global configuration specifying max_loras (cache size), max_lora_rank (adapter dimension), and sharding options. When max_lora_rank exceeds 8, vLLM automatically enables fully_sharded_loras for tensor parallelism.

  • LoRARequest (vllm/lora/request.py): Per-request descriptor containing the adapter name, a globally unique integer ID (lora_int_id ≥ 1), and the filesystem path to the adapter weights. Equality and hashing are based on lora_name, enabling automatic deduplication across requests.

  • LoRAModelRunnerMixin (vllm/v1/worker/lora_model_runner_mixin.py): Mix-in class that equips GPU and TPU model runners with LoRA capabilities. It manages the LRUCacheWorkerLoRAManager, constructs LoRAMapping objects that map tokens to adapter IDs, and handles dummy-run warm-up for CUDA graph capture.

The LRU Cache Manager

When LoRA is enabled (--enable-lora), the model runner instantiates an LRUCacheWorkerLoRAManager (vllm/lora/worker_manager.py). This manager maintains a bounded cache of adapter weights (size determined by max_loras) and implements the following critical functions:

  • add_adapter: Loads adapter weights from disk and pins them in GPU memory.
  • set_active_adapters: Builds a LoRAMapping object containing prompt_lora_mapping and token_lora_mapping arrays that tell the CUDA kernels which adapter weights to apply for each token.
  • Eviction: When the cache exceeds max_loras, the least-recently-used adapter is unloaded to free memory.

The LoRAMapping structure is defined in vllm/lora/layers.py and is passed directly

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →