LoRA Adapter Techniques Unsloth Supports for Efficient Fine‑Tuning

Unsloth supports Standard LoRA, QLoRA (4‑bit), 16‑bit LoRA, Rank‑Stabilised LoRA (RSLoRA), and fast CUDA‑accelerated LoRA kernels for memory‑efficient fine‑tuning of LLMs.

The unslothai/unsloth library provides a flexible parameter-efficient fine-tuning (PEFT) stack that implements multiple LoRA adapter techniques. These methods allow you to train large language models on limited VRAM while maintaining training speed and model quality.

Standard LoRA and 16‑Bit Fine‑Tuning

Unsloth implements Standard LoRA through the get_peft_model method in unsloth/models/llama.py. This technique injects low‑rank trainable matrices (rank r) into selected linear layers while keeping the base model frozen.

To use 16‑bit precision as the base for LoRA adapters, pass load_in_16bit=True to FastLanguageModel.from_pretrained in unsloth/models/loader.py. This loads the model in bfloat16 or float16 precision, creating a middle ground between full‑precision training and quantized approaches.

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    "unsloth/llama-3.1-8b",
    load_in_16bit=True,  # 16‑bit base for standard LoRA

)

model = model.get_peft_model(
    r=32,
    lora_alpha=64,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)

QLoRA (4‑Bit LoRA) for Maximum Memory Efficiency

QLoRA quantizes the base model to 4 bits via bitsandbytes, then attaches LoRA adapters on top. This is the default mode when load_in_4bit=True is specified in unsloth/models/loader.py.

The source code explicitly disables full‑finetuning when 4‑bit quantization is active and emits a warning ("disabling LoRA / QLoRA") to prevent configuration conflicts. This combination dramatically reduces VRAM requirements while maintaining training capability.

model, tokenizer = FastLanguageModel.from_pretrained(
    "unsloth/llama-3.1-8b-bnb-4bit",
    load_in_4bit=True,  # Enables QLoRA (default for LoRA training)

)

model = model.get_peft_model(
    r=16,
    lora_alpha=32,
    target_modules=[
        "q_proj", "k_proj", "v_proj", "o_proj",
        "gate_proj", "up_proj", "down_proj",
    ],
)

Rank‑Stabilised LoRA (RSLoRA)

RSLoRA (Rank‑Stabilised LoRA) mitigates training divergence when using very low ranks. Unsloth exposes the use_rslora boolean flag in FastLanguageModel.get_peft_model (lines 2746‑2749 in unsloth/models/llama.py) and supports it through the CLI in unsloth_cli/cli.py.

When enabled, the underlying PEFT call receives use_rslora=True, applying the scaling factor 1/√r instead of the standard 1/r to stabilize gradients during training.

model = model.get_peft_model(
    r=8,
    lora_alpha=16,
    use_rslora=True,  # Enable rank‑stabilised LoRA

    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)

Fast LoRA Kernels for Accelerated Training

Unsloth replaces the generic LoRA forward pass with custom CUDA‑accelerated kernels that eliminate extra matrix multiplication overhead. The patch_fast_lora() function in unsloth/models/_utils.py (lines 58‑63) monkey‑patches peft.tuners.lora.bnb.Linear4bit.forward to point at fast_lora_forward defined in unsloth/kernels/fast_lora.py.

When fast_lora_forwards=True (the default), all LoRA layers use this optimized kernel for significant speed improvements during training.

from unsloth.models._utils import patch_fast_lora
patch_fast_lora()  # Automatically called during model preparation

Vision and Multimodal LoRA Support

Unsloth extends LoRA adapter techniques beyond text‑only models to vision and multimodal architectures. The LoRA target lists in model files such as unsloth/models/vision.py and unsloth/models/qwen2.py include vision‑specific projection names (e.g., Q‑proj layers and vision MLPs).

This allows the same get_peft_model API to apply LoRA to multimodal models like Qwen‑VL, enabling efficient fine‑tuning of vision‑language tasks without modifying the visual encoder weights directly.

Saving and Merging LoRA Adapters

After training, Unsloth provides flexible export options through save_pretrained in unsloth/save.py (lines 44‑55). The save_method parameter controls how LoRA weights are handled:

  • save_method="lora": Saves only the adapter weights (HuggingFace‑compatible)
  • save_method="merged_16bit": Merges LoRA weights into the base model and saves as 16‑bit (useful for GGUF conversion)
  • save_method="merged_4bit_forced": Force‑merges adapters into a 4‑bit quantized model (advanced use case)

# Save adapters only

model.save_pretrained("./my_lora_adapter", save_method="lora")

# Merge into 16‑bit for inference

model.save_pretrained("./my_merged_model", save_method="merged_16bit")

Summary

  • Standard LoRA works with 16‑bit base models via load_in_16bit=True in loader.py
  • QLoRA provides maximum memory efficiency through 4‑bit quantization (load_in_4bit=True)
  • RSLoRA stabilizes training for low ranks via the use_rslora flag in llama.py
  • Fast LoRA kernels accelerate training through CUDA optimizations in fast_lora.py
  • Multimodal support extends LoRA to vision layers in models like Qwen‑VL
  • Flexible export options allow saving adapters separately or merged via save_method in save.py

Frequently Asked Questions

What is the difference between LoRA and QLoRA in Unsloth?

Standard LoRA trains adapters on a 16‑bit (bfloat16/float16) base model, while QLoRA first quantizes the base model to 4 bits using bitsandbytes, then attaches adapters. QLoRA dramatically reduces VRAM usage—often by 50% or more—making it possible to fine‑tune 70B parameter models on single 48GB GPUs. According to the source in unsloth/models/loader.py, QLoRA is the default when load_in_4bit=True is specified.

When should I use Rank‑Stabilised LoRA (RSLoRA)?

Use RSLoRA when training with very low ranks (typically r < 8) to prevent gradient divergence. The use_rslora=True flag changes the scaling factor from 1/r to 1/√r, which stabilizes training dynamics. As implemented in unsloth/models/llama.py lines 2746‑2749, this flag is passed directly to the underlying PEFT library when constructing the adapter layers.

How do I export a LoRA model for GGUF or vLLM inference?

To prepare a LoRA‑trained model for GGUF quantization or vLLM serving, merge the adapters into the base weights using save_method="merged_16bit" in model.save_pretrained(). This folds the LoRA weights into the base model and saves a standard 16‑bit checkpoint, as implemented in unsloth/save.py around line 1999. For adapter‑only deployment, use save_method="lora" to generate HuggingFace‑compatible adapter weights.

Does Unsloth's fast LoRA kernel work with all quantization levels?

Yes, the fast LoRA kernel implemented in unsloth/kernels/fast_lora.py and patched via patch_fast_lora() in unsloth/models/_utils.py works across 4‑bit, 8‑bit, and 16‑bit configurations. When fast_lora_forwards=True (default), it replaces the standard PEFT forward pass with an optimized CUDA kernel that removes matmul overhead, providing speedups regardless of the base model's precision.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →