Guide to LoRA Fine-Tuning for LLMs: Efficient Low-Rank Adaptation Explained
LoRA fine-tuning for LLMs freezes base model weights and injects lightweight trainable low-rank matrices, reducing GPU memory requirements by up to 90% while maintaining downstream task performance.
This comprehensive guide examines the production-ready LoRA implementation found in the rohitg00/ai-engineering-from-scratch repository. The core logic resides in phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py, which provides a complete toolkit for parameter-efficient fine-tuning of large language models using PyTorch.
What is LoRA and Why Use It?
LoRA (Low-Rank Adaptation) revolutionizes LLM fine-tuning by avoiding full weight updates. Instead of modifying the original pre-trained matrices $W \in \mathbb{R}^{d \times k}$, LoRA introduces trainable decomposition matrices $A$ and $B$ such that the update $\Delta W = BA$ has rank $r \ll \min(d, k)$. This approach reduces trainable parameters from billions to millions, enabling fine-tuning on consumer GPUs while preserving the base model's general knowledge.
Core Implementation Components
The lora.py module implements four critical abstractions that handle the mathematical operations and model surgery required for LoRA fine-tuning.
LoRALayer Class
The LoRALayer class encapsulates the low-rank decomposition logic. According to the source code in phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py (lines 6-18), this module initializes:
- Matrix A: Shape
(in_features × rank)with Gaussian initialization - Matrix B: Shape
(rank × out_features)initialized to zero - Scaling factor:
alpha / rapplied to the product $BA$
This design ensures the LoRA branch starts at zero and gradually learns task-specific adaptations without destabilizing the pre-trained model.
LinearWithLoRA Module
LinearWithLoRA wraps frozen nn.Linear layers and composes their outputs with the LoRA branch. As implemented in lines 20-33, the forward pass computes:
output = linear(x) + lora(x)
The base linear layer remains frozen with requires_grad=False, while only the LoRA matrices receive gradients during backpropagation.
inject_lora Function
The inject_lora utility performs model surgery by traversing the module hierarchy (lines 35-52). It accepts a list of target module names (e.g., ["0", "2"]) and replaces selected linear layers with LinearWithLoRA instances. After injection, all original parameters are frozen, leaving only the low-rank matrices trainable.
merge_lora_weights Function
After training completes, merge_lora_weights (lines 67-81) folds the learned adaptations back into the base weights using the formula:
W_merged = W_base + (alpha / r) * B @ A
This removes the LoRA modules entirely, restoring a vanilla nn.Sequential architecture optimized for inference speed.
QLoRA and 4-Bit Quantization Support
The implementation extends to QLoRA (Quantized LoRA) through the quantize_model helper and quantize_to_nf4/dequantize_from_nf4 utilities. This workflow quantizes frozen base weights to 4-bit NF4 (Normal Float 4) format, reducing memory footprint by approximately 75%, while keeping LoRA adapters in full precision (FP16/FP32) for stable training.
Complete LoRA Fine-Tuning Workflow
Follow this structured pipeline to fine-tune large language models efficiently:
- Initialize base model: Load your pre-trained transformer or create a demo architecture.
- Inject adapters: Call
inject_lora()specifying target layers and hyperparameters (rank=8,alpha=16). - Train adapters only: Use standard PyTorch optimizers; gradients flow only through the low-rank matrices.
- Optional quantization: Apply
quantize_model()for QLoRA training on memory-constrained hardware. - Merge and deploy: Run
merge_lora_weights()to create a production-ready checkpoint without inference overhead.
Practical Implementation Examples
Injecting and Training LoRA Adapters
This example demonstrates injecting LoRA into specific layers of a feed-forward network and training only the adapter parameters:
import torch
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import (
create_demo_model,
create_demo_data,
inject_lora,
train_lora,
count_parameters,
)
# 1️⃣ Build a base model
model = create_demo_model()
print("Base params:", count_parameters(model))
# 2️⃣ Inject LoRA into the first and third linear layers
lora_layers = inject_lora(model, target_modules=["0", "2"], rank=8, alpha=16)
print("LoRA injected into:", list(lora_layers.keys()))
# 3️⃣ Generate dummy data
data = create_demo_data()
# 4️⃣ Train only the LoRA adapters
losses = train_lora(model, data, epochs=5, lr=1e-3)
print("Training loss – first vs. last:", losses[0], losses[-1])
Merging Weights for Inference
After fine-tuning, merge the adapters back into the base model to eliminate runtime overhead:
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import merge_lora_weights
# Assume `model` has been fine‑tuned with LoRA
merge_lora_weights(model)
# The model is now a plain nn.Sequential without any LoRA modules
print("After merge – trainable params:", count_parameters(model)["trainable"])
Saving and Loading Adapters
Persist only the LoRA matrices for efficient storage and rapid task switching:
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import save_lora_adapter, load_lora_adapter
# Save the adapters
adapter_path = "lora_adapter.pt"
n_saved = save_lora_adapter(model, adapter_path)
print(f"Saved {n_saved} LoRA matrices")
# Load them into a fresh copy of the base model
new_model = create_demo_model()
inject_lora(new_model, target_modules=["0", "2"], rank=8, alpha=16)
load_lora_adapter(new_model, adapter_path)
Summary
- LoRA fine-tuning for LLMs reduces trainable parameters by decomposing weight updates into low-rank matrices $A$ and $B$ rather than modifying full layers.
- The implementation in
phases/11-llm-engineering/08-fine-tuning-lora/code/lora.pyprovidesLoRALayer,LinearWithLoRA, andinject_lorafor seamless model surgery. - QLoRA support via 4-bit NF4 quantization allows fine-tuning 7B+ parameter models on single consumer GPUs.
- Adapter persistence through
save_lora_adapterandload_lora_adapterenables multi-task serving from a single frozen backbone. - Weight merging with
merge_lora_weightsremoves inference overhead by integrating adaptations back into base parameters.
Frequently Asked Questions
What is the difference between LoRA and full fine-tuning?
Full fine-tuning updates every parameter in the model, requiring massive GPU memory and storage for each task-specific variant. LoRA freezes the base weights and trains only small low-rank matrices, typically reducing trainable parameters by 90-99% while achieving comparable accuracy on downstream tasks.
How does QLoRA reduce memory usage compared to standard LoRA?
QLoRA quantizes the frozen base model weights to 4-bit precision using NF4 encoding, cutting memory usage by approximately 75%. The LoRA adapters remain in 16-bit or 32-bit precision for training stability, allowing you to fine-tune models like Llama-7B on GPUs with as little as 16GB VRAM.
Can I merge LoRA weights back into the base model after training?
Yes. The merge_lora_weights() function in lora.py mathematically combines the low-rank updates $(alpha/r) \times B \times A$ with the original frozen weights, producing a standard PyTorch model without LoRA modules. This eliminates inference latency and simplifies deployment pipelines.
How do I choose the optimal rank and alpha values for LoRA fine-tuning?
Rank ($r$) controls the expressiveness of the adaptation—typical values range from 4 to 64, with 8 being a balanced starting point. Alpha ($\alpha$) scales the learning rate of the LoRA updates; setting $\alpha = 2r$ is a common heuristic. Higher ranks capture more complex adaptations but increase memory usage and risk overfitting on small datasets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →