Implementing Parameter-Efficient Finetuning with LoRA from Scratch: A Complete Guide
LoRA reduces trainable parameters by freezing pre-trained weights and injecting small, trainable low-rank matrices into linear layers, enabling efficient fine-tuning of large language models with minimal memory overhead.
The rasbt/LLMs-from-scratch repository provides a transparent, educational implementation of parameter-efficient finetuning with LoRA that works with any PyTorch model built on torch.nn.Linear. This guide examines the core mechanics implemented in appendix_e.py, demonstrating how to retrofit existing architectures for low-cost adaptation.
How LoRA Reduces Trainable Parameters
LoRA (Low-Rank Adaptation) reparameterizes weight updates using a low-rank decomposition. For any frozen weight matrix W with shape (out_dim, in_dim), the method learns two smaller matrices:
- A: shape
(in_dim, r) - B: shape
(r, out_dim)
The forward pass becomes:
y = x @ W.T + (alpha / r) * x @ A @ B
Where r is the rank (typically 4-64) and alpha is a scaling factor. This reduces trainable parameters from in_dim × out_dim to r × (in_dim + out_dim), often yielding >99% parameter reduction in large models.
Core LoRA Components in appendix_e.py
The implementation in pkg/llms_from_scratch/appendix_e.py provides three essential building blocks.
LoRALayer: Low-Rank Matrix Decomposition
The LoRALayer class (lines 10-23) holds the trainable matrices A and B, initializing A with Kaiming uniform initialization and B with zeros to ensure training starts from the original frozen weights.
LinearWithLoRA: Wrapping torch.nn.Linear
The LinearWithLoRA class (lines 25-35) wraps an existing torch.nn.Linear module. During the forward pass, it computes the original frozen linear output plus the scaled LoRA contribution (alpha/r) * xAB.
replace_linear_with_lora: Recursive Model Conversion
The replace_linear_with_lora function (lines 37-44) recursively traverses a model's module tree, replacing every Linear instance with the LinearWithLoRA wrapper while preserving the frozen original weights.
Step-by-Step Implementation Guide
Follow these steps to convert a pre-trained model for parameter-efficient finetuning.
1. Import the LoRA Utilities
from llms_from_scratch.appendix_e import LinearWithLoRA, replace_linear_with_lora
import torch
2. Freeze Pre-trained Weights and Apply LoRA
Load your model and freeze all existing parameters before injecting LoRA layers:
# Example using a model with nn.Linear layers
model = YourModel.from_pretrained("checkpoint")
model.eval()
# Freeze all base parameters
for param in model.parameters():
param.requires_grad = False
# Inject LoRA with rank=16 and alpha=16 into all Linear layers
replace_linear_with_lora(model, rank=16, alpha=16)
3. Configure the Optimizer for LoRA Parameters Only
Filter the optimizer to update only the low-rank matrices:
trainable_params = [p for p in model.parameters() if p.requires_grad]
optimizer = torch.optim.AdamW(trainable_params, lr=1e-4)
# Standard training loop
for batch in dataloader:
optimizer.zero_grad()
loss = model(batch["input_ids"], labels=batch["labels"]).loss
loss.backward()
optimizer.step()
4. Verify Trainable Parameter Count
Confirm the parameter efficiency of your setup:
total_trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total_all = sum(p.numel() for p in model.parameters())
print(f"Trainable params: {total_trainable:,} ({100 * total_trainable / total_all:.2f}%)")
5. Optional: Merge LoRA Weights for Inference
For production inference, merge the LoRA contribution back into the original weight matrix to eliminate the separate low-rank computation. The additional_experiments.py file demonstrates this merged implementation:
# Conceptual merged computation
merged_weight = linear.weight + (alpha * (A @ B)).T
output = x @ merged_weight.T + linear.bias
Practical Training Examples
The repository includes complete training pipelines demonstrating parameter-efficient finetuning in practice.
-
ch07/01_main-chapter-code/exercise_experiments.py: Contains an end-to-end fine-tuning script with command-line flags (--lora) enabling LoRA training on instruction-following tasks. -
ch06/02_bonus_additional-experiments/additional_experiments.py: Implements an alternativeLinearWithLoRAMergedclass that folds LoRA weights into the base linear layer for optimized inference speed. -
pkg/llms_from_scratch/tests/test_appendix_e.py: Validates that LoRA parameter counts match expectations and that gradients flow correctly through the low-rank pathway.
Summary
- LoRA implements parameter-efficient finetuning by freezing original weights and learning low-rank updates via matrices A and B.
- The
rasbt/LLMs-from-scratchimplementation inappendix_e.pyprovidesLoRALayer,LinearWithLoRA, andreplace_linear_with_lorafor drop-in compatibility with anynn.Linear-based architecture. - Only
r × (in_dim + out_dim)parameters become trainable instead ofin_dim × out_dim, reducing memory requirements by orders of magnitude. - The
exercise_experiments.pyscript demonstrates production-ready training, whileadditional_experiments.pyshows how to merge weights for efficient inference.
Frequently Asked Questions
What is the difference between rank and alpha in LoRA?
Rank (r) controls the dimensionality of the low-rank matrices A and B, directly impacting the number of trainable parameters and model capacity. Alpha (α) is a scaling factor applied as alpha/r in the forward pass to stabilize training; higher alpha values increase the influence of the LoRA adaptation relative to the frozen pre-trained weights.
Can I apply LoRA to specific layers only?
Yes. While replace_linear_with_lora targets all Linear layers by default, you can modify the recursive function (lines 37-44 in appendix_e.py) to filter by layer name, module type, or depth in the transformer stack—for example, applying LoRA only to attention query and value projections while excluding feed-forward layers.
How do I merge LoRA weights back into the original model?
As shown in additional_experiments.py, you can compute merged_weight = W + (alpha/r) * B.T @ A.T and assign it back to the original linear layer's weight tensor. This eliminates the separate matrix multiplication during inference, reducing latency while preserving the fine-tuned behavior.
What memory savings does LoRA provide?
LoRA reduces optimizer memory (which tracks momentum and variance for each parameter) proportionally to the parameter reduction. For a 7B parameter model with rank-16 LoRA applied to attention layers, you typically train <100M parameters instead of 7B, reducing optimizer states from ~56GB to under 1GB when using AdamW (8 bytes per parameter).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →