LoRA and QLoRA Fine-Tuning Implementation: From-Scratch PyTorch Guide
The rohitg00/ai-engineering-from-scratch repository implements Low-Rank Adaptation (LoRA) and Quantized LoRA (QLoRA) from first principles in pure PyTorch, enabling memory-efficient fine-tuning of large language models by training low-rank adapter matrices while freezing or quantizing base weights.
Efficient fine-tuning of billion-parameter models requires methods that minimize GPU memory without sacrificing performance. This article examines the LoRA and QLoRA fine-tuning implementation found in the open-source educational repository rohitg00/ai-engineering-from-scratch, specifically within phases/11_llm_engineering/08_fine_tuning_lora/code/lora.py. The codebase demonstrates how to inject trainable low-rank matrices into frozen linear layers, apply 4-bit Normal Float (NF4) quantization for QLoRA, and merge adapters back into dense weights for deployment.
Core LoRA Architecture Components
LoRALayer: Parameter-Efficient Weight Updates
The LoRALayer class represents the mathematical core of the implementation: a trainable low-rank decomposition that approximates full-rank weight updates. The layer maintains two matrices: A ∈ ℝ^(in × rank) initialized with values scaled by 1/√rank, and B ∈ ℝ^(rank × out) initialized to zeros. During the forward pass, the output is computed as scaling * (x @ A @ B), where the scaling factor equals α / r (alpha divided by rank). This initialization strategy ensures stable gradient flow at the start of training.
LinearWithLoRA: Wrapping Frozen Base Layers
The LinearWithLoRA wrapper combines a frozen base nn.Linear layer with the trainable LoRA adapter. The constructor sets requires_grad = False on all original linear parameters, ensuring only the low-rank matrices receive gradients. The forward implementation returns linear(x) + lora(x), seamlessly adding the adapter contribution to the frozen base output without modifying the underlying pre-trained weights.
Module Injection via inject_lora
The inject_lora function recursively traverses the model architecture using model.named_modules(). It checks each module name against the target_modules list (e.g., ["q_proj", "v_proj"]) and replaces matching nn.Linear layers with LinearWithLoRA instances. The function stores references to all created LoRA layers in a lora_layers dictionary, enabling later access for optimization or serialization.
QLoRA Memory Optimization Strategy
NF4 Block-wise Quantization
The quantization implementation converts frozen weights to 4-bit Normal Float format using block-wise scaling. The quantize_to_nf4 function splits tensors into 64-element blocks, computes a scaling factor for each block, and rounds values to int8 in the range [-8, 7]. The de-quantization routine restores the float tensor by reversing this scaling operation, allowing the model to maintain computation precision while storing weights at 4-bit resolution.
quantize_model Implementation
The quantize_model function applies NF4 quantization to all non-trainable parameters with dimensionality ≥ 2. It iterates over model.named_parameters() and replaces each frozen weight with its de-quantized version, keeping the model usable during training while drastically reducing memory footprint. This approach enables fine-tuning massive models on consumer GPUs by compressing the base weights while keeping LoRA adapters in full float32 or float16 precision.
Training Pipeline and Adapter Management
Selective Gradient Optimization
The train_lora function implements a parameter-efficient training loop that optimizes only the adapter weights. It constructs an AdamW optimizer over the subset p for p in model.parameters() if p.requires_grad, ensuring no gradients flow into the frozen base model or quantized weights. This selective optimization reduces memory usage for optimizer states by 90% or more compared to full fine-tuning.
Saving and Loading Adapters
The implementation provides save_lora_adapter and load_lora_adapter functions for model checkpointing. These utilities serialize the A and B matrices along with rank and alpha hyperparameters to a .pt file. This modular approach allows distributing small adapter files (megabytes) separately from massive base models (gigabytes), facilitating easy sharing and deployment of fine-tuned behaviors.
Merging Adapters for Production Inference
After training completes, the merge_lora_weights function folds the low-rank updates back into the base weight matrices to eliminate inference overhead. The implementation computes merged = (A @ B) * scaling, transposes the result to match PyTorch's [out_features, in_features] convention, and adds it directly to the original linear layer's weight tensor. Once merged, the LinearWithLoRA wrappers can be removed, yielding a standard dense model with no runtime penalty from adapter computation.
Complete Workflow Implementation
The following examples demonstrate the end-to-end usage of the LoRA and QLoRA implementation:
Inject LoRA into specific attention layers:
import torch.nn as nn
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import inject_lora
model = ... # Your nn.Module with Linear layers
target_modules = ["q_proj", "v_proj"]
lora_modules = inject_lora(model, target_modules, rank=8, alpha=16)
Apply NF4 quantization for QLoRA training:
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import quantize_model
quant_state = quantize_model(model) # Compress frozen weights to 4-bit
Train only the adapter parameters:
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import train_lora, create_demo_data
data = create_demo_data()
losses = train_lora(model, data, epochs=20, lr=1e-3, batch_size=8)
Merge adapters and save for deployment:
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import merge_lora_weights, save_lora_adapter
merge_lora_weights(model) # Fold adapters into base weights
save_lora_adapter(model, "adapter.pt") # Optional: save for later use
Summary
- The LoRA and QLoRA fine-tuning implementation in
rohitg00/ai-engineering-from-scratchprovides a complete educational framework for parameter-efficient fine-tuning without external dependencies like PEFT. - LoRALayer uses rank-decomposed matrices with scaled initialization to approximate weight updates, while LinearWithLoRA freezes base parameters and adds adapter contributions during the forward pass.
- QLoRA reduces memory consumption by quantizing frozen weights to NF4 4-bit precision using 64-element block-wise scaling, keeping only the small adapter matrices in full precision.
- The inject_lora function enables surgical targeting of specific layers (e.g., query and value projections), and merge_lora_weights eliminates inference latency by folding adapters back into the base model.
Frequently Asked Questions
What is the difference between LoRA and QLoRA in this implementation?
LoRA trains low-rank adapter matrices while keeping base weights in full precision, reducing trainable parameters by roughly 99%. QLoRA adds a quantization step where the frozen base weights are compressed to 4-bit NF4 format using the quantize_model function, cutting memory usage by an additional 75% while maintaining training stability through de-quantization during the forward pass.
How does the NF4 quantization reduce memory usage?
The quantize_to_nf4 function compresses each 64-element block of a weight tensor into 4-bit integers (range -8 to 7) plus a block-wise scaling factor. This reduces storage from 32-bit floats to 4-bit representations, shrinking the memory footprint of the frozen base model from approximately 4 bytes per parameter to 0.5 bytes per parameter, enabling fine-tuning of 7B+ models on single consumer GPUs.
Can I merge the LoRA weights back into the base model?
Yes. The merge_lora_weights function computes the full-rank update (A @ B) * scaling and adds it directly to the original linear layer's weight tensor. After merging, the model contains only standard nn.Linear layers with updated weights, eliminating the computational overhead of separate adapter modules during inference.
Which layers should I target with inject_lora?
The target_modules parameter accepts substring matches against layer names. Common targets include ["q_proj", "v_proj"] for attention query and value projections, or ["q_proj", "k_proj", "v_proj", "o_proj"] for all attention components. The implementation automatically skips non-linear layers and embeddings, applying adapters only to matching nn.Linear modules.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →