Guide to LoRA Fine-Tuning for LLMs: Efficient Low-Rank Adaptation Explained

LoRA fine-tuning for LLMs freezes base model weights and injects lightweight trainable low-rank matrices, reducing GPU memory requirements by up to 90% while maintaining downstream task performance.

This comprehensive guide examines the production-ready LoRA implementation found in the rohitg00/ai-engineering-from-scratch repository. The core logic resides in phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py, which provides a complete toolkit for parameter-efficient fine-tuning of large language models using PyTorch.

What is LoRA and Why Use It?

LoRA (Low-Rank Adaptation) revolutionizes LLM fine-tuning by avoiding full weight updates. Instead of modifying the original pre-trained matrices $W \in \mathbb{R}^{d \times k}$, LoRA introduces trainable decomposition matrices $A$ and $B$ such that the update $\Delta W = BA$ has rank $r \ll \min(d, k)$. This approach reduces trainable parameters from billions to millions, enabling fine-tuning on consumer GPUs while preserving the base model's general knowledge.

Core Implementation Components

The lora.py module implements four critical abstractions that handle the mathematical operations and model surgery required for LoRA fine-tuning.

LoRALayer Class

The LoRALayer class encapsulates the low-rank decomposition logic. According to the source code in phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py (lines 6-18), this module initializes:

  • Matrix A: Shape (in_features × rank) with Gaussian initialization
  • Matrix B: Shape (rank × out_features) initialized to zero
  • Scaling factor: alpha / r applied to the product $BA$

This design ensures the LoRA branch starts at zero and gradually learns task-specific adaptations without destabilizing the pre-trained model.

LinearWithLoRA Module

LinearWithLoRA wraps frozen nn.Linear layers and composes their outputs with the LoRA branch. As implemented in lines 20-33, the forward pass computes:

output = linear(x) + lora(x)

The base linear layer remains frozen with requires_grad=False, while only the LoRA matrices receive gradients during backpropagation.

inject_lora Function

The inject_lora utility performs model surgery by traversing the module hierarchy (lines 35-52). It accepts a list of target module names (e.g., ["0", "2"]) and replaces selected linear layers with LinearWithLoRA instances. After injection, all original parameters are frozen, leaving only the low-rank matrices trainable.

merge_lora_weights Function

After training completes, merge_lora_weights (lines 67-81) folds the learned adaptations back into the base weights using the formula:

W_merged = W_base + (alpha / r) * B @ A

This removes the LoRA modules entirely, restoring a vanilla nn.Sequential architecture optimized for inference speed.

QLoRA and 4-Bit Quantization Support

The implementation extends to QLoRA (Quantized LoRA) through the quantize_model helper and quantize_to_nf4/dequantize_from_nf4 utilities. This workflow quantizes frozen base weights to 4-bit NF4 (Normal Float 4) format, reducing memory footprint by approximately 75%, while keeping LoRA adapters in full precision (FP16/FP32) for stable training.

Complete LoRA Fine-Tuning Workflow

Follow this structured pipeline to fine-tune large language models efficiently:

  1. Initialize base model: Load your pre-trained transformer or create a demo architecture.
  2. Inject adapters: Call inject_lora() specifying target layers and hyperparameters (rank=8, alpha=16).
  3. Train adapters only: Use standard PyTorch optimizers; gradients flow only through the low-rank matrices.
  4. Optional quantization: Apply quantize_model() for QLoRA training on memory-constrained hardware.
  5. Merge and deploy: Run merge_lora_weights() to create a production-ready checkpoint without inference overhead.

Practical Implementation Examples

Injecting and Training LoRA Adapters

This example demonstrates injecting LoRA into specific layers of a feed-forward network and training only the adapter parameters:

import torch
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import (
    create_demo_model,
    create_demo_data,
    inject_lora,
    train_lora,
    count_parameters,
)

# 1️⃣ Build a base model

model = create_demo_model()
print("Base params:", count_parameters(model))

# 2️⃣ Inject LoRA into the first and third linear layers

lora_layers = inject_lora(model, target_modules=["0", "2"], rank=8, alpha=16)
print("LoRA injected into:", list(lora_layers.keys()))

# 3️⃣ Generate dummy data

data = create_demo_data()

# 4️⃣ Train only the LoRA adapters

losses = train_lora(model, data, epochs=5, lr=1e-3)
print("Training loss – first vs. last:", losses[0], losses[-1])

Merging Weights for Inference

After fine-tuning, merge the adapters back into the base model to eliminate runtime overhead:

from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import merge_lora_weights

# Assume `model` has been fine‑tuned with LoRA

merge_lora_weights(model)

# The model is now a plain nn.Sequential without any LoRA modules

print("After merge – trainable params:", count_parameters(model)["trainable"])

Saving and Loading Adapters

Persist only the LoRA matrices for efficient storage and rapid task switching:

from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import save_lora_adapter, load_lora_adapter

# Save the adapters

adapter_path = "lora_adapter.pt"
n_saved = save_lora_adapter(model, adapter_path)
print(f"Saved {n_saved} LoRA matrices")

# Load them into a fresh copy of the base model

new_model = create_demo_model()
inject_lora(new_model, target_modules=["0", "2"], rank=8, alpha=16)
load_lora_adapter(new_model, adapter_path)

Summary

  • LoRA fine-tuning for LLMs reduces trainable parameters by decomposing weight updates into low-rank matrices $A$ and $B$ rather than modifying full layers.
  • The implementation in phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py provides LoRALayer, LinearWithLoRA, and inject_lora for seamless model surgery.
  • QLoRA support via 4-bit NF4 quantization allows fine-tuning 7B+ parameter models on single consumer GPUs.
  • Adapter persistence through save_lora_adapter and load_lora_adapter enables multi-task serving from a single frozen backbone.
  • Weight merging with merge_lora_weights removes inference overhead by integrating adaptations back into base parameters.

Frequently Asked Questions

What is the difference between LoRA and full fine-tuning?

Full fine-tuning updates every parameter in the model, requiring massive GPU memory and storage for each task-specific variant. LoRA freezes the base weights and trains only small low-rank matrices, typically reducing trainable parameters by 90-99% while achieving comparable accuracy on downstream tasks.

How does QLoRA reduce memory usage compared to standard LoRA?

QLoRA quantizes the frozen base model weights to 4-bit precision using NF4 encoding, cutting memory usage by approximately 75%. The LoRA adapters remain in 16-bit or 32-bit precision for training stability, allowing you to fine-tune models like Llama-7B on GPUs with as little as 16GB VRAM.

Can I merge LoRA weights back into the base model after training?

Yes. The merge_lora_weights() function in lora.py mathematically combines the low-rank updates $(alpha/r) \times B \times A$ with the original frozen weights, producing a standard PyTorch model without LoRA modules. This eliminates inference latency and simplifies deployment pipelines.

How do I choose the optimal rank and alpha values for LoRA fine-tuning?

Rank ($r$) controls the expressiveness of the adaptation—typical values range from 4 to 64, with 8 being a balanced starting point. Alpha ($\alpha$) scales the learning rate of the LoRA updates; setting $\alpha = 2r$ is a common heuristic. Higher ranks capture more complex adaptations but increase memory usage and risk overfitting on small datasets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →