# Guide to LoRA Fine-Tuning for LLMs: Efficient Low-Rank Adaptation Explained

> Discover LoRA fine-tuning for LLMs. Freeze base weights and inject low-rank matrices to cut GPU memory by 90% while preserving performance. Learn efficient adaptation.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-19

---

**LoRA fine-tuning for LLMs freezes base model weights and injects lightweight trainable low-rank matrices, reducing GPU memory requirements by up to 90% while maintaining downstream task performance.**

This comprehensive guide examines the production-ready LoRA implementation found in the `rohitg00/ai-engineering-from-scratch` repository. The core logic resides in [`phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py), which provides a complete toolkit for parameter-efficient fine-tuning of large language models using PyTorch.

## What is LoRA and Why Use It?

**LoRA (Low-Rank Adaptation)** revolutionizes LLM fine-tuning by avoiding full weight updates. Instead of modifying the original pre-trained matrices $W \in \mathbb{R}^{d \times k}$, LoRA introduces trainable decomposition matrices $A$ and $B$ such that the update $\Delta W = BA$ has rank $r \ll \min(d, k)$. This approach reduces trainable parameters from billions to millions, enabling fine-tuning on consumer GPUs while preserving the base model's general knowledge.

## Core Implementation Components

The [`lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/lora.py) module implements four critical abstractions that handle the mathematical operations and model surgery required for LoRA fine-tuning.

### LoRALayer Class

The `LoRALayer` class encapsulates the low-rank decomposition logic. According to the source code in [`phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py) (lines 6-18), this module initializes:

- **Matrix A**: Shape `(in_features × rank)` with Gaussian initialization
- **Matrix B**: Shape `(rank × out_features)` initialized to zero
- **Scaling factor**: `alpha / r` applied to the product $BA$

This design ensures the LoRA branch starts at zero and gradually learns task-specific adaptations without destabilizing the pre-trained model.

### LinearWithLoRA Module

`LinearWithLoRA` wraps frozen `nn.Linear` layers and composes their outputs with the LoRA branch. As implemented in lines 20-33, the forward pass computes:

```python
output = linear(x) + lora(x)

```

The base linear layer remains frozen with `requires_grad=False`, while only the LoRA matrices receive gradients during backpropagation.

### inject_lora Function

The `inject_lora` utility performs model surgery by traversing the module hierarchy (lines 35-52). It accepts a list of target module names (e.g., `["0", "2"]`) and replaces selected linear layers with `LinearWithLoRA` instances. After injection, all original parameters are frozen, leaving only the low-rank matrices trainable.

### merge_lora_weights Function

After training completes, `merge_lora_weights` (lines 67-81) folds the learned adaptations back into the base weights using the formula:

```python
W_merged = W_base + (alpha / r) * B @ A

```

This removes the LoRA modules entirely, restoring a vanilla `nn.Sequential` architecture optimized for inference speed.

## QLoRA and 4-Bit Quantization Support

The implementation extends to **QLoRA** (Quantized LoRA) through the `quantize_model` helper and `quantize_to_nf4`/`dequantize_from_nf4` utilities. This workflow quantizes frozen base weights to 4-bit NF4 (Normal Float 4) format, reducing memory footprint by approximately 75%, while keeping LoRA adapters in full precision (FP16/FP32) for stable training.

## Complete LoRA Fine-Tuning Workflow

Follow this structured pipeline to fine-tune large language models efficiently:

1. **Initialize base model**: Load your pre-trained transformer or create a demo architecture.
2. **Inject adapters**: Call `inject_lora()` specifying target layers and hyperparameters (`rank=8`, `alpha=16`).
3. **Train adapters only**: Use standard PyTorch optimizers; gradients flow only through the low-rank matrices.
4. **Optional quantization**: Apply `quantize_model()` for QLoRA training on memory-constrained hardware.
5. **Merge and deploy**: Run `merge_lora_weights()` to create a production-ready checkpoint without inference overhead.

## Practical Implementation Examples

### Injecting and Training LoRA Adapters

This example demonstrates injecting LoRA into specific layers of a feed-forward network and training only the adapter parameters:

```python
import torch
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import (
    create_demo_model,
    create_demo_data,
    inject_lora,
    train_lora,
    count_parameters,
)

# 1️⃣ Build a base model

model = create_demo_model()
print("Base params:", count_parameters(model))

# 2️⃣ Inject LoRA into the first and third linear layers

lora_layers = inject_lora(model, target_modules=["0", "2"], rank=8, alpha=16)
print("LoRA injected into:", list(lora_layers.keys()))

# 3️⃣ Generate dummy data

data = create_demo_data()

# 4️⃣ Train only the LoRA adapters

losses = train_lora(model, data, epochs=5, lr=1e-3)
print("Training loss – first vs. last:", losses[0], losses[-1])

```

### Merging Weights for Inference

After fine-tuning, merge the adapters back into the base model to eliminate runtime overhead:

```python
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import merge_lora_weights

# Assume `model` has been fine‑tuned with LoRA

merge_lora_weights(model)

# The model is now a plain nn.Sequential without any LoRA modules

print("After merge – trainable params:", count_parameters(model)["trainable"])

```

### Saving and Loading Adapters

Persist only the LoRA matrices for efficient storage and rapid task switching:

```python
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import save_lora_adapter, load_lora_adapter

# Save the adapters

adapter_path = "lora_adapter.pt"
n_saved = save_lora_adapter(model, adapter_path)
print(f"Saved {n_saved} LoRA matrices")

# Load them into a fresh copy of the base model

new_model = create_demo_model()
inject_lora(new_model, target_modules=["0", "2"], rank=8, alpha=16)
load_lora_adapter(new_model, adapter_path)

```

## Summary

- **LoRA fine-tuning for LLMs** reduces trainable parameters by decomposing weight updates into low-rank matrices $A$ and $B$ rather than modifying full layers.
- The implementation in [`phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py) provides `LoRALayer`, `LinearWithLoRA`, and `inject_lora` for seamless model surgery.
- **QLoRA support** via 4-bit NF4 quantization allows fine-tuning 7B+ parameter models on single consumer GPUs.
- **Adapter persistence** through `save_lora_adapter` and `load_lora_adapter` enables multi-task serving from a single frozen backbone.
- **Weight merging** with `merge_lora_weights` removes inference overhead by integrating adaptations back into base parameters.

## Frequently Asked Questions

### What is the difference between LoRA and full fine-tuning?

Full fine-tuning updates every parameter in the model, requiring massive GPU memory and storage for each task-specific variant. **LoRA freezes the base weights** and trains only small low-rank matrices, typically reducing trainable parameters by 90-99% while achieving comparable accuracy on downstream tasks.

### How does QLoRA reduce memory usage compared to standard LoRA?

QLoRA quantizes the frozen base model weights to 4-bit precision using NF4 encoding, cutting memory usage by approximately 75%. The LoRA adapters remain in 16-bit or 32-bit precision for training stability, allowing you to fine-tune models like Llama-7B on GPUs with as little as 16GB VRAM.

### Can I merge LoRA weights back into the base model after training?

Yes. The `merge_lora_weights()` function in [`lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/lora.py) mathematically combines the low-rank updates $(alpha/r) \times B \times A$ with the original frozen weights, producing a standard PyTorch model without LoRA modules. This eliminates inference latency and simplifies deployment pipelines.

### How do I choose the optimal rank and alpha values for LoRA fine-tuning?

**Rank ($r$)** controls the expressiveness of the adaptation—typical values range from 4 to 64, with 8 being a balanced starting point. **Alpha ($\alpha$)** scales the learning rate of the LoRA updates; setting $\alpha = 2r$ is a common heuristic. Higher ranks capture more complex adaptations but increase memory usage and risk overfitting on small datasets.