# LoRA and QLoRA Fine-Tuning Implementation: From-Scratch PyTorch Guide

> Learn LoRA and QLoRA fine-tuning from scratch with PyTorch. This guide details memory-efficient LLM adaptation by training low-rank matrices, freezing or quantizing base weights.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-25

---

**The `rohitg00/ai-engineering-from-scratch` repository implements Low-Rank Adaptation (LoRA) and Quantized LoRA (QLoRA) from first principles in pure PyTorch, enabling memory-efficient fine-tuning of large language models by training low-rank adapter matrices while freezing or quantizing base weights.**

Efficient fine-tuning of billion-parameter models requires methods that minimize GPU memory without sacrificing performance. This article examines the **LoRA and QLoRA fine-tuning implementation** found in the open-source educational repository `rohitg00/ai-engineering-from-scratch`, specifically within [`phases/11_llm_engineering/08_fine_tuning_lora/code/lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11_llm_engineering/08_fine_tuning_lora/code/lora.py). The codebase demonstrates how to inject trainable low-rank matrices into frozen linear layers, apply 4-bit Normal Float (NF4) quantization for QLoRA, and merge adapters back into dense weights for deployment.

## Core LoRA Architecture Components

### LoRALayer: Parameter-Efficient Weight Updates

The `LoRALayer` class represents the mathematical core of the implementation: a trainable low-rank decomposition that approximates full-rank weight updates. The layer maintains two matrices: **A** ∈ ℝ^(in × rank) initialized with values scaled by `1/√rank`, and **B** ∈ ℝ^(rank × out) initialized to zeros. During the forward pass, the output is computed as `scaling * (x @ A @ B)`, where the scaling factor equals `α / r` (alpha divided by rank). This initialization strategy ensures stable gradient flow at the start of training.

### LinearWithLoRA: Wrapping Frozen Base Layers

The `LinearWithLoRA` wrapper combines a frozen base `nn.Linear` layer with the trainable LoRA adapter. The constructor sets `requires_grad = False` on all original linear parameters, ensuring only the low-rank matrices receive gradients. The forward implementation returns `linear(x) + lora(x)`, seamlessly adding the adapter contribution to the frozen base output without modifying the underlying pre-trained weights.

### Module Injection via inject_lora

The `inject_lora` function recursively traverses the model architecture using `model.named_modules()`. It checks each module name against the `target_modules` list (e.g., `["q_proj", "v_proj"]`) and replaces matching `nn.Linear` layers with `LinearWithLoRA` instances. The function stores references to all created LoRA layers in a `lora_layers` dictionary, enabling later access for optimization or serialization.

## QLoRA Memory Optimization Strategy

### NF4 Block-wise Quantization

The quantization implementation converts frozen weights to 4-bit Normal Float format using block-wise scaling. The `quantize_to_nf4` function splits tensors into 64-element blocks, computes a scaling factor for each block, and rounds values to int8 in the range `[-8, 7]`. The de-quantization routine restores the float tensor by reversing this scaling operation, allowing the model to maintain computation precision while storing weights at 4-bit resolution.

### quantize_model Implementation

The `quantize_model` function applies NF4 quantization to all non-trainable parameters with dimensionality ≥ 2. It iterates over `model.named_parameters()` and replaces each frozen weight with its de-quantized version, keeping the model usable during training while drastically reducing memory footprint. This approach enables fine-tuning massive models on consumer GPUs by compressing the base weights while keeping LoRA adapters in full `float32` or `float16` precision.

## Training Pipeline and Adapter Management

### Selective Gradient Optimization

The `train_lora` function implements a parameter-efficient training loop that optimizes only the adapter weights. It constructs an `AdamW` optimizer over the subset `p for p in model.parameters() if p.requires_grad`, ensuring no gradients flow into the frozen base model or quantized weights. This selective optimization reduces memory usage for optimizer states by 90% or more compared to full fine-tuning.

### Saving and Loading Adapters

The implementation provides `save_lora_adapter` and `load_lora_adapter` functions for model checkpointing. These utilities serialize the `A` and `B` matrices along with `rank` and `alpha` hyperparameters to a `.pt` file. This modular approach allows distributing small adapter files (megabytes) separately from massive base models (gigabytes), facilitating easy sharing and deployment of fine-tuned behaviors.

## Merging Adapters for Production Inference

After training completes, the `merge_lora_weights` function folds the low-rank updates back into the base weight matrices to eliminate inference overhead. The implementation computes `merged = (A @ B) * scaling`, transposes the result to match PyTorch's `[out_features, in_features]` convention, and adds it directly to the original linear layer's weight tensor. Once merged, the `LinearWithLoRA` wrappers can be removed, yielding a standard dense model with no runtime penalty from adapter computation.

## Complete Workflow Implementation

The following examples demonstrate the end-to-end usage of the LoRA and QLoRA implementation:

Inject LoRA into specific attention layers:

```python
import torch.nn as nn
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import inject_lora

model = ...  # Your nn.Module with Linear layers

target_modules = ["q_proj", "v_proj"]
lora_modules = inject_lora(model, target_modules, rank=8, alpha=16)

```

Apply NF4 quantization for QLoRA training:

```python
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import quantize_model

quant_state = quantize_model(model)  # Compress frozen weights to 4-bit

```

Train only the adapter parameters:

```python
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import train_lora, create_demo_data

data = create_demo_data()
losses = train_lora(model, data, epochs=20, lr=1e-3, batch_size=8)

```

Merge adapters and save for deployment:

```python
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import merge_lora_weights, save_lora_adapter

merge_lora_weights(model)  # Fold adapters into base weights

save_lora_adapter(model, "adapter.pt")  # Optional: save for later use

```

## Summary

- The **LoRA and QLoRA fine-tuning implementation** in `rohitg00/ai-engineering-from-scratch` provides a complete educational framework for parameter-efficient fine-tuning without external dependencies like PEFT.
- **LoRALayer** uses rank-decomposed matrices with scaled initialization to approximate weight updates, while **LinearWithLoRA** freezes base parameters and adds adapter contributions during the forward pass.
- **QLoRA** reduces memory consumption by quantizing frozen weights to NF4 4-bit precision using 64-element block-wise scaling, keeping only the small adapter matrices in full precision.
- The **inject_lora** function enables surgical targeting of specific layers (e.g., query and value projections), and **merge_lora_weights** eliminates inference latency by folding adapters back into the base model.

## Frequently Asked Questions

### What is the difference between LoRA and QLoRA in this implementation?

**LoRA** trains low-rank adapter matrices while keeping base weights in full precision, reducing trainable parameters by roughly 99%. **QLoRA** adds a quantization step where the frozen base weights are compressed to 4-bit NF4 format using the `quantize_model` function, cutting memory usage by an additional 75% while maintaining training stability through de-quantization during the forward pass.

### How does the NF4 quantization reduce memory usage?

The `quantize_to_nf4` function compresses each 64-element block of a weight tensor into 4-bit integers (range -8 to 7) plus a block-wise scaling factor. This reduces storage from 32-bit floats to 4-bit representations, shrinking the memory footprint of the frozen base model from approximately 4 bytes per parameter to 0.5 bytes per parameter, enabling fine-tuning of 7B+ models on single consumer GPUs.

### Can I merge the LoRA weights back into the base model?

Yes. The `merge_lora_weights` function computes the full-rank update `(A @ B) * scaling` and adds it directly to the original linear layer's weight tensor. After merging, the model contains only standard `nn.Linear` layers with updated weights, eliminating the computational overhead of separate adapter modules during inference.

### Which layers should I target with inject_lora?

The `target_modules` parameter accepts substring matches against layer names. Common targets include `["q_proj", "v_proj"]` for attention query and value projections, or `["q_proj", "k_proj", "v_proj", "o_proj"]` for all attention components. The implementation automatically skips non-linear layers and embeddings, applying adapters only to matching `nn.Linear` modules.