# How to Apply LoRA Fine-Tuning to LLMs: A Complete Implementation Guide

> Learn how to apply LoRA fine tuning to LLMs. Freeze original weights and inject trainable low rank matrices to train models efficiently and maintain performance. Complete implementation guide.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**LoRA fine-tuning freezes original model weights and injects trainable low-rank matrices, reducing trainable parameters by up to 10,000× while maintaining downstream task performance.**

Fine-tuning large language models traditionally requires updating billions of parameters, but Low-Rank Adaptation (LoRA) offers an efficient alternative. This guide examines the practical implementation found in the `rohitg00/ai-engineering-from-scratch` repository, specifically within [`phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py), demonstrating how to adapt massive models with minimal computational resources.

## Understanding LoRA Architecture

The LoRA approach decomposes weight updates into low-rank matrices rather than modifying the original pre-trained weights. According to the implementation in [`lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/lora.py), the method freezes base parameters and introduces trainable rank decomposition matrices to each transformer layer, dramatically cutting memory and compute requirements.

### LoRALayer: The Low-Rank Building Block

The `LoRALayer` class (lines 6–18) initializes two matrices: **A** with shape `in_features × rank` and **B** with shape `rank × out_features`. During the forward pass, the implementation computes the product of these matrices and scales the output by **α/r**, where `alpha` controls the learning rate scaling relative to the rank `r`. This scaling factor ensures stable training regardless of the chosen rank dimension.

### LinearWithLoRA: Combining Frozen and Trainable Weights

The `LinearWithLoRA` wrapper (lines 20–33) modifies standard `nn.Linear` layers by preserving frozen base weights while adding the LoRA pathway. The forward computation follows `linear(x) + LoRA(x)`, ensuring the original pre-trained knowledge remains intact while the low-rank adaptation learns task-specific adjustments.

## Implementing LoRA Fine-Tuning

The `inject_lora` function (lines 35–52) traverses the model hierarchy, identifies target modules by name, and replaces selected `nn.Linear` layers with `LinearWithLoRA` instances. This function automatically freezes all non-LoRA parameters, ensuring only the low-rank matrices receive gradients during backpropagation.

Here is the complete workflow to apply LoRA fine-tuning to a feed-forward model:

```python
import torch
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import (
    create_demo_model,
    create_demo_data,
    inject_lora,
    train_lora,
    count_parameters,
)

# 1️⃣ Build a base model

model = create_demo_model()
print("Base params:", count_parameters(model))

# 2️⃣ Inject LoRA into the first and third linear layers

lora_layers = inject_lora(model, target_modules=["0", "2"], rank=8, alpha=16)
print("LoRA injected into:", list(lora_layers.keys()))

# 3️⃣ Generate dummy data

data = create_demo_data()

# 4️⃣ Train only the LoRA adapters

losses = train_lora(model, data, epochs=5, lr=1e-3)
print("Training loss – first vs. last:", losses[0], losses[-1])

```

The `target_modules` parameter accepts layer names or indices, allowing precise control over which components receive adapters. Common practice targets attention projections and feed-forward layers while excluding embedding and normalization layers.

## QLoRA: 4-Bit Quantization for Extreme Memory Efficiency

For scenarios with severe memory constraints, the repository implements QLoRA support through `quantize_to_nf4` and `quantize_model` functions. These utilities convert frozen base weights to 4-bit Normal Float 4 (NF4) format, reducing storage requirements while keeping LoRA adapters in full precision for stable training.

The quantization workflow stores the massive base model in low-precision memory, computes forward passes with dequantized weights, and backpropagates only through the small, full-precision adapters. This combination enables fine-tuning 70-billion parameter models on single consumer GPUs.

## Adapter Persistence and Deployment

### Merging Adapters for Inference

After training completes, the `merge_lora_weights` function (lines 67–81) folds the trained low-rank updates back into the original linear weights by computing `W_new = W + (alpha/r) * B * A`. This eliminates inference overhead by removing LoRA modules entirely, resulting in a standard `nn.Sequential` model with no additional latency.

```python
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import merge_lora_weights

# Assume `model` has been fine‑tuned with LoRA

merge_lora_weights(model)

# The model is now a plain nn.Sequential without any LoRA modules

print("After merge – trainable params:", count_parameters(model)["trainable"])

```

### Saving and Loading Adapters

The `save_lora_adapter` and `load_lora_adapter` functions serialize only the low-rank matrices, creating tiny adapter files (typically megabytes versus gigabytes for full models). This architecture enables serving multiple specialized tasks from a single frozen backbone by hot-swapping adapter weights.

```python
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import save_lora_adapter, load_lora_adapter

# Save the adapters

adapter_path = "lora_adapter.pt"
n_saved = save_lora_adapter(model, adapter_path)
print(f"Saved {n_saved} LoRA matrices")

# Load them into a fresh copy of the base model

new_model = create_demo_model()
inject_lora(new_model, target_modules=["0", "2"], rank=8, alpha=16)
load_lora_adapter(new_model, adapter_path)

```

## Summary

- **LoRA freezes base model weights** and injects trainable low-rank matrices **A** and **B**, reducing trainable parameters by orders of magnitude.
- The complete implementation resides in [`phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py), providing `inject_lora`, `merge_lora_weights`, and quantization utilities.
- `LinearWithLoRA` wraps frozen layers and combines base outputs with low-rank updates during the forward pass.
- **QLoRA** combines 4-bit quantization with LoRA adapters, enabling fine-tuning massive models on limited hardware.
- Adapter persistence functions allow saving only the small LoRA weights, facilitating efficient multi-task serving and rapid task switching.

## Frequently Asked Questions

### What is the difference between LoRA and full fine-tuning?

Full fine-tuning updates every parameter in the pre-trained model, requiring massive memory for gradients and optimizer states. LoRA freezes the original weights and trains only small injection matrices, typically reducing trainable parameters by 10,000× while achieving comparable accuracy on downstream tasks. As implemented in the repository, only the **A** and **B** matrices in `LoRALayer` receive gradients.

### How do I choose the rank and alpha hyperparameters?

**Rank** (`r`) controls the expressiveness of the low-rank update—common values range from 4 to 64, with higher ranks capturing more complex adaptations but requiring more memory. **Alpha** (`α`) scales the learning rate of the LoRA update; the implementation scales the product by `alpha/r`, so setting `alpha=16` with `rank=8` effectively doubles the learning rate relative to the base model. Start with `rank=8` and `alpha=16`, then increase rank if underfitting occurs.

### Can LoRA adapters be combined with quantized models?

Yes, the repository supports QLoRA through `quantize_model` and `quantize_to_nf4` functions. These utilities convert the frozen base weights to 4-bit NF4 format while keeping LoRA adapters in full precision (FP16 or FP32). This combination allows fine-tuning models that would otherwise exceed GPU memory limits.

### When should I merge LoRA weights versus keeping them separate?

**Merge** using `merge_lora_weights` when deploying to production environments where inference latency matters, as merged models eliminate the computational overhead of computing `A @ B` during each forward pass. **Keep adapters separate** when serving multiple specialized tasks from one base model, allowing rapid task switching by loading different adapter files without reloading the massive backbone.