# Implementing Parameter-Efficient Finetuning with LoRA from Scratch: A Complete Guide

> Learn to implement parameter-efficient finetuning with LoRA from scratch. This guide explains LoRA's low-rank matrices for efficient LLM fine-tuning with minimal memory.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-12

---

**LoRA reduces trainable parameters by freezing pre-trained weights and injecting small, trainable low-rank matrices into linear layers, enabling efficient fine-tuning of large language models with minimal memory overhead.**

The `rasbt/LLMs-from-scratch` repository provides a transparent, educational implementation of parameter-efficient finetuning with LoRA that works with any PyTorch model built on `torch.nn.Linear`. This guide examines the core mechanics implemented in [`appendix_e.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/appendix_e.py), demonstrating how to retrofit existing architectures for low-cost adaptation.

## How LoRA Reduces Trainable Parameters

LoRA (Low-Rank Adaptation) reparameterizes weight updates using a low-rank decomposition. For any frozen weight matrix **W** with shape `(out_dim, in_dim)`, the method learns two smaller matrices:

- **A**: shape `(in_dim, r)`
- **B**: shape `(r, out_dim)`

The forward pass becomes:

```python
y = x @ W.T + (alpha / r) * x @ A @ B

```

Where `r` is the **rank** (typically 4-64) and `alpha` is a scaling factor. This reduces trainable parameters from `in_dim × out_dim` to `r × (in_dim + out_dim)`, often yielding **>99% parameter reduction** in large models.

## Core LoRA Components in appendix_e.py

The implementation in [`pkg/llms_from_scratch/appendix_e.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/appendix_e.py) provides three essential building blocks.

### LoRALayer: Low-Rank Matrix Decomposition

The `LoRALayer` class (lines 10-23) holds the trainable matrices **A** and **B**, initializing **A** with Kaiming uniform initialization and **B** with zeros to ensure training starts from the original frozen weights.

### LinearWithLoRA: Wrapping torch.nn.Linear

The `LinearWithLoRA` class (lines 25-35) wraps an existing `torch.nn.Linear` module. During the forward pass, it computes the original frozen linear output plus the scaled LoRA contribution `(alpha/r) * xAB`.

### replace_linear_with_lora: Recursive Model Conversion

The `replace_linear_with_lora` function (lines 37-44) recursively traverses a model's module tree, replacing every `Linear` instance with the `LinearWithLoRA` wrapper while preserving the frozen original weights.

## Step-by-Step Implementation Guide

Follow these steps to convert a pre-trained model for parameter-efficient finetuning.

### 1. Import the LoRA Utilities

```python
from llms_from_scratch.appendix_e import LinearWithLoRA, replace_linear_with_lora
import torch

```

### 2. Freeze Pre-trained Weights and Apply LoRA

Load your model and freeze all existing parameters before injecting LoRA layers:

```python

# Example using a model with nn.Linear layers

model = YourModel.from_pretrained("checkpoint")
model.eval()

# Freeze all base parameters

for param in model.parameters():
    param.requires_grad = False

# Inject LoRA with rank=16 and alpha=16 into all Linear layers

replace_linear_with_lora(model, rank=16, alpha=16)

```

### 3. Configure the Optimizer for LoRA Parameters Only

Filter the optimizer to update only the low-rank matrices:

```python
trainable_params = [p for p in model.parameters() if p.requires_grad]
optimizer = torch.optim.AdamW(trainable_params, lr=1e-4)

# Standard training loop

for batch in dataloader:
    optimizer.zero_grad()
    loss = model(batch["input_ids"], labels=batch["labels"]).loss
    loss.backward()
    optimizer.step()

```

### 4. Verify Trainable Parameter Count

Confirm the parameter efficiency of your setup:

```python
total_trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total_all = sum(p.numel() for p in model.parameters())
print(f"Trainable params: {total_trainable:,} ({100 * total_trainable / total_all:.2f}%)")

```

### 5. Optional: Merge LoRA Weights for Inference

For production inference, merge the LoRA contribution back into the original weight matrix to eliminate the separate low-rank computation. The [`additional_experiments.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/additional_experiments.py) file demonstrates this merged implementation:

```python

# Conceptual merged computation

merged_weight = linear.weight + (alpha * (A @ B)).T
output = x @ merged_weight.T + linear.bias

```

## Practical Training Examples

The repository includes complete training pipelines demonstrating parameter-efficient finetuning in practice.

- **[`ch07/01_main-chapter-code/exercise_experiments.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/exercise_experiments.py)**: Contains an end-to-end fine-tuning script with command-line flags (`--lora`) enabling LoRA training on instruction-following tasks.

- **[`ch06/02_bonus_additional-experiments/additional_experiments.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch06/02_bonus_additional-experiments/additional_experiments.py)**: Implements an alternative `LinearWithLoRAMerged` class that folds LoRA weights into the base linear layer for optimized inference speed.

- **[`pkg/llms_from_scratch/tests/test_appendix_e.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/tests/test_appendix_e.py)**: Validates that LoRA parameter counts match expectations and that gradients flow correctly through the low-rank pathway.

## Summary

- **LoRA** implements parameter-efficient finetuning by freezing original weights and learning low-rank updates via matrices **A** and **B**.
- The `rasbt/LLMs-from-scratch` implementation in [`appendix_e.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/appendix_e.py) provides `LoRALayer`, `LinearWithLoRA`, and `replace_linear_with_lora` for drop-in compatibility with any `nn.Linear`-based architecture.
- Only `r × (in_dim + out_dim)` parameters become trainable instead of `in_dim × out_dim`, reducing memory requirements by orders of magnitude.
- The [`exercise_experiments.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/exercise_experiments.py) script demonstrates production-ready training, while [`additional_experiments.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/additional_experiments.py) shows how to merge weights for efficient inference.

## Frequently Asked Questions

### What is the difference between rank and alpha in LoRA?

**Rank (`r`)** controls the dimensionality of the low-rank matrices **A** and **B**, directly impacting the number of trainable parameters and model capacity. **Alpha (`α`)** is a scaling factor applied as `alpha/r` in the forward pass to stabilize training; higher alpha values increase the influence of the LoRA adaptation relative to the frozen pre-trained weights.

### Can I apply LoRA to specific layers only?

Yes. While `replace_linear_with_lora` targets all `Linear` layers by default, you can modify the recursive function (lines 37-44 in [`appendix_e.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/appendix_e.py)) to filter by layer name, module type, or depth in the transformer stack—for example, applying LoRA only to attention query and value projections while excluding feed-forward layers.

### How do I merge LoRA weights back into the original model?

As shown in [`additional_experiments.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/additional_experiments.py), you can compute `merged_weight = W + (alpha/r) * B.T @ A.T` and assign it back to the original linear layer's weight tensor. This eliminates the separate matrix multiplication during inference, reducing latency while preserving the fine-tuned behavior.

### What memory savings does LoRA provide?

LoRA reduces optimizer memory (which tracks momentum and variance for each parameter) proportionally to the parameter reduction. For a 7B parameter model with rank-16 LoRA applied to attention layers, you typically train <100M parameters instead of 7B, reducing optimizer states from ~56GB to under 1GB when using AdamW (8 bytes per parameter).