Implementing Parameter-Efficient Finetuning with LoRA from Scratch: A Complete Guide

LoRA reduces trainable parameters by freezing pre-trained weights and injecting small, trainable low-rank matrices into linear layers, enabling efficient fine-tuning of large language models with minimal memory overhead.

The rasbt/LLMs-from-scratch repository provides a transparent, educational implementation of parameter-efficient finetuning with LoRA that works with any PyTorch model built on torch.nn.Linear. This guide examines the core mechanics implemented in appendix_e.py, demonstrating how to retrofit existing architectures for low-cost adaptation.

How LoRA Reduces Trainable Parameters

LoRA (Low-Rank Adaptation) reparameterizes weight updates using a low-rank decomposition. For any frozen weight matrix W with shape (out_dim, in_dim), the method learns two smaller matrices:

  • A: shape (in_dim, r)
  • B: shape (r, out_dim)

The forward pass becomes:

y = x @ W.T + (alpha / r) * x @ A @ B

Where r is the rank (typically 4-64) and alpha is a scaling factor. This reduces trainable parameters from in_dim × out_dim to r × (in_dim + out_dim), often yielding >99% parameter reduction in large models.

Core LoRA Components in appendix_e.py

The implementation in pkg/llms_from_scratch/appendix_e.py provides three essential building blocks.

LoRALayer: Low-Rank Matrix Decomposition

The LoRALayer class (lines 10-23) holds the trainable matrices A and B, initializing A with Kaiming uniform initialization and B with zeros to ensure training starts from the original frozen weights.

LinearWithLoRA: Wrapping torch.nn.Linear

The LinearWithLoRA class (lines 25-35) wraps an existing torch.nn.Linear module. During the forward pass, it computes the original frozen linear output plus the scaled LoRA contribution (alpha/r) * xAB.

replace_linear_with_lora: Recursive Model Conversion

The replace_linear_with_lora function (lines 37-44) recursively traverses a model's module tree, replacing every Linear instance with the LinearWithLoRA wrapper while preserving the frozen original weights.

Step-by-Step Implementation Guide

Follow these steps to convert a pre-trained model for parameter-efficient finetuning.

1. Import the LoRA Utilities

from llms_from_scratch.appendix_e import LinearWithLoRA, replace_linear_with_lora
import torch

2. Freeze Pre-trained Weights and Apply LoRA

Load your model and freeze all existing parameters before injecting LoRA layers:


# Example using a model with nn.Linear layers

model = YourModel.from_pretrained("checkpoint")
model.eval()

# Freeze all base parameters

for param in model.parameters():
    param.requires_grad = False

# Inject LoRA with rank=16 and alpha=16 into all Linear layers

replace_linear_with_lora(model, rank=16, alpha=16)

3. Configure the Optimizer for LoRA Parameters Only

Filter the optimizer to update only the low-rank matrices:

trainable_params = [p for p in model.parameters() if p.requires_grad]
optimizer = torch.optim.AdamW(trainable_params, lr=1e-4)

# Standard training loop

for batch in dataloader:
    optimizer.zero_grad()
    loss = model(batch["input_ids"], labels=batch["labels"]).loss
    loss.backward()
    optimizer.step()

4. Verify Trainable Parameter Count

Confirm the parameter efficiency of your setup:

total_trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total_all = sum(p.numel() for p in model.parameters())
print(f"Trainable params: {total_trainable:,} ({100 * total_trainable / total_all:.2f}%)")

5. Optional: Merge LoRA Weights for Inference

For production inference, merge the LoRA contribution back into the original weight matrix to eliminate the separate low-rank computation. The additional_experiments.py file demonstrates this merged implementation:


# Conceptual merged computation

merged_weight = linear.weight + (alpha * (A @ B)).T
output = x @ merged_weight.T + linear.bias

Practical Training Examples

The repository includes complete training pipelines demonstrating parameter-efficient finetuning in practice.

Summary

  • LoRA implements parameter-efficient finetuning by freezing original weights and learning low-rank updates via matrices A and B.
  • The rasbt/LLMs-from-scratch implementation in appendix_e.py provides LoRALayer, LinearWithLoRA, and replace_linear_with_lora for drop-in compatibility with any nn.Linear-based architecture.
  • Only r × (in_dim + out_dim) parameters become trainable instead of in_dim × out_dim, reducing memory requirements by orders of magnitude.
  • The exercise_experiments.py script demonstrates production-ready training, while additional_experiments.py shows how to merge weights for efficient inference.

Frequently Asked Questions

What is the difference between rank and alpha in LoRA?

Rank (r) controls the dimensionality of the low-rank matrices A and B, directly impacting the number of trainable parameters and model capacity. Alpha (α) is a scaling factor applied as alpha/r in the forward pass to stabilize training; higher alpha values increase the influence of the LoRA adaptation relative to the frozen pre-trained weights.

Can I apply LoRA to specific layers only?

Yes. While replace_linear_with_lora targets all Linear layers by default, you can modify the recursive function (lines 37-44 in appendix_e.py) to filter by layer name, module type, or depth in the transformer stack—for example, applying LoRA only to attention query and value projections while excluding feed-forward layers.

How do I merge LoRA weights back into the original model?

As shown in additional_experiments.py, you can compute merged_weight = W + (alpha/r) * B.T @ A.T and assign it back to the original linear layer's weight tensor. This eliminates the separate matrix multiplication during inference, reducing latency while preserving the fine-tuned behavior.

What memory savings does LoRA provide?

LoRA reduces optimizer memory (which tracks momentum and variance for each parameter) proportionally to the parameter reduction. For a 7B parameter model with rank-16 LoRA applied to attention layers, you typically train <100M parameters instead of 7B, reducing optimizer states from ~56GB to under 1GB when using AdamW (8 bytes per parameter).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →