How to Apply LoRA Fine-Tuning to LLMs: A Complete Implementation Guide

LoRA fine-tuning freezes original model weights and injects trainable low-rank matrices, reducing trainable parameters by up to 10,000× while maintaining downstream task performance.

Fine-tuning large language models traditionally requires updating billions of parameters, but Low-Rank Adaptation (LoRA) offers an efficient alternative. This guide examines the practical implementation found in the rohitg00/ai-engineering-from-scratch repository, specifically within phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py, demonstrating how to adapt massive models with minimal computational resources.

Understanding LoRA Architecture

The LoRA approach decomposes weight updates into low-rank matrices rather than modifying the original pre-trained weights. According to the implementation in lora.py, the method freezes base parameters and introduces trainable rank decomposition matrices to each transformer layer, dramatically cutting memory and compute requirements.

LoRALayer: The Low-Rank Building Block

The LoRALayer class (lines 6–18) initializes two matrices: A with shape in_features × rank and B with shape rank × out_features. During the forward pass, the implementation computes the product of these matrices and scales the output by α/r, where alpha controls the learning rate scaling relative to the rank r. This scaling factor ensures stable training regardless of the chosen rank dimension.

LinearWithLoRA: Combining Frozen and Trainable Weights

The LinearWithLoRA wrapper (lines 20–33) modifies standard nn.Linear layers by preserving frozen base weights while adding the LoRA pathway. The forward computation follows linear(x) + LoRA(x), ensuring the original pre-trained knowledge remains intact while the low-rank adaptation learns task-specific adjustments.

Implementing LoRA Fine-Tuning

The inject_lora function (lines 35–52) traverses the model hierarchy, identifies target modules by name, and replaces selected nn.Linear layers with LinearWithLoRA instances. This function automatically freezes all non-LoRA parameters, ensuring only the low-rank matrices receive gradients during backpropagation.

Here is the complete workflow to apply LoRA fine-tuning to a feed-forward model:

import torch
from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import (
    create_demo_model,
    create_demo_data,
    inject_lora,
    train_lora,
    count_parameters,
)

# 1️⃣ Build a base model

model = create_demo_model()
print("Base params:", count_parameters(model))

# 2️⃣ Inject LoRA into the first and third linear layers

lora_layers = inject_lora(model, target_modules=["0", "2"], rank=8, alpha=16)
print("LoRA injected into:", list(lora_layers.keys()))

# 3️⃣ Generate dummy data

data = create_demo_data()

# 4️⃣ Train only the LoRA adapters

losses = train_lora(model, data, epochs=5, lr=1e-3)
print("Training loss – first vs. last:", losses[0], losses[-1])

The target_modules parameter accepts layer names or indices, allowing precise control over which components receive adapters. Common practice targets attention projections and feed-forward layers while excluding embedding and normalization layers.

QLoRA: 4-Bit Quantization for Extreme Memory Efficiency

For scenarios with severe memory constraints, the repository implements QLoRA support through quantize_to_nf4 and quantize_model functions. These utilities convert frozen base weights to 4-bit Normal Float 4 (NF4) format, reducing storage requirements while keeping LoRA adapters in full precision for stable training.

The quantization workflow stores the massive base model in low-precision memory, computes forward passes with dequantized weights, and backpropagates only through the small, full-precision adapters. This combination enables fine-tuning 70-billion parameter models on single consumer GPUs.

Adapter Persistence and Deployment

Merging Adapters for Inference

After training completes, the merge_lora_weights function (lines 67–81) folds the trained low-rank updates back into the original linear weights by computing W_new = W + (alpha/r) * B * A. This eliminates inference overhead by removing LoRA modules entirely, resulting in a standard nn.Sequential model with no additional latency.

from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import merge_lora_weights

# Assume `model` has been fine‑tuned with LoRA

merge_lora_weights(model)

# The model is now a plain nn.Sequential without any LoRA modules

print("After merge – trainable params:", count_parameters(model)["trainable"])

Saving and Loading Adapters

The save_lora_adapter and load_lora_adapter functions serialize only the low-rank matrices, creating tiny adapter files (typically megabytes versus gigabytes for full models). This architecture enables serving multiple specialized tasks from a single frozen backbone by hot-swapping adapter weights.

from phases.11_llm_engineering.08_fine_tuning_lora.code.lora import save_lora_adapter, load_lora_adapter

# Save the adapters

adapter_path = "lora_adapter.pt"
n_saved = save_lora_adapter(model, adapter_path)
print(f"Saved {n_saved} LoRA matrices")

# Load them into a fresh copy of the base model

new_model = create_demo_model()
inject_lora(new_model, target_modules=["0", "2"], rank=8, alpha=16)
load_lora_adapter(new_model, adapter_path)

Summary

  • LoRA freezes base model weights and injects trainable low-rank matrices A and B, reducing trainable parameters by orders of magnitude.
  • The complete implementation resides in phases/11-llm-engineering/08-fine-tuning-lora/code/lora.py, providing inject_lora, merge_lora_weights, and quantization utilities.
  • LinearWithLoRA wraps frozen layers and combines base outputs with low-rank updates during the forward pass.
  • QLoRA combines 4-bit quantization with LoRA adapters, enabling fine-tuning massive models on limited hardware.
  • Adapter persistence functions allow saving only the small LoRA weights, facilitating efficient multi-task serving and rapid task switching.

Frequently Asked Questions

What is the difference between LoRA and full fine-tuning?

Full fine-tuning updates every parameter in the pre-trained model, requiring massive memory for gradients and optimizer states. LoRA freezes the original weights and trains only small injection matrices, typically reducing trainable parameters by 10,000× while achieving comparable accuracy on downstream tasks. As implemented in the repository, only the A and B matrices in LoRALayer receive gradients.

How do I choose the rank and alpha hyperparameters?

Rank (r) controls the expressiveness of the low-rank update—common values range from 4 to 64, with higher ranks capturing more complex adaptations but requiring more memory. Alpha (α) scales the learning rate of the LoRA update; the implementation scales the product by alpha/r, so setting alpha=16 with rank=8 effectively doubles the learning rate relative to the base model. Start with rank=8 and alpha=16, then increase rank if underfitting occurs.

Can LoRA adapters be combined with quantized models?

Yes, the repository supports QLoRA through quantize_model and quantize_to_nf4 functions. These utilities convert the frozen base weights to 4-bit NF4 format while keeping LoRA adapters in full precision (FP16 or FP32). This combination allows fine-tuning models that would otherwise exceed GPU memory limits.

When should I merge LoRA weights versus keeping them separate?

Merge using merge_lora_weights when deploying to production environments where inference latency matters, as merged models eliminate the computational overhead of computing A @ B during each forward pass. Keep adapters separate when serving multiple specialized tasks from one base model, allowing rapid task switching by loading different adapter files without reloading the massive backbone.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →