# How to Use LoRA for Parameter-Efficient Fine-Tuning of Agent Models: Complete Implementation Guide

> Master LoRA for parameter-efficient fine-tuning of agent models. This guide shows how to adapt large models on single GPUs, reducing trainable parameters significantly. Get the complete implementation guide.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-06

---

**LoRA (Low-Rank Adaptation) reduces trainable parameters to 0.1–2% by adding small adapter matrices instead of updating full model weights, enabling fine-tuning of large agent models on single GPUs.**

In the *AI-Agent-Book* repository by [bojieli](https://github.com/bojieli/ai-agent-book), LoRA is implemented across multiple agent fine-tuning experiments. This guide walks through the complete workflow—from adapter configuration to inference—using actual source code from speech agents, distillation pipelines, and SFT trainers.

## What LoRA Does for Agent Models

LoRA freezes the pretrained LLM's original weights and inserts **trainable low-rank matrices** (A and B) into each attention and feed-forward layer. During **parameter-efficient fine-tuning**, only these adapter parameters receive gradients.

This approach delivers three critical advantages for production agent systems:

- **Memory efficiency**: Fit 7–30B parameter models on consumer GPUs
- **Fast deployment**: Adapter files (typically 10–100 MB) load instantly without transferring multi-gigabyte base models
- **Modular capabilities**: Swap LoRA adapters per task or user while sharing one base model

## Step-by-Step LoRA Implementation

### 1. Load Base Model with Adapter Support

In [`chapter8/prompt-distillation/train_sft_trl.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/prompt-distillation/train_sft_trl.py) (lines 59–71), the base model loads through a helper that prepares it for PEFT wrapping:

```python
from train_sft_trl import prepare_model_and_tokenizer

model, tokenizer, peft_cfg = prepare_model_and_tokenizer(
    model_name="meta-llama/Llama-2-7b-chat-hf",
    use_lora=True,        # Enable LoRA pathway

    lora_rank=32,         # Rank r: internal dimension of adapter matrices

    lora_alpha=16,        # Scaling factor: how strongly adapters influence output

)

```

The `prepare_model_and_tokenizer` function (lines 21–30) constructs a `LoraConfig` with **target modules** automatically detected—typically `["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]` for Llama-style architectures.

### 2. Configure LoRA Hyperparameters

The configuration in [`train_sft_trl.py`](https://github.com/bojieli/ai-agent-book/blob/main/train_sft_trl.py) demonstrates production-ready defaults:

```python
from peft import LoraConfig, get_peft_model

lora_config = LoraConfig(
    r=32,                  # Rank: higher = more expressive, more parameters

    lora_alpha=16,         # Alpha: scaling applied to adapter outputs

    target_modules=None,   # Auto-detected if None

    lora_dropout=0.0,      # Dropout: often 0 for small datasets

    bias="none",           # Whether to train bias terms

    task_type="CAUSAL_LM",
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()  # Shows: 0.5% trainable

```

Key insight from the repository: **learning rates for LoRA training run approximately 10× higher** than full fine-tuning. The scripts use `2e-4` versus `2e-5` for full-FT, as noted in training configuration comments.

### 3. Train with SFTTrainer

The `train_model` function (lines 45–55) initiates adapter-only training:

```python
from train_sft_trl import train_model

trainer = train_model(
    model=model,
    tokenizer=tokenizer,
    train_dataset=my_dataset,
    output_dir="out",
    num_train_epochs=1,
    per_device_train_batch_size=4,
    learning_rate=2e-4,   # 10× full-FT rate; only adapters update

    logging_steps=10,
)

trainer.train()

```

The **SFTTrainer from 🤗 TRL** (or Unsloth's optimized trainer) automatically excludes frozen parameters from the optimizer state, cutting memory usage by ~70% compared to full fine-tuning.

### 4. Save Adapter Weights

After training, persist only the adapter matrices—never duplicate the base model. From [`chapter7/sesame/sesame_csm_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/sesame/sesame_csm_sft_unsloth.py) (lines 238–240):

```python

# Save only adapter parameters (~50 MB instead of 14 GB)

model.save_pretrained("lora_model")      # adapter_model.bin + adapter_config.json

tokenizer.save_pretrained("lora_model")  # tokenizer files for completeness

```

The output directory contains:
- [`adapter_config.json`](https://github.com/bojieli/ai-agent-book/blob/main/adapter_config.json): LoRA hyperparameters (r, alpha, target modules)
- `adapter_model.bin`: Trainable weights (A and B matrices)

### 5. Load Adapters for Agent Inference

Deploy the fine-tuned agent by loading base weights once, then injecting adapters. From [`chapter7/sesame/inference.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/sesame/inference.py) (lines 40–43):

```python
from sesame.inference import load_model

model, processor = load_model(
    base_model_name="bojieli/tts-csm-1b",
    lora_path="lora_model",   # Directory with adapter files

    load_in_4bit=False,       # Or True for additional memory savings

)

```

The `PeftModel.from_pretrained` mechanism (used internally) merges adapter weights into the computation graph without modifying base tensors. This enables **hot-swapping adapters** for multi-tenant agent serving.

## Complete Inference Example

Run the fine-tuned voice agent with standard generation APIs:

```python
import soundfile as sf

output = model.generate(
    **processor(text="Hello, I am your new voice assistant!"),
    max_new_tokens=125,
)

audio = processor.decode(output, skip_special_tokens=True)
sf.write("greeting.wav", audio["audio"], samplerate=24000)

```

The LoRA adapter modifies how the model processes prompts and generates tokens, while the core agent pipeline (tool calling, memory management, safety filters) remains unchanged.

## Repository File Reference

| File | Purpose | Key LoRA Functions |
|------|---------|------------------|
| [`chapter8/prompt-distillation/train_sft_trl.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/prompt-distillation/train_sft_trl.py) | General SFT training with TRL | `prepare_model_and_tokenizer()`, `train_model()` |
| [`chapter7/sesame/sesame_csm_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/sesame/sesame_csm_sft_unsloth.py) | Speech model fine-tuning with Unsloth | Optimized LoRA saving (lines 238–240) |
| [`chapter7/sesame/inference.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/sesame/inference.py) | TTS agent inference | `load_model()` with LoRA injection |
| [`chapter7/orpheus/orpheus_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/orpheus/orpheus_sft_unsloth.py) | Alternative speech model | Same LoRA pattern, different base model |
| [`chapter7/cot-distillation/train_student.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/cot-distillation/train_student.py) | Chain-of-thought distillation | LoRA with rank-0 disabling for ablation |

## Summary

- **LoRA enables parameter-efficient fine-tuning** by training only 0.1–2% of parameters through low-rank adapter matrices
- **Initialize adapters** via `LoraConfig` and `get_peft_model()` before training
- **Scale learning rates 10× higher** than full fine-tuning for adapter convergence
- **Save only adapter files**—tiny artifacts that slot into any compatible base model
- **Deploy with `PeftModel.from_pretrained`** for hot-swappable, multi-tenant agent architectures

## Frequently Asked Questions

### What rank should I use for LoRA fine-tuning?

Start with **rank r=16 or r=32** for most agent tasks. Higher ranks (64–128) improve expressiveness on complex instruction-following but increase parameters linearly. The *AI-Agent-Book* repository uses r=32 for speech models and r=16 for prompt distillation experiments.

### Can I merge LoRA adapters back into the base model?

Yes—use `model.merge_and_unload()` to bake adapters into base weights for deployment scenarios requiring zero inference overhead. However, this sacrifices the modularity benefits of separate adapters.

### Does LoRA support training with quantization?

Absolutely. The repository demonstrates **QLoRA** (4-bit base models + 16-bit adapters) in [`sesame_csm_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/sesame_csm_sft_unsloth.py), enabling fine-tuning of 70B parameter models on single 24GB GPUs. Pass `load_in_4bit=True` during base model loading.

### How do I disable LoRA for ablation studies?

Set `lora_rank=0` or pass `use_lora=False` to `prepare_model_and_tokenizer()`—the training pipeline falls back to full fine-tuning or frozen feature extraction. The [`cot-distillation/train_student.py`](https://github.com/bojieli/ai-agent-book/blob/main/cot-distillation/train_student.py) script uses this pattern for controlled experiments.