# How to Use Unsloth for Accelerated Fine-Tuning of AI Agents

> Accelerate AI agent fine-tuning with Unsloth. Slash training time in half and cut VRAM by 30% using optimized CUDA kernels and LoRA. Seamlessly integrates with Hugging Face Trainer.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-24

---

**Unsloth accelerates AI agent fine-tuning by combining optimized CUDA kernels, minimal gradient checkpointing, and LoRA adapters, reducing VRAM consumption by approximately 30% and cutting training time in half while maintaining full compatibility with standard Hugging Face Trainer workflows.**

The `bojieli/ai-agent-book` repository provides production-ready examples demonstrating how to leverage Unsloth—a lightweight wrapper around Hugging Face Transformers—to efficiently fine-tune large language models for AI agent applications. By automatically selecting the fastest kernels and implementing a specialized gradient checkpointing mode, Unsloth enables practitioners to train multi-billion parameter models on consumer hardware without rewriting their existing training infrastructure.

## Core Optimizations That Drive Speed

Unsloth achieves its performance gains through three primary mechanisms that operate transparently beneath the standard training API.

### Hardware-Optimized Kernel Selection

The library automatically detects available CUDA capabilities and routes operations through the fastest available implementations. This includes custom kernels such as `cut_cross_entropy` that improve token throughput by 15–25% over vanilla Transformers, selecting the optimal path based on the specific GPU architecture.

### The "Unsloth" Gradient Checkpointing Mode

Unlike standard gradient checkpointing, Unsloth's `use_gradient_checkpointing="unsloth"` mode stores only the minimal activation slices required for backpropagation. This specialized approach cuts VRAM usage by approximately 30% while maintaining the same throughput, effectively doubling feasible batch sizes on memory-constrained hardware.

### LoRA-First Architecture

Unsloth enforces a parameter-efficient fine-tuning strategy where `FastModel.get_peft_model` creates LoRA-adapted copies of base models. By freezing the full pre-trained weights and training only 0.1–10% of parameters (typically adapter matrices in `q_proj`, `v_proj`, and MLP layers), the optimizer processes dramatically smaller gradient tensors, accelerating convergence.

## The Standard Unsloth Workflow

All examples in the repository follow a consistent four-step pattern that mirrors standard Transformers workflows while injecting Unsloth's optimizations underneath.

First, import `FastModel` from the `unsloth` package and load your base checkpoint using `FastModel.from_pretrained`, specifying the concrete model class (e.g., `CsmForConditionalGeneration` or `LlamaForCausalLM`), maximum sequence length, and quantization settings.

Second, wrap the model with LoRA adapters via `FastModel.get_peft_model`, passing the rank `r`, `target_modules`, `lora_alpha`, and crucially setting `use_gradient_checkpointing="unsloth"` to activate the memory-saving checkpointing mode.

Third, prepare your dataset using standard Hugging Face `Dataset` objects, applying your model's specific processor to handle multimodal inputs (text and audio for TTS models, for instance).

Finally, instantiate a standard `Trainer` and call `train()`. The heavy lifting—optimizer state management, mixed-precision scaling, and kernel dispatch—happens automatically within Unsloth's internals.

## End-to-End Code Examples

### Fine-Tuning Sesame CSM for Text-to-Speech

The file [`chapter8/sesame/sesame_csm_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/sesame/sesame_csm_sft_unsloth.py) demonstrates fine-tuning the 1B parameter `unsloth/csm-1b` model on custom text-to-speech data. The implementation loads the model with its specific processor, applies LoRA adapters to attention and MLP layers, and processes the `MrDragonFox/Elise` dataset with audio resampling.

```python
from unsloth import FastModel, is_bfloat16_supported
from transformers import TrainingArguments, Trainer
from datasets import load_dataset, Audio

# Load base model with specific auto_model class

model, processor = FastModel.from_pretrained(
    model_name="unsloth/csm-1b",
    max_seq_length=2048,
    dtype=None,  # auto-detect fp16/bf16

    auto_model=CsmForConditionalGeneration,
    load_in_4bit=False,
)

# Apply LoRA with Unsloth's optimized checkpointing

model = FastModel.get_peft_model(
    model,
    r=32,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
    lora_alpha=32,
    lora_dropout=0,
    bias="none",
    use_gradient_checkpointing="unsloth",  # Critical for VRAM reduction

    random_state=3407,
)

# Standard Trainer configuration

trainer = Trainer(
    model=model,
    train_dataset=train_ds,
    args=TrainingArguments(
        per_device_train_batch_size=2,
        gradient_accumulation_steps=4,
        max_steps=200,
        learning_rate=2e-4,
        fp16=not is_bfloat16_supported(),
        bf16=is_bfloat16_supported(),
        optim="adamw_8bit",
        weight_decay=0.01,
    ),
)
trainer.train()

```

### Scaling to 7B Models with 4-Bit Quantization

For larger models like `unsloth/mistral-7b-v0.3`, the repository demonstrates aggressive quantization in [`chapter8/continued-pretraining/continued-pretrain.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/continued-pretraining/continued-pretrain.py). Setting `load_in_4bit=True` reduces memory footprint by an additional 2×, allowing 7B parameter models to fit within approximately 10GB VRAM.

```python
from unsloth import FastModel, is_bfloat16_supported
from transformers import AutoModelForCausalLM, TrainingArguments, Trainer

model, processor = FastModel.from_pretrained(
    model_name="unsloth/mistral-7b-v0.3",
    max_seq_length=2048,
    dtype=None,
    auto_model=AutoModelForCausalLM,
    load_in_4bit=True,  # Enable 4-bit quantization

)

model = FastModel.get_peft_model(
    model,
    r=16,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    lora_alpha=16,
    use_gradient_checkpointing="unsloth",
)

trainer = Trainer(
    model=model,
    train_dataset=train_ds,
    args=TrainingArguments(
        per_device_train_batch_size=1,  # Necessary for 16GB VRAM with 7B model

        max_steps=500,
        learning_rate=3e-4,
        fp16=not is_bfloat16_supported(),
        bf16=is_bfloat16_supported(),
    ),
)
trainer.train()

```

### Swapping Base Models: Sesame to Orpheus

The repository highlights Unsloth's model-agnostic API in [`chapter8/orpheus/orpheus_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/orpheus/orpheus_sft_unsloth.py). By changing only the `model_name` constant from `unsloth/csm-1b` to `unsloth/orpheus-3b-0.1-ft`, the identical training pipeline adapts to a different architecture and parameter count without modifying the data processing or LoRA configuration logic.

```python

# Single-line change to switch architectures

BASE_MODEL = "unsloth/orpheus-3b-0.1-ft"  # Previously "unsloth/csm-1b"

# Remaining preprocessing, LoRA setup, and Trainer code remain identical

# to the Sesame example in chapter8/sesame/sesame_csm_sft_unsloth.py

```

## Key Implementation Files

The `bojieli/ai-agent-book` repository contains several reference implementations that demonstrate production patterns:

- **[`chapter8/sesame/sesame_csm_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/sesame/sesame_csm_sft_unsloth.py)**: Complete TTS fine-tuning pipeline showing multimodal dataset preparation with `AutoProcessor` and audio-specific preprocessing.
- **[`chapter8/orpheus/orpheus_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/orpheus/orpheus_sft_unsloth.py)**: 3B parameter TTS example illustrating how to adapt the Sesame workflow to larger checkpoints.
- **[`chapter8/continued-pretraining/continued-pretrain.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/continued-pretraining/continued-pretrain.py)**: Generic script supporting any `unsloth/*` model (Mistral-7B, Llama-3-8B, Gemma-7B) with command-line argument parsing for `--base_model`.
- **[`chapter8/sesame/inference.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/sesame/inference.py)**: Post-training inference utilities demonstrating how to load fine-tuned `FastModel` checkpoints for audio generation.

## Summary

- **Unsloth** wraps Hugging Face Transformers to accelerate fine-tuning through optimized CUDA kernels and specialized gradient checkpointing.
- The **`use_gradient_checkpointing="unsloth"`** parameter reduces VRAM usage by ~30% compared to standard checkpointing.
- **4-bit quantization** via `load_in_4bit=True` enables training 7B parameter models on single consumer GPUs with ~10GB VRAM.
- **LoRA adapters** created through `FastModel.get_peft_model` train only 0.1–10% of parameters, drastically reducing optimizer overhead.
- The repository provides ready-to-adapt scripts in `chapter8/` that handle text-to-speech and continued pretraining tasks without modifying core training loops.

## Frequently Asked Questions

### What hardware requirements are needed to use Unsloth for fine-tuning?

Unsloth supports NVIDIA GPUs with CUDA capability. With 4-bit quantization enabled, you can fine-tune 7B parameter models on a single GPU with 16GB VRAM (such as an RTX 4090). For 1–3B parameter models, 8GB VRAM is typically sufficient. The library automatically detects whether your hardware supports bfloat16 and falls back to float16 when necessary.

### How does Unsloth differ from standard PEFT and Transformers workflows?

Unsloth maintains API compatibility with Hugging Face Transformers and PEFT while substituting optimized kernels underneath. The primary differences are the use of `FastModel.from_pretrained` instead of `AutoModel.from_pretrained`, the requirement to specify the concrete model class in `auto_model`, and the `use_gradient_checkpointing="unsloth"` flag which activates custom memory management not available in standard gradient checkpointing.

### Can I use Unsloth with custom datasets and existing training scripts?

Yes. Unsloth accepts standard Hugging Face `Dataset` objects and works with the standard `Trainer` class. You can adapt existing scripts by replacing the model loading lines with `FastModel.from_pretrained` and adding the LoRA wrapper via `FastModel.get_peft_model`. The data preprocessing, collation, and training loop logic remains identical to standard Transformers workflows.

### Does Unsloth support inference after fine-tuning?

Yes. The repository includes inference examples in [`chapter8/sesame/inference.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/sesame/inference.py) and [`chapter8/orpheus/inference.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/orpheus/inference.py) demonstrating how to load fine-tuned checkpoints using the same `FastModel` API. You can load saved adapters using `FastModel.from_pretrained` with the adapter path, then generate outputs using the standard `generate()` method or model-specific inference calls.