How to Use the Unsloth Backend for Faster Training in Soup

Set backend: unsloth in your soup.yaml configuration and install the optional unsloth package to enable optimized 4-bit quantization and 2× faster training throughput on compatible GPUs.

Soup is an open-source framework for fine-tuning large language models that supports pluggable training backends. When you configure the Unsloth backend for faster training in Soup, the framework bypasses the standard transformers loading path and instead leverages specialized utilities in src/soup_cli/utils/unsloth.py that handle quantization, LoRA patching, and kernel optimization in a single pass.

Installation and Hardware Requirements

Before enabling the Unsloth backend, install the optional dependency and verify your environment meets the CUDA requirements.

pip install "unsloth @ git+https://github.com/unslothai/unsloth.git"

Unsloth requires a GPU with CUDA 11.8 or higher and a recent PyTorch build. The backend performs runtime detection via unsloth.is_unsloth_available() in src/soup_cli/utils/unsloth.py (lines 6-13), raising a clear error if the package is missing before attempting model loading.

Configuring the Unsloth Backend

To activate the optimized path, set the backend field to unsloth in your configuration file. The src/soup_cli/config/schema.py module validates this setting and related flags such as unsloth_bnb_4bit.

backend: unsloth
model: meta-llama/Llama-2-7b-chat-hf
max_seq_length: 2048
training:
  unsloth_bnb_4bit: true          # Enable 4-bit BitsAndBytes quantization

  quantization: 4bit
  lora_r: 64
  lora_alpha: 16
  lora_dropout: 0.05
  target_modules: auto

Validation helpers like validate_unsloth_bnb_4bit_compat strictly enforce that quantization-related options are only accepted when backend="unsloth". If you attempt to use unsloth_bnb_4bit: true with backend: transformers, the schema validation rejects the configuration.

How the Unsloth Backend Works

The Unsloth integration replaces Soup's generic model loading with optimized routines that combine quantization and LoRA application into a single step.

Backend Detection and Version Querying

The src/soup_cli/utils/unsloth.py module provides two utility functions for environment validation:

  • is_unsloth_available(): Attempts to import the unsloth package and returns a boolean (lines 6-13)
  • get_unsloth_version(): Returns the installed version string or None if the package is absent (lines 16-23)

You can verify installation manually:

from soup_cli.utils.unsloth import is_unsloth_available, get_unsloth_version

print(is_unsloth_available())  # True

print(get_unsloth_version())   # '0.4.2' or similar

Optimized Model Loading with LoRA Integration

The load_model_and_tokenizer() function (lines 26-80 in src/soup_cli/utils/unsloth.py) handles the entire initialization sequence:

  1. Calls unsloth.FastLanguageModel.from_pretrained with the requested quantization (default 4bit) and sequence length
  2. Resolves LoRA target modules (defaults: ["q_proj","k_proj","v_proj","o_proj","gate_proj","up_proj","down_proj"])
  3. Invokes FastLanguageModel.get_peft_model to attach LoRA adapters immediately

This consolidation eliminates the double-pass over model weights that occurs in the standard transformers + peft workflow.

from soup_cli.utils.unsloth import load_model_and_tokenizer

model, tokenizer = load_model_and_tokenizer(
    model_name="meta-llama/Llama-2-7b-chat-hf",
    max_seq_length=2048,
    quantization="4bit",
    lora_r=64,
    lora_alpha=16,
    lora_dropout=0.05,
    target_modules="auto",
)

Trainer Integration

Both the SFT and DPO trainer wrappers (src/soup_cli/trainer/sft.py and src/soup_cli/trainer/dpo.py) implement a private _setup_unsloth method. When backend="unsloth" is detected, this method delegates to load_model_and_tokenizer() instead of the standard transformers path.

The trainers receive a model that already has LoRA adapters attached, skipping the separate peft-based patching step required by the default backend. To run training:

soup train --config soup.yaml

Performance Impact

By handling quantization, LoRA patching, and kernel optimization internally, the Unsloth backend reduces memory-bandwidth pressure during the training loop. Benchmarks in the repository's benchmarks/ directory demonstrate up to 2× faster token-per-second throughput on supported GPUs compared to the generic transformers path.

The performance gain stems from avoiding redundant weight copies and leveraging fused kernels that apply adapters during the forward pass rather than as a separate transformation layer.

Switching Back to the Transformers Backend

To revert to the default behavior, change the backend field and omit Unsloth-specific flags:

backend: transformers
training:
  quantization: none
  lora_r: 64
  # Do not include unsloth_bnb_4bit or other Unsloth-specific options

Summary

  • Set backend: unsloth in soup.yaml to enable the optimized training path
  • Install Unsloth via pip with CUDA 11.8+ requirements before use
  • Configuration validation in src/soup_cli/config/schema.py ensures incompatible options are rejected for the wrong backend
  • Model loading occurs through load_model_and_tokenizer() in src/soup_cli/utils/unsloth.py, which uses FastLanguageModel for single-pass initialization
  • Trainer wrappers in sft.py and dpo.py automatically handle the backend via _setup_unsloth
  • Performance gains reach up to 2× faster throughput by eliminating double-pass weight processing and optimizing memory bandwidth

Frequently Asked Questions

What hardware is required to use the Unsloth backend in Soup?

The Unsloth backend requires a CUDA-capable GPU with Compute Capability 7.0 or higher and CUDA 11.8+. The is_unsloth_available() function in src/soup_cli/utils/unsloth.py verifies that the package can be imported, but you must ensure your PyTorch installation matches your CUDA version before training begins.

Can I use the same LoRA configuration between Unsloth and transformers backends?

Yes, the LoRA hyperparameters (lora_r, lora_alpha, lora_dropout) use identical schemas across both backends. However, the Unsloth backend automatically selects optimal default target modules (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj) when target_modules: auto is specified, whereas the transformers backend may require explicit module specification depending on your model architecture.

Why does Soup reject my configuration when I set unsloth_bnb_4bit: true?

The schema validation in src/soup_cli/config/schema.py enforces that unsloth_bnb_4bit and related quantization flags are only valid when backend: unsloth is explicitly declared. If you receive a validation error, verify that your YAML file specifies backend: unsloth and not the default backend: transformers.

How do I verify that training is actually using the Unsloth optimized kernels?

Check your training logs for successful initialization messages from load_model_and_tokenizer(), or programmatically verify the backend before training:

from soup_cli.utils.unsloth import is_unsloth_available, get_unsloth_version

assert is_unsloth_available(), "Unsloth not installed"
print(f"Using Unsloth version: {get_unsloth_version()}")

During training, you should observe higher GPU utilization and approximately 2× improved tokens-per-second compared to equivalent runs with backend: transformers in the benchmarks/ directory comparisons.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →