How to Use the Unsloth Backend for Faster Training in Soup
Set backend: unsloth in your soup.yaml configuration and install the optional unsloth package to enable optimized 4-bit quantization and 2× faster training throughput on compatible GPUs.
Soup is an open-source framework for fine-tuning large language models that supports pluggable training backends. When you configure the Unsloth backend for faster training in Soup, the framework bypasses the standard transformers loading path and instead leverages specialized utilities in src/soup_cli/utils/unsloth.py that handle quantization, LoRA patching, and kernel optimization in a single pass.
Installation and Hardware Requirements
Before enabling the Unsloth backend, install the optional dependency and verify your environment meets the CUDA requirements.
pip install "unsloth @ git+https://github.com/unslothai/unsloth.git"
Unsloth requires a GPU with CUDA 11.8 or higher and a recent PyTorch build. The backend performs runtime detection via unsloth.is_unsloth_available() in src/soup_cli/utils/unsloth.py (lines 6-13), raising a clear error if the package is missing before attempting model loading.
Configuring the Unsloth Backend
To activate the optimized path, set the backend field to unsloth in your configuration file. The src/soup_cli/config/schema.py module validates this setting and related flags such as unsloth_bnb_4bit.
backend: unsloth
model: meta-llama/Llama-2-7b-chat-hf
max_seq_length: 2048
training:
unsloth_bnb_4bit: true # Enable 4-bit BitsAndBytes quantization
quantization: 4bit
lora_r: 64
lora_alpha: 16
lora_dropout: 0.05
target_modules: auto
Validation helpers like validate_unsloth_bnb_4bit_compat strictly enforce that quantization-related options are only accepted when backend="unsloth". If you attempt to use unsloth_bnb_4bit: true with backend: transformers, the schema validation rejects the configuration.
How the Unsloth Backend Works
The Unsloth integration replaces Soup's generic model loading with optimized routines that combine quantization and LoRA application into a single step.
Backend Detection and Version Querying
The src/soup_cli/utils/unsloth.py module provides two utility functions for environment validation:
is_unsloth_available(): Attempts to import theunslothpackage and returns a boolean (lines 6-13)get_unsloth_version(): Returns the installed version string orNoneif the package is absent (lines 16-23)
You can verify installation manually:
from soup_cli.utils.unsloth import is_unsloth_available, get_unsloth_version
print(is_unsloth_available()) # True
print(get_unsloth_version()) # '0.4.2' or similar
Optimized Model Loading with LoRA Integration
The load_model_and_tokenizer() function (lines 26-80 in src/soup_cli/utils/unsloth.py) handles the entire initialization sequence:
- Calls
unsloth.FastLanguageModel.from_pretrainedwith the requested quantization (default 4bit) and sequence length - Resolves LoRA target modules (defaults:
["q_proj","k_proj","v_proj","o_proj","gate_proj","up_proj","down_proj"]) - Invokes
FastLanguageModel.get_peft_modelto attach LoRA adapters immediately
This consolidation eliminates the double-pass over model weights that occurs in the standard transformers + peft workflow.
from soup_cli.utils.unsloth import load_model_and_tokenizer
model, tokenizer = load_model_and_tokenizer(
model_name="meta-llama/Llama-2-7b-chat-hf",
max_seq_length=2048,
quantization="4bit",
lora_r=64,
lora_alpha=16,
lora_dropout=0.05,
target_modules="auto",
)
Trainer Integration
Both the SFT and DPO trainer wrappers (src/soup_cli/trainer/sft.py and src/soup_cli/trainer/dpo.py) implement a private _setup_unsloth method. When backend="unsloth" is detected, this method delegates to load_model_and_tokenizer() instead of the standard transformers path.
The trainers receive a model that already has LoRA adapters attached, skipping the separate peft-based patching step required by the default backend. To run training:
soup train --config soup.yaml
Performance Impact
By handling quantization, LoRA patching, and kernel optimization internally, the Unsloth backend reduces memory-bandwidth pressure during the training loop. Benchmarks in the repository's benchmarks/ directory demonstrate up to 2× faster token-per-second throughput on supported GPUs compared to the generic transformers path.
The performance gain stems from avoiding redundant weight copies and leveraging fused kernels that apply adapters during the forward pass rather than as a separate transformation layer.
Switching Back to the Transformers Backend
To revert to the default behavior, change the backend field and omit Unsloth-specific flags:
backend: transformers
training:
quantization: none
lora_r: 64
# Do not include unsloth_bnb_4bit or other Unsloth-specific options
Summary
- Set
backend: unslothinsoup.yamlto enable the optimized training path - Install Unsloth via pip with CUDA 11.8+ requirements before use
- Configuration validation in
src/soup_cli/config/schema.pyensures incompatible options are rejected for the wrong backend - Model loading occurs through
load_model_and_tokenizer()insrc/soup_cli/utils/unsloth.py, which usesFastLanguageModelfor single-pass initialization - Trainer wrappers in
sft.pyanddpo.pyautomatically handle the backend via_setup_unsloth - Performance gains reach up to 2× faster throughput by eliminating double-pass weight processing and optimizing memory bandwidth
Frequently Asked Questions
What hardware is required to use the Unsloth backend in Soup?
The Unsloth backend requires a CUDA-capable GPU with Compute Capability 7.0 or higher and CUDA 11.8+. The is_unsloth_available() function in src/soup_cli/utils/unsloth.py verifies that the package can be imported, but you must ensure your PyTorch installation matches your CUDA version before training begins.
Can I use the same LoRA configuration between Unsloth and transformers backends?
Yes, the LoRA hyperparameters (lora_r, lora_alpha, lora_dropout) use identical schemas across both backends. However, the Unsloth backend automatically selects optimal default target modules (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj) when target_modules: auto is specified, whereas the transformers backend may require explicit module specification depending on your model architecture.
Why does Soup reject my configuration when I set unsloth_bnb_4bit: true?
The schema validation in src/soup_cli/config/schema.py enforces that unsloth_bnb_4bit and related quantization flags are only valid when backend: unsloth is explicitly declared. If you receive a validation error, verify that your YAML file specifies backend: unsloth and not the default backend: transformers.
How do I verify that training is actually using the Unsloth optimized kernels?
Check your training logs for successful initialization messages from load_model_and_tokenizer(), or programmatically verify the backend before training:
from soup_cli.utils.unsloth import is_unsloth_available, get_unsloth_version
assert is_unsloth_available(), "Unsloth not installed"
print(f"Using Unsloth version: {get_unsloth_version()}")
During training, you should observe higher GPU utilization and approximately 2× improved tokens-per-second compared to equivalent runs with backend: transformers in the benchmarks/ directory comparisons.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →