# How to Use the Unsloth Backend for Faster Training in Soup

> Accelerate Soup training with Unsloth backend. Easily configure your soup.yaml and install unsloth for 4-bit quantization and 2x faster GPU training.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Set `backend: unsloth` in your [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) configuration and install the optional `unsloth` package to enable optimized 4-bit quantization and 2× faster training throughput on compatible GPUs.**

Soup is an open-source framework for fine-tuning large language models that supports pluggable training backends. When you configure the Unsloth backend for faster training in Soup, the framework bypasses the standard `transformers` loading path and instead leverages specialized utilities in [`src/soup_cli/utils/unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/unsloth.py) that handle quantization, LoRA patching, and kernel optimization in a single pass.

## Installation and Hardware Requirements

Before enabling the Unsloth backend, install the optional dependency and verify your environment meets the CUDA requirements.

```bash
pip install "unsloth @ git+https://github.com/unslothai/unsloth.git"

```

**Unsloth requires a GPU with CUDA 11.8 or higher** and a recent PyTorch build. The backend performs runtime detection via `unsloth.is_unsloth_available()` in [`src/soup_cli/utils/unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/unsloth.py) (lines 6-13), raising a clear error if the package is missing before attempting model loading.

## Configuring the Unsloth Backend

To activate the optimized path, set the **backend** field to `unsloth` in your configuration file. The [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) module validates this setting and related flags such as `unsloth_bnb_4bit`.

```yaml
backend: unsloth
model: meta-llama/Llama-2-7b-chat-hf
max_seq_length: 2048
training:
  unsloth_bnb_4bit: true          # Enable 4-bit BitsAndBytes quantization

  quantization: 4bit
  lora_r: 64
  lora_alpha: 16
  lora_dropout: 0.05
  target_modules: auto

```

Validation helpers like `validate_unsloth_bnb_4bit_compat` strictly enforce that quantization-related options are only accepted when `backend="unsloth"`. If you attempt to use `unsloth_bnb_4bit: true` with `backend: transformers`, the schema validation rejects the configuration.

## How the Unsloth Backend Works

The Unsloth integration replaces Soup's generic model loading with optimized routines that combine quantization and LoRA application into a single step.

### Backend Detection and Version Querying

The [`src/soup_cli/utils/unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/unsloth.py) module provides two utility functions for environment validation:

- **`is_unsloth_available()`**: Attempts to import the `unsloth` package and returns a boolean (lines 6-13)
- **`get_unsloth_version()`**: Returns the installed version string or `None` if the package is absent (lines 16-23)

You can verify installation manually:

```python
from soup_cli.utils.unsloth import is_unsloth_available, get_unsloth_version

print(is_unsloth_available())  # True

print(get_unsloth_version())   # '0.4.2' or similar

```

### Optimized Model Loading with LoRA Integration

The `load_model_and_tokenizer()` function (lines 26-80 in [`src/soup_cli/utils/unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/unsloth.py)) handles the entire initialization sequence:

1. Calls `unsloth.FastLanguageModel.from_pretrained` with the requested quantization (default **4bit**) and sequence length
2. Resolves LoRA target modules (defaults: `["q_proj","k_proj","v_proj","o_proj","gate_proj","up_proj","down_proj"]`)
3. Invokes `FastLanguageModel.get_peft_model` to attach LoRA adapters immediately

This consolidation eliminates the double-pass over model weights that occurs in the standard `transformers` + `peft` workflow.

```python
from soup_cli.utils.unsloth import load_model_and_tokenizer

model, tokenizer = load_model_and_tokenizer(
    model_name="meta-llama/Llama-2-7b-chat-hf",
    max_seq_length=2048,
    quantization="4bit",
    lora_r=64,
    lora_alpha=16,
    lora_dropout=0.05,
    target_modules="auto",
)

```

## Trainer Integration

Both the SFT and DPO trainer wrappers ([`src/soup_cli/trainer/sft.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/sft.py) and [`src/soup_cli/trainer/dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/dpo.py)) implement a private `_setup_unsloth` method. When `backend="unsloth"` is detected, this method delegates to `load_model_and_tokenizer()` instead of the standard transformers path.

The trainers receive a model that already has LoRA adapters attached, **skipping the separate `peft`-based patching step** required by the default backend. To run training:

```bash
soup train --config soup.yaml

```

## Performance Impact

By handling quantization, LoRA patching, and kernel optimization internally, the Unsloth backend reduces memory-bandwidth pressure during the training loop. Benchmarks in the repository's `benchmarks/` directory demonstrate up to **2× faster token-per-second throughput** on supported GPUs compared to the generic `transformers` path.

The performance gain stems from avoiding redundant weight copies and leveraging fused kernels that apply adapters during the forward pass rather than as a separate transformation layer.

## Switching Back to the Transformers Backend

To revert to the default behavior, change the backend field and omit Unsloth-specific flags:

```yaml
backend: transformers
training:
  quantization: none
  lora_r: 64
  # Do not include unsloth_bnb_4bit or other Unsloth-specific options

```

## Summary

- **Set `backend: unsloth`** in [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) to enable the optimized training path
- **Install Unsloth** via pip with CUDA 11.8+ requirements before use
- **Configuration validation** in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) ensures incompatible options are rejected for the wrong backend
- **Model loading** occurs through `load_model_and_tokenizer()` in [`src/soup_cli/utils/unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/unsloth.py), which uses `FastLanguageModel` for single-pass initialization
- **Trainer wrappers** in [`sft.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/sft.py) and [`dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/dpo.py) automatically handle the backend via `_setup_unsloth`
- **Performance gains** reach up to 2× faster throughput by eliminating double-pass weight processing and optimizing memory bandwidth

## Frequently Asked Questions

### What hardware is required to use the Unsloth backend in Soup?

The Unsloth backend requires a CUDA-capable GPU with Compute Capability 7.0 or higher and CUDA 11.8+. The `is_unsloth_available()` function in [`src/soup_cli/utils/unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/unsloth.py) verifies that the package can be imported, but you must ensure your PyTorch installation matches your CUDA version before training begins.

### Can I use the same LoRA configuration between Unsloth and transformers backends?

Yes, the LoRA hyperparameters (`lora_r`, `lora_alpha`, `lora_dropout`) use identical schemas across both backends. However, the Unsloth backend automatically selects optimal default target modules (`q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`) when `target_modules: auto` is specified, whereas the transformers backend may require explicit module specification depending on your model architecture.

### Why does Soup reject my configuration when I set `unsloth_bnb_4bit: true`?

The schema validation in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) enforces that `unsloth_bnb_4bit` and related quantization flags are only valid when `backend: unsloth` is explicitly declared. If you receive a validation error, verify that your YAML file specifies `backend: unsloth` and not the default `backend: transformers`.

### How do I verify that training is actually using the Unsloth optimized kernels?

Check your training logs for successful initialization messages from `load_model_and_tokenizer()`, or programmatically verify the backend before training:

```python
from soup_cli.utils.unsloth import is_unsloth_available, get_unsloth_version

assert is_unsloth_available(), "Unsloth not installed"
print(f"Using Unsloth version: {get_unsloth_version()}")

```

During training, you should observe higher GPU utilization and approximately 2× improved tokens-per-second compared to equivalent runs with `backend: transformers` in the `benchmarks/` directory comparisons.