# Which Model Architectures Does Unsloth Optimize For? A Complete Technical Guide

> Discover which model architectures Unsloth optimizes including Llama, Mistral, Qwen, Gemma, and more. Accelerate your AI development with cutting-edge performance.

- Repository: [Unsloth AI/unsloth](https://github.com/unslothai/unsloth)
- Tags: deep-dive
- Published: 2026-03-20

---

**Unsloth optimizes seven major model families: Llama, Mistral, Qwen, Gemma, Phi, and DeepSeek, with specific support for variants ranging from 1B vision models to 70B+ parameter architectures.**

Unsloth accelerates large language model training and inference by replacing standard PyTorch operations with hand-optimized CUDA and ROCm kernels. The library maintains a dedicated model registry at `unsloth/registry` that maps specific architecture classes to these optimized implementations, delivering 2-3x faster training and significantly lower memory consumption compared to vanilla Hugging Face Transformers.

## How Unsloth Registers Model Architectures

Unsloth implements a metadata-driven registration system through the `register_models()` function in [`unsloth/registry/__init__.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/__init__.py). When the package imports, this dispatcher loads `ModelMeta` objects that describe every supported architecture.

Each model family has a dedicated registry file (e.g., [`_llama.py`](https://github.com/unslothai/unsloth/blob/main/_llama.py), [`_qwen.py`](https://github.com/unslothai/unsloth/blob/main/_qwen.py)) that defines:
- **Architecture variants**: Base, instruct, and vision-capable checkpoints
- **Parameter sizes**: Specific support for 1B, 3B, 7B, 32B, 72B, and larger variants
- **Kernel mappings**: Which optimized CUDA kernels (RMSNorm, SwiGLU, attention) apply to each variant

If a model's family is absent from the registry, Unsloth automatically falls back to standard Transformers implementations without custom acceleration.

## Complete List of Model Architectures Optimized by Unsloth

### Llama Architecture Family

The Llama registry at [`unsloth/registry/_llama.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_llama.py) supports the complete Llama 3.x ecosystem:

- **Llama 3.1**: 8B parameter variants including base and instruction-tuned checkpoints
- **Llama 3.2**: Text-only models (1B and 3B) and vision-capable multimodal variants (11B and 90B)
- **Implementation**: Uses optimized RMSNorm kernels, grouped-query attention implementations, and rotary positional embedding acceleration

### Mistral Architecture Family

Mistral support is defined in [`unsloth/registry/_mistral.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_mistral.py):

- **Mistral-Small**: 24B parameter models with both base and instruction-tuned variants
- **Version tracking**: Supports multiple release versions (2503, 2501) with specific kernel optimizations for each checkpoint format
- **Sliding window attention**: Custom CUDA kernels implement Mistral's sliding window attention pattern efficiently

### Qwen Architecture Family

The Qwen registry ([`unsloth/registry/_qwen.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_qwen.py)) provides extensive coverage of Alibaba's model ecosystem:

- **Qwen 2.5**: Standard LLM variants (3B, 7B) with base and instruction-tuned versions
- **Qwen 2.5-VL**: Vision-language models supporting 3B, 7B, 32B, and 72B parameter sizes
- **Reasoning models**: QwQ (32B) and QVQ-Preview (72B) with specialized kernels for chain-of-thought generation patterns
- **Multimodal optimizations**: Custom vision projection layer accelerations for image processing pipelines

### Gemma Architecture Family

Google's Gemma models are registered in [`unsloth/registry/_gemma.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_gemma.py):

- **Gemma 3**: Complete size range from 1B to 27B parameters, including both base and instruction-tuned checkpoints
- **Multimodal support**: Vision capabilities across all parameter sizes with optimized image encoder kernels
- **Knowledge distillation**: Specialized optimizations for the smaller 1B and 4B variants designed for edge deployment

### Phi Architecture Family

Microsoft's Phi models are handled in [`unsloth/registry/_phi.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_phi.py):

- **Phi-4**: 1B parameter variants including base models and "mini-instruct" versions optimized for chat applications
- **Long-context optimizations**: Custom attention kernel implementations for Phi-4's extended context window capabilities

### DeepSeek Architecture Family

The DeepSeek registry ([`unsloth/registry/_deepseek.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_deepseek.py)) provides comprehensive support for both base and distilled variants:

- **DeepSeek V3**: Base V3 and V3-0324 checkpoint variants
- **DeepSeek R1**: Full R1 series with various sizes, R1-Zero training configurations, and R1-Distill variants
- **Distillation support**: 
  - R1-Distill-Llama (8B and 70B)
  - R1-Distill-Qwen (1.5B, 7B, 14B, and 32B)
- **Mixture-of-Experts optimizations**: Custom kernels for DeepSeek's MoE architecture with efficient expert routing

## Loading Optimized Models: Practical Code Examples

### Loading Llama 3.2 with 4-bit Quantization

```python
import unsloth
from unsloth import LlamaForCausalLM

model_id = "unsloth/Llama-3.2-1B-Instruct-bnb-4bit"
model, tokenizer = unsloth.load_pretrained(
    model_id,
    quantization="unsloth",   # Activates Unsloth's 4-bit kernels

    device_map="auto"
)

# Verify architecture registration

print(model.config.architectures)   # → ['LlamaForCausalLM']

```

### Fine-tuning Mistral with LoRA

```python
from unsloth import MistralForCausalLM, LoRAConfig, Trainer, TrainingArguments

model_id = "unsloth/Mistral-Small-24B-Instruct-bnb-4bit"
model, tokenizer = unsloth.load_pretrained(model_id, quantization="unsloth")

# Configure LoRA for optimized attention layers

lora_cfg = LoRAConfig(r=8, lora_alpha=16, target_modules=["q_proj", "v_proj"])
model = get_peft_config(model, lora_cfg)

training_args = TrainingArguments(
    output_dir="outputs",
    per_device_train_batch_size=4,
    learning_rate=2e-4,
    num_train_epochs=3,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=my_dataset,
    tokenizer=tokenizer,
)

trainer.train()

```

### Running Qwen 2.5-VL Multimodal Inference

```python
model_id = "unsloth/Qwen-2.5-7B-VL-bnb-4bit"
model, processor = unsloth.load_pretrained(
    model_id,
    quantization="unsloth",
    device_map="auto",
    trust_remote_code=True  # Required for Qwen-VL custom processor

)

# Process image and text

image = processor.image_processor("path/to/image.jpg")
inputs = processor(text="Describe this picture:", images=image, return_tensors="pt")
output = model.generate(**inputs)
print(processor.decode(output[0], skip_special_tokens=True))

```

## Registry Implementation Details

The architecture support is implemented through a modular registry system:

- **[`unsloth/registry/__init__.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/__init__.py)**: Contains the central `register_models()` dispatcher that aggregates all architecture definitions
- **[`unsloth/registry/_llama.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_llama.py)**: Llama 3.1 and 3.2 variants including vision models
- **[`unsloth/registry/_mistral.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_mistral.py)**: Mistral-Small 24B configurations
- **[`unsloth/registry/_qwen.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_qwen.py)**: Qwen 2.5, Qwen-VL, QwQ, and QVQ series
- **[`unsloth/registry/_gemma.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_gemma.py)**: Gemma 3 1B through 27B variants
- **[`unsloth/registry/_phi.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_phi.py)**: Phi-4 1B models
- **[`unsloth/registry/_deepseek.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_deepseek.py)**: DeepSeek V3, R1, and distilled variants

When `unsloth` is imported, the [`__init__.py`](https://github.com/unslothai/unsloth/blob/main/__init__.py) file triggers the registration process, patching the appropriate model classes with optimized kernels for RMSNorm, attention mechanisms, SwiGLU activation, and quantization operations.

## Summary

- Unsloth optimizes **seven core architecture families**: Llama, Mistral, Qwen, Gemma, Phi, and DeepSeek through its modular registry system in `unsloth/registry`.
- **Vision and multimodal support** is available for Llama 3.2, Qwen 2.5-VL, and Gemma 3 with specialized projection layer optimizations.
- **DeepSeek coverage** includes both base models (V3, R1) and extensive distillation variants (Llama and Qwen-based).
- Registration occurs automatically via `register_models()` in [`unsloth/registry/__init__.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/__init__.py), which maps architecture classes to custom CUDA/ROCm kernels for 2-3x training speedup.

## Frequently Asked Questions

### Does Unsloth support vision-language models?

Yes, Unsloth provides optimized kernels for multimodal architectures. Specifically, the library supports Llama 3.2 Vision (11B and 90B), Qwen 2.5-VL (3B, 7B, 32B, 72B), and Gemma 3 multimodal variants (1B through 27B). These implementations include custom vision projection layer accelerations and optimized attention mechanisms for handling both image and text inputs.

### How does Unsloth handle DeepSeek's Mixture-of-Experts architecture?

Unsloth implements custom CUDA kernels specifically for DeepSeek's MoE routing mechanisms in [`unsloth/registry/_deepseek.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_deepseek.py). The optimization covers DeepSeek V3 and R1 series, including efficient expert routing algorithms that minimize memory overhead during training. Additionally, Unsloth provides specialized support for distilled variants like R1-Distill-Llama (8B, 70B) and R1-Distill-Qwen (1.5B, 7B, 14B, 32B), ensuring these student models retain the optimizations of their base architectures.

### What happens if I try to use a model not in the Unsloth registry?

If a model architecture is absent from the Unsloth registry in `unsloth/registry`, the library automatically falls back to standard Hugging Face Transformers implementations without custom kernel acceleration. While the model will still function, you will not benefit from Unsloth's 2-3x training speedup or reduced memory consumption. To check if your model is supported, verify that its architecture class appears in the output of `register_models()` or inspect the specific registry files (e.g., [`_llama.py`](https://github.com/unslothai/unsloth/blob/main/_llama.py), [`_qwen.py`](https://github.com/unslothai/unsloth/blob/main/_qwen.py)) for your model family.

### Does Unsloth support Phi-4 models from Microsoft?

Yes, Unsloth includes dedicated support for Microsoft's Phi-4 architecture through [`unsloth/registry/_phi.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/registry/_phi.py). The registry currently covers the 1B parameter variants, including both base models and the "mini-instruct" versions optimized for chat applications. These implementations include custom attention kernel optimizations specifically designed for Phi-4's extended context window capabilities, allowing efficient fine-tuning even on consumer-grade hardware.