Which Model Architectures Does Unsloth Optimize For? A Complete Technical Guide

Unsloth optimizes seven major model families: Llama, Mistral, Qwen, Gemma, Phi, and DeepSeek, with specific support for variants ranging from 1B vision models to 70B+ parameter architectures.

Unsloth accelerates large language model training and inference by replacing standard PyTorch operations with hand-optimized CUDA and ROCm kernels. The library maintains a dedicated model registry at unsloth/registry that maps specific architecture classes to these optimized implementations, delivering 2-3x faster training and significantly lower memory consumption compared to vanilla Hugging Face Transformers.

How Unsloth Registers Model Architectures

Unsloth implements a metadata-driven registration system through the register_models() function in unsloth/registry/__init__.py. When the package imports, this dispatcher loads ModelMeta objects that describe every supported architecture.

Each model family has a dedicated registry file (e.g., _llama.py, _qwen.py) that defines:

  • Architecture variants: Base, instruct, and vision-capable checkpoints
  • Parameter sizes: Specific support for 1B, 3B, 7B, 32B, 72B, and larger variants
  • Kernel mappings: Which optimized CUDA kernels (RMSNorm, SwiGLU, attention) apply to each variant

If a model's family is absent from the registry, Unsloth automatically falls back to standard Transformers implementations without custom acceleration.

Complete List of Model Architectures Optimized by Unsloth

Llama Architecture Family

The Llama registry at unsloth/registry/_llama.py supports the complete Llama 3.x ecosystem:

  • Llama 3.1: 8B parameter variants including base and instruction-tuned checkpoints
  • Llama 3.2: Text-only models (1B and 3B) and vision-capable multimodal variants (11B and 90B)
  • Implementation: Uses optimized RMSNorm kernels, grouped-query attention implementations, and rotary positional embedding acceleration

Mistral Architecture Family

Mistral support is defined in unsloth/registry/_mistral.py:

  • Mistral-Small: 24B parameter models with both base and instruction-tuned variants
  • Version tracking: Supports multiple release versions (2503, 2501) with specific kernel optimizations for each checkpoint format
  • Sliding window attention: Custom CUDA kernels implement Mistral's sliding window attention pattern efficiently

Qwen Architecture Family

The Qwen registry (unsloth/registry/_qwen.py) provides extensive coverage of Alibaba's model ecosystem:

  • Qwen 2.5: Standard LLM variants (3B, 7B) with base and instruction-tuned versions
  • Qwen 2.5-VL: Vision-language models supporting 3B, 7B, 32B, and 72B parameter sizes
  • Reasoning models: QwQ (32B) and QVQ-Preview (72B) with specialized kernels for chain-of-thought generation patterns
  • Multimodal optimizations: Custom vision projection layer accelerations for image processing pipelines

Gemma Architecture Family

Google's Gemma models are registered in unsloth/registry/_gemma.py:

  • Gemma 3: Complete size range from 1B to 27B parameters, including both base and instruction-tuned checkpoints
  • Multimodal support: Vision capabilities across all parameter sizes with optimized image encoder kernels
  • Knowledge distillation: Specialized optimizations for the smaller 1B and 4B variants designed for edge deployment

Phi Architecture Family

Microsoft's Phi models are handled in unsloth/registry/_phi.py:

  • Phi-4: 1B parameter variants including base models and "mini-instruct" versions optimized for chat applications
  • Long-context optimizations: Custom attention kernel implementations for Phi-4's extended context window capabilities

DeepSeek Architecture Family

The DeepSeek registry (unsloth/registry/_deepseek.py) provides comprehensive support for both base and distilled variants:

  • DeepSeek V3: Base V3 and V3-0324 checkpoint variants
  • DeepSeek R1: Full R1 series with various sizes, R1-Zero training configurations, and R1-Distill variants
  • Distillation support:
    • R1-Distill-Llama (8B and 70B)
    • R1-Distill-Qwen (1.5B, 7B, 14B, and 32B)
  • Mixture-of-Experts optimizations: Custom kernels for DeepSeek's MoE architecture with efficient expert routing

Loading Optimized Models: Practical Code Examples

Loading Llama 3.2 with 4-bit Quantization

import unsloth
from unsloth import LlamaForCausalLM

model_id = "unsloth/Llama-3.2-1B-Instruct-bnb-4bit"
model, tokenizer = unsloth.load_pretrained(
    model_id,
    quantization="unsloth",   # Activates Unsloth's 4-bit kernels

    device_map="auto"
)

# Verify architecture registration

print(model.config.architectures)   # → ['LlamaForCausalLM']

Fine-tuning Mistral with LoRA

from unsloth import MistralForCausalLM, LoRAConfig, Trainer, TrainingArguments

model_id = "unsloth/Mistral-Small-24B-Instruct-bnb-4bit"
model, tokenizer = unsloth.load_pretrained(model_id, quantization="unsloth")

# Configure LoRA for optimized attention layers

lora_cfg = LoRAConfig(r=8, lora_alpha=16, target_modules=["q_proj", "v_proj"])
model = get_peft_config(model, lora_cfg)

training_args = TrainingArguments(
    output_dir="outputs",
    per_device_train_batch_size=4,
    learning_rate=2e-4,
    num_train_epochs=3,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=my_dataset,
    tokenizer=tokenizer,
)

trainer.train()

Running Qwen 2.5-VL Multimodal Inference

model_id = "unsloth/Qwen-2.5-7B-VL-bnb-4bit"
model, processor = unsloth.load_pretrained(
    model_id,
    quantization="unsloth",
    device_map="auto",
    trust_remote_code=True  # Required for Qwen-VL custom processor

)

# Process image and text

image = processor.image_processor("path/to/image.jpg")
inputs = processor(text="Describe this picture:", images=image, return_tensors="pt")
output = model.generate(**inputs)
print(processor.decode(output[0], skip_special_tokens=True))

Registry Implementation Details

The architecture support is implemented through a modular registry system:

When unsloth is imported, the __init__.py file triggers the registration process, patching the appropriate model classes with optimized kernels for RMSNorm, attention mechanisms, SwiGLU activation, and quantization operations.

Summary

  • Unsloth optimizes seven core architecture families: Llama, Mistral, Qwen, Gemma, Phi, and DeepSeek through its modular registry system in unsloth/registry.
  • Vision and multimodal support is available for Llama 3.2, Qwen 2.5-VL, and Gemma 3 with specialized projection layer optimizations.
  • DeepSeek coverage includes both base models (V3, R1) and extensive distillation variants (Llama and Qwen-based).
  • Registration occurs automatically via register_models() in unsloth/registry/__init__.py, which maps architecture classes to custom CUDA/ROCm kernels for 2-3x training speedup.

Frequently Asked Questions

Does Unsloth support vision-language models?

Yes, Unsloth provides optimized kernels for multimodal architectures. Specifically, the library supports Llama 3.2 Vision (11B and 90B), Qwen 2.5-VL (3B, 7B, 32B, 72B), and Gemma 3 multimodal variants (1B through 27B). These implementations include custom vision projection layer accelerations and optimized attention mechanisms for handling both image and text inputs.

How does Unsloth handle DeepSeek's Mixture-of-Experts architecture?

Unsloth implements custom CUDA kernels specifically for DeepSeek's MoE routing mechanisms in unsloth/registry/_deepseek.py. The optimization covers DeepSeek V3 and R1 series, including efficient expert routing algorithms that minimize memory overhead during training. Additionally, Unsloth provides specialized support for distilled variants like R1-Distill-Llama (8B, 70B) and R1-Distill-Qwen (1.5B, 7B, 14B, 32B), ensuring these student models retain the optimizations of their base architectures.

What happens if I try to use a model not in the Unsloth registry?

If a model architecture is absent from the Unsloth registry in unsloth/registry, the library automatically falls back to standard Hugging Face Transformers implementations without custom kernel acceleration. While the model will still function, you will not benefit from Unsloth's 2-3x training speedup or reduced memory consumption. To check if your model is supported, verify that its architecture class appears in the output of register_models() or inspect the specific registry files (e.g., _llama.py, _qwen.py) for your model family.

Does Unsloth support Phi-4 models from Microsoft?

Yes, Unsloth includes dedicated support for Microsoft's Phi-4 architecture through unsloth/registry/_phi.py. The registry currently covers the 1B parameter variants, including both base models and the "mini-instruct" versions optimized for chat applications. These implementations include custom attention kernel optimizations specifically designed for Phi-4's extended context window capabilities, allowing efficient fine-tuning even on consumer-grade hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →