unsloth
Unified web UI for training and running open models like Qwen, DeepSeek, gpt-oss and Gemma locally.
Unsloth achieves 2x faster model training with optimized kernels, padding free packing, and custom Triton implementations eliminating bottlenecks and wasted compute.
Unsloth CLI Options for Model Training and Inference: A Complete GuideExplore unsloth CLI options for model training and inference. Discover 40+ fine-tuning and 10+ text generation parameters without Python scripting. Master unsloth for efficient AI.
How Unsloth Handles the Safetensors Format for Model Weights: A Complete Technical GuideUnsloth prioritizes safetensors model weights with dedicated protocols and smart conflict resolution. Learn its technical approach for efficient AI.
Benefits of Using Unsloth's Custom SwigLU and GeGLU Kernels for Faster LLM InferenceBoost LLM inference speed with Unsloth's custom SwigLU and GeGLU kernels. Experience 2-3x faster performance and reduced memory traffic for optimized AI applications.
How Unsloth's FastLoRA Kernel Differs from Standard LoRA Implementations: Fused CUDA Optimization GuideDiscover how Unsloth's FastLoRA kernel fuses computations into a single CUDA launch, overcoming memory bottlenecks and improving 4-bit quantization performance over standard LoRA.
How the cross_entropy_loss Kernel Optimizes Training in UnslothUnsloth's cross_entropy_loss kernel speeds up training 2-3x by running entirely on GPU. Discover how this Triton implementation optimizes AI model training and eliminates bottlenecks.
How Unsloth Implements FP8 Training: Architecture, Benefits, and Code ExamplesDiscover how Unsloth implements FP8 training using custom Triton kernels and torch-ao. Achieve 50% VRAM reduction and faster throughput on H100 GPUs with a simple flag.
Which Model Architectures Does Unsloth Optimize For? A Complete Technical GuideDiscover which model architectures Unsloth optimizes including Llama, Mistral, Qwen, Gemma, and more. Accelerate your AI development with cutting-edge performance.
How Unsloth Monitors Training Progress: Loss, GPU Usage, and Real-Time MetricsUnsloth monitors training progress with real-time metrics like loss and GPU usage. Discover how our custom TrainerCallback tracks performance and optimizes your models.
How to Use Unsloth Data Recipes for Fine-Tuning: A Complete Guide to Data Preparation UtilitiesMaster Unsloth Data Recipes to transform raw documents into Hugging Face datasets. Learn data preparation utilities for efficient model fine-tuning with this comprehensive guide.
Can Unsloth Be Used for Pre-training Large Language Models?Discover how Unsloth accelerates large language model pre-training. Learn to use FastLanguageModel and UnslothTrainer for 2-3x faster training with gradient checkpointing and mixed precision.
How Unsloth Optimizes RoPE Embedding Computations in Triton KernelsDiscover how Unsloth optimizes RoPE embedding computations with Triton kernels for over 10x speedup. Learn about parallel attention heads, dynamic block sizes, and lightweight autograd.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →