unsloth

Unified web UI for training and running open models like Qwen, DeepSeek, gpt-oss and Gemma locally.

24 articles 57.1k View on GitHub ↗
24 articles
How Unsloth Achieves 2x Faster Training Speeds on Supported Models

Unsloth achieves 2x faster model training with optimized kernels, padding free packing, and custom Triton implementations eliminating bottlenecks and wasted compute.

performance
Mar 20, 2026
Unsloth CLI Options for Model Training and Inference: A Complete Guide

Explore unsloth CLI options for model training and inference. Discover 40+ fine-tuning and 10+ text generation parameters without Python scripting. Master unsloth for efficient AI.

how-to-guide
Mar 20, 2026
How Unsloth Handles the Safetensors Format for Model Weights: A Complete Technical Guide

Unsloth prioritizes safetensors model weights with dedicated protocols and smart conflict resolution. Learn its technical approach for efficient AI.

deep-dive
Mar 20, 2026
Benefits of Using Unsloth's Custom SwigLU and GeGLU Kernels for Faster LLM Inference

Boost LLM inference speed with Unsloth's custom SwigLU and GeGLU kernels. Experience 2-3x faster performance and reduced memory traffic for optimized AI applications.

performance
Mar 20, 2026
How Unsloth's FastLoRA Kernel Differs from Standard LoRA Implementations: Fused CUDA Optimization Guide

Discover how Unsloth's FastLoRA kernel fuses computations into a single CUDA launch, overcoming memory bottlenecks and improving 4-bit quantization performance over standard LoRA.

deep-dive
Mar 20, 2026
How the cross_entropy_loss Kernel Optimizes Training in Unsloth

Unsloth's cross_entropy_loss kernel speeds up training 2-3x by running entirely on GPU. Discover how this Triton implementation optimizes AI model training and eliminates bottlenecks.

internals
Mar 20, 2026
How Unsloth Implements FP8 Training: Architecture, Benefits, and Code Examples

Discover how Unsloth implements FP8 training using custom Triton kernels and torch-ao. Achieve 50% VRAM reduction and faster throughput on H100 GPUs with a simple flag.

deep-dive
Mar 20, 2026
Which Model Architectures Does Unsloth Optimize For? A Complete Technical Guide

Discover which model architectures Unsloth optimizes including Llama, Mistral, Qwen, Gemma, and more. Accelerate your AI development with cutting-edge performance.

deep-dive
Mar 20, 2026
How Unsloth Monitors Training Progress: Loss, GPU Usage, and Real-Time Metrics

Unsloth monitors training progress with real-time metrics like loss and GPU usage. Discover how our custom TrainerCallback tracks performance and optimizes your models.

internals
Mar 20, 2026
How to Use Unsloth Data Recipes for Fine-Tuning: A Complete Guide to Data Preparation Utilities

Master Unsloth Data Recipes to transform raw documents into Hugging Face datasets. Learn data preparation utilities for efficient model fine-tuning with this comprehensive guide.

how-to-guide
Mar 20, 2026
Can Unsloth Be Used for Pre-training Large Language Models?

Discover how Unsloth accelerates large language model pre-training. Learn to use FastLanguageModel and UnslothTrainer for 2-3x faster training with gradient checkpointing and mixed precision.

how-to-guide
Mar 20, 2026
How Unsloth Optimizes RoPE Embedding Computations in Triton Kernels

Discover how Unsloth optimizes RoPE embedding computations with Triton kernels for over 10x speedup. Learn about parallel attention heads, dynamic block sizes, and lightweight autograd.

internals
Mar 20, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →