Benefits of Using Unsloth's Custom SwigLU and GeGLU Kernels for Faster LLM Inference
Unsloth's custom SwigLU and GeGLU kernels fuse activation functions with MLP projections into single Triton GPU kernels, reducing memory traffic by up to 30% and delivering 2-3× faster inference compared to standard PyTorch implementations.
The unslothai/unsloth repository provides optimized Triton-based replacements for standard transformer MLP activation layers. These custom kernels eliminate redundant memory passes by combining the gating mechanism, activation function, and up-projection into single fused operations, significantly accelerating large language model training and inference.
Architectural Advantages of Fused Kernels
Unsloth's kernel architecture prioritizes memory efficiency and compute fusion over naive layer-by-layer execution. The implementation in unsloth/kernels/swiglu.py and unsloth/kernels/geglu.py fundamentally restructures how MLP layers process hidden states.
Fused Gate and Activation Operations
The SwigLU kernel computes the SiLU-gated activation (f = e * sigmoid(e)) and immediately multiplies it by the up-projection (h = f * g) within a single Triton kernel launch. This eliminates intermediate storage of activation outputs.
Similarly, the GeGLU kernels (both exact and approximate variants) evaluate the GELU-style activation (f = 0.5·e·(1 + erf(...)) or its tanh-based approximation) and fuse it with the up-projection in one pass. According to the source code in unsloth/kernels/geglu.py, the geglu_exact_forward_kernel handles the exact implementation while an approximate version uses triton_tanh for faster execution.
Reduced Memory Traffic and Bandwidth
By loading the gate tensor e and up-projection tensor g once and storing only the final result h, these kernels cut the number of global memory reads and writes roughly in half compared to standard PyTorch implementations. This architectural choice is critical for large hidden-state tensors common in modern LLMs.
The kernels achieve approximately 30% fewer global memory transactions per MLP layer, directly translating to lower latency during inference.
Triton-Generated GPU Optimization
The kernels leverage Triton to generate highly-tuned GPU code for CUDA, HIP, and Intel XPU backends. The implementation automatically selects 32-bit indexing for smaller tensors and 64-bit indexing (LONG_INDEXING) for tensors exceeding 2³¹ elements, ensuring correctness for massive models without manual intervention.
This dynamic indexing logic guarantees safe memory access across all device types while maintaining optimal performance.
Automatic Mixed-Precision Support
Unsloth's kernels handle FP16, BF16, and FP32 tensors transparently. The implementation casts to float32 only where numerically necessary (such as during sigmoid or error function calculations) and writes results back in the original dtype. This preserves the model's precision budget while maximizing throughput.
The approximate GeGLU kernel specifically uses a Triton-implemented tanh that maps to the best available libdevice implementation for the target GPU architecture.
LoRA-Aware Fast Path Integration
In unsloth/kernels/fast_lora.py, the SwigLU and GeGLU kernels integrate with LoRA_MLP.apply and specialized functions like apply_lora_mlp_swiglu, apply_lora_mlp_geglu_exact, and apply_lora_mlp_geglu_approx. This integration allows LoRA adapter weights to be fused into the same kernel execution, avoiding separate matrix multiplication calls for adapter weights.
Quantization-Friendly Execution
The kernels accept pre-quantized weight tensors (including 4-bit NF4 formats) and automatically invoke fast_dequantize when needed. This keeps the entire forward pass in the optimized fast path, enabling efficient 4-bit inference without additional kernel launch overhead.
Adaptive Block Sizing
The BLOCK_SIZE parameter computes as the next power-of-two of the tensor size, capped at MAX_FUSED_SIZE = 65536. This automatic adaptation accommodates various hidden-size configurations (such as 4096 or 8192 dimensions) without requiring manual tuning.
Performance Impact and Speed Improvements
The architectural optimizations deliver measurable performance gains:
- Memory bandwidth reduction: Up to ~30% fewer global memory transactions per MLP layer
- Kernel launch overhead: Fused forward and backward passes require only one Triton launch per MLP instead of three separate calls (gate, up, down)
- Empirical speed-ups: The Unsloth project reports 2-3× faster inference on RTX 4090 hardware compared to vanilla PyTorch implementations using
nn.GELUornn.SiLU
These benefits compound across transformer layers, significantly reducing inference latency for production deployments.
Implementation Examples
Using SwigLU in Llama Models
The Llama implementation in unsloth/models/llama.py replaces standard MLP forward passes with the optimized kernel:
from unsloth.kernels import fast_swiglu_inference
def forward(self, hidden_states):
# Standard MLP forward replaced by the fused SwigLU kernel
hidden_states = fast_swiglu_inference(self.mlp, hidden_states)
return hidden_states
Source: fast_swiglu_inference definition in unsloth/models/llama.py (lines 612-637)
Adding LoRA Adapters with SwigLU
For fine-tuning scenarios, combine LoRA adapters with the fused kernel:
from unsloth.kernels.fast_lora import apply_lora_mlp_swiglu
class LoRA_SwigLU_MLP(nn.Module):
def __init__(self, base_mlp):
super().__init__()
self.base = base_mlp
def forward(self, X):
# LoRA weights are fused inside the custom kernel
return apply_lora_mlp_swiglu(self.base, X)
Source: LoRA-aware SwigLU kernel usage in unsloth/kernels/fast_lora.py (lines 35-62)
Using GeGLU in Gemma2 Models
For Gemma2 architectures requiring exact GeGLU activation:
from unsloth.kernels import fast_geglu_inference
def forward(self, hidden_states):
hidden_states = fast_geglu_inference(self.mlp, hidden_states)
return hidden_states
Source: fast_geglu_inference in unsloth/models/gemma2.py and geglu_exact_forward_kernel in unsloth/kernels/geglu.py (lines 63-78)
LoRA with Approximate GeGLU
When using the faster approximate variant:
from unsloth.kernels.fast_lora import apply_lora_mlp_geglu_approx
def forward(self, X):
return apply_lora_mlp_geglu_approx(self, X)
Source: Approximate GeGLU LoRA path in unsloth/kernels/fast_lora.py (lines 96-104)
Key Source Files and Kernel Locations
Understanding the repository structure helps when customizing or debugging kernel behavior:
unsloth/kernels/swiglu.py: Triton-based forward and backward kernels for SwigLU activationunsloth/kernels/geglu.py: Exact and approximate GeGLU implementations (erf-based and tanh-based)unsloth/kernels/fast_lora.py: LoRA-aware wrappers that inject custom kernels into the computation graphunsloth/models/llama.py: Demonstrates SwigLU integration viafast_swiglu_inferenceunsloth/models/gemma.py&gemma2.py: GeGLU usage examples for exact and approximate variantsunsloth/kernels/utils.py: Helper functions (calculate_settings, device handling) underlying the custom kernels
Summary
Unsloth's custom SwigLU and GeGLU kernels provide substantial performance benefits for transformer inference and training:
- Fused operations eliminate intermediate memory storage by combining activation, gating, and projection in single kernel launches
- Reduced memory bandwidth consumption (up to 30% fewer global memory transactions) enables higher throughput on bandwidth-constrained GPUs
- Triton optimization delivers automatic device-specific tuning for CUDA, HIP, and XPU backends with dynamic 32/64-bit indexing
- Mixed-precision and quantization support maintains FP16/BF16/FP32 compatibility while handling 4-bit NF4 weights efficiently
- LoRA integration allows adapter weights to fuse into kernel execution without separate matrix multiplication overhead
- Empirical 2-3× speedups on consumer and datacenter GPUs compared to standard PyTorch MLP implementations
Frequently Asked Questions
What is the difference between SwigLU and GeGLU in Unsloth?
SwigLU uses the SiLU activation function (sigmoid-weighted linear unit) as the gating mechanism, computing f = e * sigmoid(e) before multiplying by the up-projection. GeGLU uses the Gaussian Error Linear Unit (GELU) for gating, with implementations available in both exact form using the error function (erf) and an approximate form using tanh. According to the source code in unsloth/kernels/geglu.py, the exact kernel provides numerical precision while the approximate variant prioritizes speed.
How do Unsloth's custom kernels improve inference speed?
The kernels improve speed through memory bandwidth reduction and kernel fusion. By loading gate and up-projection tensors once and writing only the final result, they cut memory traffic by roughly half compared to naive implementations. Additionally, fusing the forward and backward passes into single Triton launches eliminates kernel overhead. As reported in the Unsloth project benchmarks, this architecture delivers 2-3× faster inference on hardware like the RTX 4090.
Can I use Unsloth kernels with quantized models?
Yes. The kernels in unsloth/kernels/swiglu.py and unsloth/kernels/geglu.py accept pre-quantized weight tensors, including 4-bit NF4 formats. The implementation automatically calls fast_dequantize when encountering quantized weights, maintaining the entire forward pass within the optimized kernel path without requiring separate dequantization steps.
Do these kernels support LoRA fine-tuning?
Yes. The unsloth/kernels/fast_lora.py module provides specialized functions including apply_lora_mlp_swiglu, apply_lora_mlp_geglu_exact, and apply_lora_mlp_geglu_approx. These wrappers inject LoRA adapter weights directly into the custom kernel execution, avoiding separate matrix multiplication calls for adapter weights and maintaining the performance benefits of the fused kernels during fine-tuning.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →