# Benefits of Using Unsloth's Custom SwigLU and GeGLU Kernels for Faster LLM Inference

> Boost LLM inference speed with Unsloth's custom SwigLU and GeGLU kernels. Experience 2-3x faster performance and reduced memory traffic for optimized AI applications.

- Repository: [Unsloth AI/unsloth](https://github.com/unslothai/unsloth)
- Tags: performance
- Published: 2026-03-20

---

**Unsloth's custom SwigLU and GeGLU kernels fuse activation functions with MLP projections into single Triton GPU kernels, reducing memory traffic by up to 30% and delivering 2-3× faster inference compared to standard PyTorch implementations.**

The `unslothai/unsloth` repository provides optimized Triton-based replacements for standard transformer MLP activation layers. These custom kernels eliminate redundant memory passes by combining the gating mechanism, activation function, and up-projection into single fused operations, significantly accelerating large language model training and inference.

## Architectural Advantages of Fused Kernels

Unsloth's kernel architecture prioritizes memory efficiency and compute fusion over naive layer-by-layer execution. The implementation in [`unsloth/kernels/swiglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/swiglu.py) and [`unsloth/kernels/geglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/geglu.py) fundamentally restructures how MLP layers process hidden states.

### Fused Gate and Activation Operations

The **SwigLU** kernel computes the SiLU-gated activation (`f = e * sigmoid(e)`) and immediately multiplies it by the up-projection (`h = f * g`) within a single Triton kernel launch. This eliminates intermediate storage of activation outputs.

Similarly, the **GeGLU** kernels (both exact and approximate variants) evaluate the GELU-style activation (`f = 0.5·e·(1 + erf(...))` or its tanh-based approximation) and fuse it with the up-projection in one pass. According to the source code in [`unsloth/kernels/geglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/geglu.py), the `geglu_exact_forward_kernel` handles the exact implementation while an approximate version uses `triton_tanh` for faster execution.

### Reduced Memory Traffic and Bandwidth

By loading the gate tensor `e` and up-projection tensor `g` once and storing only the final result `h`, these kernels cut the number of global memory reads and writes roughly in half compared to standard PyTorch implementations. This architectural choice is critical for large hidden-state tensors common in modern LLMs.

The kernels achieve approximately **30% fewer global memory transactions** per MLP layer, directly translating to lower latency during inference.

### Triton-Generated GPU Optimization

The kernels leverage Triton to generate highly-tuned GPU code for CUDA, HIP, and Intel XPU backends. The implementation automatically selects **32-bit indexing** for smaller tensors and **64-bit indexing** (`LONG_INDEXING`) for tensors exceeding 2³¹ elements, ensuring correctness for massive models without manual intervention.

This dynamic indexing logic guarantees safe memory access across all device types while maintaining optimal performance.

### Automatic Mixed-Precision Support

Unsloth's kernels handle **FP16, BF16, and FP32** tensors transparently. The implementation casts to `float32` only where numerically necessary (such as during sigmoid or error function calculations) and writes results back in the original dtype. This preserves the model's precision budget while maximizing throughput.

The approximate GeGLU kernel specifically uses a Triton-implemented `tanh` that maps to the best available libdevice implementation for the target GPU architecture.

### LoRA-Aware Fast Path Integration

In [`unsloth/kernels/fast_lora.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/fast_lora.py), the SwigLU and GeGLU kernels integrate with `LoRA_MLP.apply` and specialized functions like `apply_lora_mlp_swiglu`, `apply_lora_mlp_geglu_exact`, and `apply_lora_mlp_geglu_approx`. This integration allows LoRA adapter weights to be fused into the same kernel execution, avoiding separate matrix multiplication calls for adapter weights.

### Quantization-Friendly Execution

The kernels accept pre-quantized weight tensors (including 4-bit NF4 formats) and automatically invoke `fast_dequantize` when needed. This keeps the entire forward pass in the optimized fast path, enabling efficient 4-bit inference without additional kernel launch overhead.

### Adaptive Block Sizing

The `BLOCK_SIZE` parameter computes as the next power-of-two of the tensor size, capped at `MAX_FUSED_SIZE = 65536`. This automatic adaptation accommodates various hidden-size configurations (such as 4096 or 8192 dimensions) without requiring manual tuning.

## Performance Impact and Speed Improvements

The architectural optimizations deliver measurable performance gains:

- **Memory bandwidth reduction**: Up to ~30% fewer global memory transactions per MLP layer
- **Kernel launch overhead**: Fused forward and backward passes require only one Triton launch per MLP instead of three separate calls (gate, up, down)
- **Empirical speed-ups**: The Unsloth project reports **2-3× faster inference** on RTX 4090 hardware compared to vanilla PyTorch implementations using `nn.GELU` or `nn.SiLU`

These benefits compound across transformer layers, significantly reducing inference latency for production deployments.

## Implementation Examples

### Using SwigLU in Llama Models

The Llama implementation in [`unsloth/models/llama.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/llama.py) replaces standard MLP forward passes with the optimized kernel:

```python
from unsloth.kernels import fast_swiglu_inference

def forward(self, hidden_states):
    # Standard MLP forward replaced by the fused SwigLU kernel

    hidden_states = fast_swiglu_inference(self.mlp, hidden_states)
    return hidden_states

```

*Source: `fast_swiglu_inference` definition in [`unsloth/models/llama.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/llama.py) (lines 612-637)*

### Adding LoRA Adapters with SwigLU

For fine-tuning scenarios, combine LoRA adapters with the fused kernel:

```python
from unsloth.kernels.fast_lora import apply_lora_mlp_swiglu

class LoRA_SwigLU_MLP(nn.Module):
    def __init__(self, base_mlp):
        super().__init__()
        self.base = base_mlp

    def forward(self, X):
        # LoRA weights are fused inside the custom kernel

        return apply_lora_mlp_swiglu(self.base, X)

```

*Source: LoRA-aware SwigLU kernel usage in [`unsloth/kernels/fast_lora.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/fast_lora.py) (lines 35-62)*

### Using GeGLU in Gemma2 Models

For Gemma2 architectures requiring exact GeGLU activation:

```python
from unsloth.kernels import fast_geglu_inference

def forward(self, hidden_states):
    hidden_states = fast_geglu_inference(self.mlp, hidden_states)
    return hidden_states

```

*Source: `fast_geglu_inference` in [`unsloth/models/gemma2.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/gemma2.py) and `geglu_exact_forward_kernel` in [`unsloth/kernels/geglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/geglu.py) (lines 63-78)*

### LoRA with Approximate GeGLU

When using the faster approximate variant:

```python
from unsloth.kernels.fast_lora import apply_lora_mlp_geglu_approx

def forward(self, X):
    return apply_lora_mlp_geglu_approx(self, X)

```

*Source: Approximate GeGLU LoRA path in [`unsloth/kernels/fast_lora.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/fast_lora.py) (lines 96-104)*

## Key Source Files and Kernel Locations

Understanding the repository structure helps when customizing or debugging kernel behavior:

- **[`unsloth/kernels/swiglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/swiglu.py)**: Triton-based forward and backward kernels for SwigLU activation
- **[`unsloth/kernels/geglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/geglu.py)**: Exact and approximate GeGLU implementations (erf-based and tanh-based)
- **[`unsloth/kernels/fast_lora.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/fast_lora.py)**: LoRA-aware wrappers that inject custom kernels into the computation graph
- **[`unsloth/models/llama.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/llama.py)**: Demonstrates SwigLU integration via `fast_swiglu_inference`
- **[`unsloth/models/gemma.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/gemma.py) & [`gemma2.py`](https://github.com/unslothai/unsloth/blob/main/gemma2.py)**: GeGLU usage examples for exact and approximate variants
- **[`unsloth/kernels/utils.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/utils.py)**: Helper functions (`calculate_settings`, device handling) underlying the custom kernels

## Summary

Unsloth's custom SwigLU and GeGLU kernels provide substantial performance benefits for transformer inference and training:

- **Fused operations** eliminate intermediate memory storage by combining activation, gating, and projection in single kernel launches
- **Reduced memory bandwidth** consumption (up to 30% fewer global memory transactions) enables higher throughput on bandwidth-constrained GPUs
- **Triton optimization** delivers automatic device-specific tuning for CUDA, HIP, and XPU backends with dynamic 32/64-bit indexing
- **Mixed-precision and quantization support** maintains FP16/BF16/FP32 compatibility while handling 4-bit NF4 weights efficiently
- **LoRA integration** allows adapter weights to fuse into kernel execution without separate matrix multiplication overhead
- **Empirical 2-3× speedups** on consumer and datacenter GPUs compared to standard PyTorch MLP implementations

## Frequently Asked Questions

### What is the difference between SwigLU and GeGLU in Unsloth?

**SwigLU** uses the SiLU activation function (sigmoid-weighted linear unit) as the gating mechanism, computing `f = e * sigmoid(e)` before multiplying by the up-projection. **GeGLU** uses the Gaussian Error Linear Unit (GELU) for gating, with implementations available in both exact form using the error function (`erf`) and an approximate form using `tanh`. According to the source code in [`unsloth/kernels/geglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/geglu.py), the exact kernel provides numerical precision while the approximate variant prioritizes speed.

### How do Unsloth's custom kernels improve inference speed?

The kernels improve speed through **memory bandwidth reduction** and **kernel fusion**. By loading gate and up-projection tensors once and writing only the final result, they cut memory traffic by roughly half compared to naive implementations. Additionally, fusing the forward and backward passes into single Triton launches eliminates kernel overhead. As reported in the Unsloth project benchmarks, this architecture delivers **2-3× faster inference** on hardware like the RTX 4090.

### Can I use Unsloth kernels with quantized models?

Yes. The kernels in [`unsloth/kernels/swiglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/swiglu.py) and [`unsloth/kernels/geglu.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/geglu.py) accept pre-quantized weight tensors, including 4-bit NF4 formats. The implementation automatically calls `fast_dequantize` when encountering quantized weights, maintaining the entire forward pass within the optimized kernel path without requiring separate dequantization steps.

### Do these kernels support LoRA fine-tuning?

Yes. The [`unsloth/kernels/fast_lora.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/kernels/fast_lora.py) module provides specialized functions including `apply_lora_mlp_swiglu`, `apply_lora_mlp_geglu_exact`, and `apply_lora_mlp_geglu_approx`. These wrappers inject LoRA adapter weights directly into the custom kernel execution, avoiding separate matrix multiplication calls for adapter weights and maintaining the performance benefits of the fused kernels during fine-tuning.