# How FlashAttention, NATTEN, and CuTe DSL Attention Backends Work in LTX-2

> Explore how LTX-2's interchangeable attention backends like FlashAttention, NATTEN, and CuTe DSL optimize performance. Discover automatic selection and fallback mechanisms for seamless GPU integration.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: internals
- Published: 2026-08-20

---

**LTX-2 implements four interchangeable attention backends—FlashAttention 3, FlashAttention 4, NATTEN, and CuTe DSL—that are automatically selected based on GPU architecture and availability, with fallback to standard SDPA.**

The LTX-2 video generation model delegates its most compute-intensive attention operations to highly optimized kernels. Understanding how these **attention backends** switch and interact is essential for performance tuning on different hardware.

## Overview of the Four Attention Backends

| Backend | Implementation Location | GPU Requirement | Best For |
|---------|------------------------|-----------------|----------|
| **FlashAttention 3** | [`ltx_core/model/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/transformer/attention.py) | Hopper (sm 90) | Standard transformer layers on H100 |
| **FlashAttention 4** | Same file | Hopper (sm 90) or Blackwell (sm 100) | Latest generation, fused operations |
| **NATTEN** | [`ltx_core/model/video_vae/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/attention.py) | Any CUDA with `natten` installed | 3-D neighborhood attention in video VAE |
| **CuTe DSL** | [`ltx_core/model/video_vae/transformer/dsl_kernels/attn.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/dsl_kernels/attn.py) | Blackwell with DSL compiled | Fused NA + stage-5 block operations |

## FlashAttention 3 and FlashAttention 4: CUDA-Only Fused Attention

The `FlashAttention3` and `FlashAttention4` classes in [`ltx_core/model/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/transformer/attention.py) wrap the `flash-attn` and `flash-attn-4` Python packages respectively. Both execute attention in a single fused kernel to minimize memory bandwidth bottlenecks.

### Key Implementation Details

- **FlashAttention 3** calls `flash_attn_interface` and converts tensors to the dtype of `v` before computation
- **FlashAttention 4** uses `flash_attn_4_func` with identical tensor layout expectations
- Both classes reshape outputs to match the expected `(batch, seq, heads, dim)` format

The selection logic resides in `_select_primary_attention()` (lines 84–92):

```python
from ltx_core.model.transformer.attention import FlashAttention3, FlashAttention4

# Force FlashAttention 3 (Hopper required)

fa3 = FlashAttention3()
output = fa3(q, k, v, heads=8)

# Force FlashAttention 4 (Hopper or Blackwell)

fa4 = FlashAttention4()
output = fa4(q, k, v, heads=8)

```

## NATTEN: 3-D Neighborhood Attention for Video VAE

**NATTEN** (Neighborhood Attention) is the default backend for 3-D spatial-temporal attention in LTX-2's video VAE. Located in [`ltx_core/model/video_vae/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/attention.py), the `NattenAttention` class wraps `natten.na3d`—a CUTLASS-based kernel optimized for local attention patterns.

### How NATTEN Integrates

The import is guarded with availability checking (lines 23–30):

```python
try:
    import natten
    _NATTEN_AVAILABLE = True
except ImportError:
    _NATTEN_AVAILABLE = False
    # Error message directs users to specific wheel:

    # uv pip install "natten==0.21.7+torch2130cu132"

```

The `NattenAttention.__call__` method (lines 62–82) enforces uniform dtype across query, key, and value tensors before delegating to `natten.na3d`. The kernel automatically selects the best implementation (e.g., "hopper-fna" on H100) unless overridden.

```python
from ltx_core.model.video_vae.transformer.attention import NattenAttention

na = NattenAttention()

# attn_module is a NeighborhoodAttention3D instance

output = na(attn_module, q, k, v)

```

## CuTe DSL: Custom Blackwell Kernel Fusion

**CuTe DSL** represents LTX-2's most specialized backend. Implemented in [`ltx_core/model/video_vae/transformer/dsl_kernels/attn.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/dsl_kernels/attn.py), it executes a custom kernel that fuses neighborhood attention with the subsequent stage-5 block operation.

### DSL Kernel Architecture

The `DSLAttention` class forwards to a custom Torch library op registered as `ltx_core::na_attention_dsl` (lines 51–62). A fake implementation (`_na_attention_dsl_fake`, lines 65–76) enables tracing and compilation when the real kernel is unavailable.

Critical to operation is the **softmax bound buffer**, retrieved via `require_softmax_bound()` (lines 33–48). This pre-computed bound enables numerical stability in the fused kernel.

```python
from ltx_core.model.video_vae.transformer.attention import DSLAttention

dsl = DSLAttention()
if dsl._available():  # Checks na_dsl_available()

    output = dsl(attn_module, q, k, v)

```

CuTe DSL only activates when:
1. `BLACKWELL_DSL` mode is configured (see [`config.py`](https://github.com/Lightricks/LTX-2/blob/main/config.py) lines 37–58)
2. The DSL kernel is compiled and `na_dsl_available()` returns `True`

## Automatic Backend Selection Logic

The `automatic_attention()` and `automatic_masked_attention()` helpers (lines 9–20 of [`video_vae/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/video_vae/transformer/attention.py)) implement runtime backend selection:

1. **Detect GPU architecture** via `torch.cuda.get_device_capability()`
2. **Hopper (sm 90)**: Prefer FlashAttention 3 → FlashAttention 4 → SDPA
3. **Blackwell (sm 100)**: Prefer FlashAttention 4 → CuTe DSL (if configured) → SDPA
4. **macOS**: Use `MPSSdpaAttention` for Apple Silicon
5. **Fallback**: Full SDPA priority list (`_SDPA_FULL_PRIORITY`)

The selected callable is cached with `@functools.cache` so all `AttentionOps` instances share the same backend.

## Runtime Backend Switching

Model authors can inspect and override the backend after instantiation:

```python
from ltx_core.model.video_vae.transformer.attention import (
    NeighborhoodAttention3D, NattenAttention, DSLAttention
)

attn = NeighborhoodAttention3D(dim=256, kernel_size=(3, 3, 3))
print(attn.attention_function.label)   # "NattenAttention"

# Upgrade to DSL on Blackwell if available

if DSLAttention()._available():
    attn.attention_function = DSLAttention()
    print(attn.attention_function.label)   # "DSLAttention"

```

## Summary

- **FlashAttention 3/4** provide fused, memory-efficient attention for standard transformer layers on Hopper and Blackwell GPUs
- **NATTEN** delivers hardware-optimized 3-D neighborhood attention as the video VAE default, with automatic kernel selection
- **CuTe DSL** enables maximum fusion by combining NA computation with downstream operations, requiring Blackwell and explicit configuration
- All backends are **automatically selected** based on `torch.cuda.get_device_capability()` and package availability, with `@functools.cache` ensuring consistent behavior across modules

## Frequently Asked Questions

### How does LTX-2 choose between FlashAttention 3 and FlashAttention 4?

LTX-2 queries `torch.cuda.get_device_capability()` at runtime. On Hopper (sm 90), it prefers FlashAttention 3 if available, otherwise FlashAttention 4. On Blackwell (sm 100), it prefers FlashAttention 4 directly. Both are CUDA-only and require their respective `flash-attn` packages to be installed.

### What happens if NATTEN is not installed?

The import guard in [`attention.py`](https://github.com/Lightricks/LTX-2/blob/main/attention.py) lines 23–30 catches the `ImportError` and sets `_NATTEN_AVAILABLE = False`. The error message provides the exact installation command: `uv pip install "natten==0.21.7+torch2130cu132"`. Without NATTEN, the video VAE falls back to alternative attention implementations.

### When should I use CuTe DSL instead of NATTEN?

Use CuTe DSL when running on Blackwell GPUs with the `BLACKWELL_DSL` configuration enabled and when maximum fusion is required. The DSL kernel fuses neighborhood attention with the stage-5 block, reducing kernel launch overhead. It requires compilation of the custom `ltx_core::na_attention_dsl` op and is not compatible with older GPU architectures.