How FlashAttention, NATTEN, and CuTe DSL Attention Backends Work in LTX-2
LTX-2 implements four interchangeable attention backends—FlashAttention 3, FlashAttention 4, NATTEN, and CuTe DSL—that are automatically selected based on GPU architecture and availability, with fallback to standard SDPA.
The LTX-2 video generation model delegates its most compute-intensive attention operations to highly optimized kernels. Understanding how these attention backends switch and interact is essential for performance tuning on different hardware.
Overview of the Four Attention Backends
| Backend | Implementation Location | GPU Requirement | Best For |
|---|---|---|---|
| FlashAttention 3 | ltx_core/model/transformer/attention.py |
Hopper (sm 90) | Standard transformer layers on H100 |
| FlashAttention 4 | Same file | Hopper (sm 90) or Blackwell (sm 100) | Latest generation, fused operations |
| NATTEN | ltx_core/model/video_vae/transformer/attention.py |
Any CUDA with natten installed |
3-D neighborhood attention in video VAE |
| CuTe DSL | ltx_core/model/video_vae/transformer/dsl_kernels/attn.py |
Blackwell with DSL compiled | Fused NA + stage-5 block operations |
FlashAttention 3 and FlashAttention 4: CUDA-Only Fused Attention
The FlashAttention3 and FlashAttention4 classes in ltx_core/model/transformer/attention.py wrap the flash-attn and flash-attn-4 Python packages respectively. Both execute attention in a single fused kernel to minimize memory bandwidth bottlenecks.
Key Implementation Details
- FlashAttention 3 calls
flash_attn_interfaceand converts tensors to the dtype ofvbefore computation - FlashAttention 4 uses
flash_attn_4_funcwith identical tensor layout expectations - Both classes reshape outputs to match the expected
(batch, seq, heads, dim)format
The selection logic resides in _select_primary_attention() (lines 84–92):
from ltx_core.model.transformer.attention import FlashAttention3, FlashAttention4
# Force FlashAttention 3 (Hopper required)
fa3 = FlashAttention3()
output = fa3(q, k, v, heads=8)
# Force FlashAttention 4 (Hopper or Blackwell)
fa4 = FlashAttention4()
output = fa4(q, k, v, heads=8)
NATTEN: 3-D Neighborhood Attention for Video VAE
NATTEN (Neighborhood Attention) is the default backend for 3-D spatial-temporal attention in LTX-2's video VAE. Located in ltx_core/model/video_vae/transformer/attention.py, the NattenAttention class wraps natten.na3d—a CUTLASS-based kernel optimized for local attention patterns.
How NATTEN Integrates
The import is guarded with availability checking (lines 23–30):
try:
import natten
_NATTEN_AVAILABLE = True
except ImportError:
_NATTEN_AVAILABLE = False
# Error message directs users to specific wheel:
# uv pip install "natten==0.21.7+torch2130cu132"
The NattenAttention.__call__ method (lines 62–82) enforces uniform dtype across query, key, and value tensors before delegating to natten.na3d. The kernel automatically selects the best implementation (e.g., "hopper-fna" on H100) unless overridden.
from ltx_core.model.video_vae.transformer.attention import NattenAttention
na = NattenAttention()
# attn_module is a NeighborhoodAttention3D instance
output = na(attn_module, q, k, v)
CuTe DSL: Custom Blackwell Kernel Fusion
CuTe DSL represents LTX-2's most specialized backend. Implemented in ltx_core/model/video_vae/transformer/dsl_kernels/attn.py, it executes a custom kernel that fuses neighborhood attention with the subsequent stage-5 block operation.
DSL Kernel Architecture
The DSLAttention class forwards to a custom Torch library op registered as ltx_core::na_attention_dsl (lines 51–62). A fake implementation (_na_attention_dsl_fake, lines 65–76) enables tracing and compilation when the real kernel is unavailable.
Critical to operation is the softmax bound buffer, retrieved via require_softmax_bound() (lines 33–48). This pre-computed bound enables numerical stability in the fused kernel.
from ltx_core.model.video_vae.transformer.attention import DSLAttention
dsl = DSLAttention()
if dsl._available(): # Checks na_dsl_available()
output = dsl(attn_module, q, k, v)
CuTe DSL only activates when:
BLACKWELL_DSLmode is configured (seeconfig.pylines 37–58)- The DSL kernel is compiled and
na_dsl_available()returnsTrue
Automatic Backend Selection Logic
The automatic_attention() and automatic_masked_attention() helpers (lines 9–20 of video_vae/transformer/attention.py) implement runtime backend selection:
- Detect GPU architecture via
torch.cuda.get_device_capability() - Hopper (sm 90): Prefer FlashAttention 3 → FlashAttention 4 → SDPA
- Blackwell (sm 100): Prefer FlashAttention 4 → CuTe DSL (if configured) → SDPA
- macOS: Use
MPSSdpaAttentionfor Apple Silicon - Fallback: Full SDPA priority list (
_SDPA_FULL_PRIORITY)
The selected callable is cached with @functools.cache so all AttentionOps instances share the same backend.
Runtime Backend Switching
Model authors can inspect and override the backend after instantiation:
from ltx_core.model.video_vae.transformer.attention import (
NeighborhoodAttention3D, NattenAttention, DSLAttention
)
attn = NeighborhoodAttention3D(dim=256, kernel_size=(3, 3, 3))
print(attn.attention_function.label) # "NattenAttention"
# Upgrade to DSL on Blackwell if available
if DSLAttention()._available():
attn.attention_function = DSLAttention()
print(attn.attention_function.label) # "DSLAttention"
Summary
- FlashAttention 3/4 provide fused, memory-efficient attention for standard transformer layers on Hopper and Blackwell GPUs
- NATTEN delivers hardware-optimized 3-D neighborhood attention as the video VAE default, with automatic kernel selection
- CuTe DSL enables maximum fusion by combining NA computation with downstream operations, requiring Blackwell and explicit configuration
- All backends are automatically selected based on
torch.cuda.get_device_capability()and package availability, with@functools.cacheensuring consistent behavior across modules
Frequently Asked Questions
How does LTX-2 choose between FlashAttention 3 and FlashAttention 4?
LTX-2 queries torch.cuda.get_device_capability() at runtime. On Hopper (sm 90), it prefers FlashAttention 3 if available, otherwise FlashAttention 4. On Blackwell (sm 100), it prefers FlashAttention 4 directly. Both are CUDA-only and require their respective flash-attn packages to be installed.
What happens if NATTEN is not installed?
The import guard in attention.py lines 23–30 catches the ImportError and sets _NATTEN_AVAILABLE = False. The error message provides the exact installation command: uv pip install "natten==0.21.7+torch2130cu132". Without NATTEN, the video VAE falls back to alternative attention implementations.
When should I use CuTe DSL instead of NATTEN?
Use CuTe DSL when running on Blackwell GPUs with the BLACKWELL_DSL configuration enabled and when maximum fusion is required. The DSL kernel fuses neighborhood attention with the stage-5 block, reducing kernel launch overhead. It requires compilation of the custom ltx_core::na_attention_dsl op and is not compatible with older GPU architectures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →