How FlashAttention, NATTEN, and CuTe DSL Attention Backends Work in LTX-2

LTX-2 implements four interchangeable attention backends—FlashAttention 3, FlashAttention 4, NATTEN, and CuTe DSL—that are automatically selected based on GPU architecture and availability, with fallback to standard SDPA.

The LTX-2 video generation model delegates its most compute-intensive attention operations to highly optimized kernels. Understanding how these attention backends switch and interact is essential for performance tuning on different hardware.

Overview of the Four Attention Backends

Backend Implementation Location GPU Requirement Best For
FlashAttention 3 ltx_core/model/transformer/attention.py Hopper (sm 90) Standard transformer layers on H100
FlashAttention 4 Same file Hopper (sm 90) or Blackwell (sm 100) Latest generation, fused operations
NATTEN ltx_core/model/video_vae/transformer/attention.py Any CUDA with natten installed 3-D neighborhood attention in video VAE
CuTe DSL ltx_core/model/video_vae/transformer/dsl_kernels/attn.py Blackwell with DSL compiled Fused NA + stage-5 block operations

FlashAttention 3 and FlashAttention 4: CUDA-Only Fused Attention

The FlashAttention3 and FlashAttention4 classes in ltx_core/model/transformer/attention.py wrap the flash-attn and flash-attn-4 Python packages respectively. Both execute attention in a single fused kernel to minimize memory bandwidth bottlenecks.

Key Implementation Details

  • FlashAttention 3 calls flash_attn_interface and converts tensors to the dtype of v before computation
  • FlashAttention 4 uses flash_attn_4_func with identical tensor layout expectations
  • Both classes reshape outputs to match the expected (batch, seq, heads, dim) format

The selection logic resides in _select_primary_attention() (lines 84–92):

from ltx_core.model.transformer.attention import FlashAttention3, FlashAttention4

# Force FlashAttention 3 (Hopper required)

fa3 = FlashAttention3()
output = fa3(q, k, v, heads=8)

# Force FlashAttention 4 (Hopper or Blackwell)

fa4 = FlashAttention4()
output = fa4(q, k, v, heads=8)

NATTEN: 3-D Neighborhood Attention for Video VAE

NATTEN (Neighborhood Attention) is the default backend for 3-D spatial-temporal attention in LTX-2's video VAE. Located in ltx_core/model/video_vae/transformer/attention.py, the NattenAttention class wraps natten.na3d—a CUTLASS-based kernel optimized for local attention patterns.

How NATTEN Integrates

The import is guarded with availability checking (lines 23–30):

try:
    import natten
    _NATTEN_AVAILABLE = True
except ImportError:
    _NATTEN_AVAILABLE = False
    # Error message directs users to specific wheel:

    # uv pip install "natten==0.21.7+torch2130cu132"

The NattenAttention.__call__ method (lines 62–82) enforces uniform dtype across query, key, and value tensors before delegating to natten.na3d. The kernel automatically selects the best implementation (e.g., "hopper-fna" on H100) unless overridden.

from ltx_core.model.video_vae.transformer.attention import NattenAttention

na = NattenAttention()

# attn_module is a NeighborhoodAttention3D instance

output = na(attn_module, q, k, v)

CuTe DSL: Custom Blackwell Kernel Fusion

CuTe DSL represents LTX-2's most specialized backend. Implemented in ltx_core/model/video_vae/transformer/dsl_kernels/attn.py, it executes a custom kernel that fuses neighborhood attention with the subsequent stage-5 block operation.

DSL Kernel Architecture

The DSLAttention class forwards to a custom Torch library op registered as ltx_core::na_attention_dsl (lines 51–62). A fake implementation (_na_attention_dsl_fake, lines 65–76) enables tracing and compilation when the real kernel is unavailable.

Critical to operation is the softmax bound buffer, retrieved via require_softmax_bound() (lines 33–48). This pre-computed bound enables numerical stability in the fused kernel.

from ltx_core.model.video_vae.transformer.attention import DSLAttention

dsl = DSLAttention()
if dsl._available():  # Checks na_dsl_available()

    output = dsl(attn_module, q, k, v)

CuTe DSL only activates when:

  1. BLACKWELL_DSL mode is configured (see config.py lines 37–58)
  2. The DSL kernel is compiled and na_dsl_available() returns True

Automatic Backend Selection Logic

The automatic_attention() and automatic_masked_attention() helpers (lines 9–20 of video_vae/transformer/attention.py) implement runtime backend selection:

  1. Detect GPU architecture via torch.cuda.get_device_capability()
  2. Hopper (sm 90): Prefer FlashAttention 3 → FlashAttention 4 → SDPA
  3. Blackwell (sm 100): Prefer FlashAttention 4 → CuTe DSL (if configured) → SDPA
  4. macOS: Use MPSSdpaAttention for Apple Silicon
  5. Fallback: Full SDPA priority list (_SDPA_FULL_PRIORITY)

The selected callable is cached with @functools.cache so all AttentionOps instances share the same backend.

Runtime Backend Switching

Model authors can inspect and override the backend after instantiation:

from ltx_core.model.video_vae.transformer.attention import (
    NeighborhoodAttention3D, NattenAttention, DSLAttention
)

attn = NeighborhoodAttention3D(dim=256, kernel_size=(3, 3, 3))
print(attn.attention_function.label)   # "NattenAttention"

# Upgrade to DSL on Blackwell if available

if DSLAttention()._available():
    attn.attention_function = DSLAttention()
    print(attn.attention_function.label)   # "DSLAttention"

Summary

  • FlashAttention 3/4 provide fused, memory-efficient attention for standard transformer layers on Hopper and Blackwell GPUs
  • NATTEN delivers hardware-optimized 3-D neighborhood attention as the video VAE default, with automatic kernel selection
  • CuTe DSL enables maximum fusion by combining NA computation with downstream operations, requiring Blackwell and explicit configuration
  • All backends are automatically selected based on torch.cuda.get_device_capability() and package availability, with @functools.cache ensuring consistent behavior across modules

Frequently Asked Questions

How does LTX-2 choose between FlashAttention 3 and FlashAttention 4?

LTX-2 queries torch.cuda.get_device_capability() at runtime. On Hopper (sm 90), it prefers FlashAttention 3 if available, otherwise FlashAttention 4. On Blackwell (sm 100), it prefers FlashAttention 4 directly. Both are CUDA-only and require their respective flash-attn packages to be installed.

What happens if NATTEN is not installed?

The import guard in attention.py lines 23–30 catches the ImportError and sets _NATTEN_AVAILABLE = False. The error message provides the exact installation command: uv pip install "natten==0.21.7+torch2130cu132". Without NATTEN, the video VAE falls back to alternative attention implementations.

When should I use CuTe DSL instead of NATTEN?

Use CuTe DSL when running on Blackwell GPUs with the BLACKWELL_DSL configuration enabled and when maximum fusion is required. The DSL kernel fuses neighborhood attention with the stage-5 block, reducing kernel launch overhead. It requires compilation of the custom ltx_core::na_attention_dsl op and is not compatible with older GPU architectures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →