DiffVAE Backend Options in LTX-2: NATTEN, Triton, Eager SDPA & Blackwell DSL Explained

LTX-2 supports four DiffVAE backends: NATTEN (CUDA kernel library), Triton (JIT-compiled fallback), Eager SDPA (pure PyTorch fallback), and Blackwell DSL (CuTe DSL kernel for Blackwell GPUs), selectable via the NAttentionKind enum. According to the Lightricks/LTX-2 source code, the DiffVAE architecture uses a pluggable neighborhood-attention system that automatically falls back through these options based on hardware availability and user configuration.

The video-latent diffusion VAE in LTX-2 relies heavily on 3D neighborhood attention for efficient spatiotemporal processing. The backend selection mechanism—implemented across config.py, attention.py, and apply.py—lets users prioritize speed (NATTEN), compatibility (Triton/SDPA), or cutting-edge hardware acceleration (Blackwell DSL). Each backend targets different deployment scenarios, from consumer GPUs to datacenter Blackwell clusters.

The Four DiffVAE Backend Options

LTX-2 defines all backend choices in ltx_core/model/video_vae/transformer/config.py through the NAttentionKind enum. Here's how each option works:

NATTEN: The Primary CUDA Kernel Backend

NAttentionKind.NATTEN is the default high-performance path when the natten Python package is installed. This backend calls into optimized CUDA kernels, with optional kernel-level selection (e.g., "cutlass-fna") via the backend argument of NattenAttention.

  • Source location: ltx_core/model/video_vae/transformer/config.py, lines 15-21
  • Best for: Production inference on CUDA GPUs with natten wheels available
  • Performance: Fastest general-purpose option; kernel auto-tuned for spatial and temporal dimensions

The NATTEN backend is what most users will run in practice, provided they install the underlying natten dependency.

Triton: The JIT-Compiled Fallback

NAttentionKind.TRITON activates when natten is unavailable but Triton is present. LTX-2's implementation in fallback_na/triton_na3d.py exposes the same Python API as NATTEN, just compiled through the Triton JIT compiler at runtime.

This fallback ensures the DiffVAE model remains functional across diverse hardware without requiring recompilation.

Eager SDPA: The Device-Agnostic Fallback

NAttentionKind.EAGER_SDPA is the last-resort fallback using PyTorch's native scaled_dot_product_attention. It requires no external dependencies beyond PyTorch itself.

This backend guarantees the DiffVAE model will run anywhere PyTorch does, though latencies may be unacceptable for real-time video generation.

Blackwell DSL: The Specialized Datacenter Kernel

NAttentionKind.BLACKWELL_DSL is an experimental, hardware-specific backend for NVIDIA Blackwell (B100/B200) GPUs. It compiles a fused CuTe DSL kernel that handles the complete neighborhood-attention block, including a deferred stage-5 computation unique to the DiffVAE architecture.

This backend is not selected automatically—users must request it through the preset system.

How LTX-2 Selects Your DiffVAE Backend

The backend selection pipeline involves three stages, each implemented in specific source files:

1. User Preset Resolution (DiffVAEMode)

Users start with high-level presets defined in config.py. Calling DiffVAEMode.resolve() (lines 70-101) expands these into concrete DiffVAEConfig objects:

Preset Backend Behavior
CHUNKED_EAGER Chunked processing, no compilation, NATTEN if available
CHUNKED_COMPILE Chunked processing with torch.compile, NATTEN preferred
COMBINED_COMPILE Combined spatial-temporal compilation, NATTEN preferred
BLACKWELL_DSL Forces Blackwell DSL kernel, ignores NATTEN availability
from ltx_core.model.video_vae.transformer.config import DiffVAEMode

# Select Blackwell DSL path explicitly

vae_cfg = DiffVAEMode.BLACKWELL_DSL.resolve()
print(vae_cfg)

# DiffVAEConfig(

#     block=DiffVAEBlockKind.BLACKWELL_DSL,

#     w_chunks=1,

#     natten_backend=None,

#     attention=NAttentionKind.BLACKWELL_DSL,

#     compile_blocks=True,

#     compile_det_stages=True

# )

2. Backend Injection (configure_natten_backend)

The function configure_natten_backend(module_root, backend) in ltx_core/model/video_vae/transformer/apply.py (line 108) walks the module tree and writes the selected backend into each NeighborhoodAttention3D instance:

from ltx_core.model.video_vae.transformer.attention import NAttentionKind
from ltx_core.model.video_vae.transformer.apply import configure_natten_backend

# Force Triton fallback across all NA layers

configure_natten_backend(decoder, backend=NAttentionKind.TRITON.value)

For BLACKWELL_DSL requests, enable_blackwell_dsl replaces the NA modules entirely with DSL implementations rather than just setting a backend flag.

3. Runtime Auto-Detection

Each NeighborhoodAttention3D layer stores its backend in self.natten_backend. When this is None, the attention_function (defined around line 85 in attention.py) performs runtime probing:

  1. Call natten_available() to check for the natten library
  2. If missing, attempt to import Triton fallback
  3. If Triton unavailable, fall back to EAGER_SDPA
from ltx_core.model.video_vae.transformer.attention import NeighborhoodAttention3D

na_layer = NeighborhoodAttention3D(dim=256, kernel_size=(3,3,3))
print(f"Backend selected: {na_layer.natten_backend}")

# None → triggers auto-detection at first forward pass

Practical Backend Configuration Examples

Force Eager SDPA for Debugging

from ltx_core.model.video_vae.transformer.attention import NAttentionKind
from ltx_core.model.video_vae.transformer.apply import configure_natten_backend

configure_natten_backend(decoder, backend=NAttentionKind.EAGER_SDPA.value)

# All NA layers now use pure PyTorch—slower but fully inspectable

Verify Actual Runtime Backend

import torch
from ltx_core.model.video_vae.transformer.attention import NeighborhoodAttention3D

# Create layer and run dummy forward to trigger backend selection

na = NeighborhoodAttention3D(dim=512, kernel_size=(3, 3, 3), num_heads=8).cuda()
dummy = torch.randn(1, 512, 8, 8, 8).cuda()
_ = na(dummy)  # Probing happens here

print(f"Resolved backend: {na.natten_backend}")  # e.g., "natten", "triton", or "eager_sdpa"

Blackwell DSL Activation Checklist

from ltx_core.model.video_vae.transformer.config import DiffVAEMode
from ltx_core.model.video_vae.transformer.dsl_kernels.apply_dsl import enable_blackwell_dsl

# 1. Verify Blackwell architecture

assert torch.cuda.get_device_capability() == (10, 0)  # SM100 = Blackwell

# 2. Use the preset (includes DSL enablement)

config = DiffVAEMode.BLACKWELL_DSL.resolve()

# 3. Or manually enable on existing model

enable_blackwell_dsl(model)  # Swaps NA modules for DSL kernels

Key Source Files for Backend Implementation

File Responsibility
ltx_core/model/video_vae/transformer/config.py NAttentionKind enum, DiffVAEMode presets, DiffVAEConfig dataclass
ltx_core/model/video_vae/transformer/attention.py NeighborhoodAttention3D module, NattenAttention callable, runtime probing
ltx_core/model/video_vae/transformer/apply.py configure_natten_backend(), model-wide backend injection
ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py enable_blackwell_dsl(), Blackwell kernel swap logic
ltx_core/model/video_vae/transformer/fallback_na/__init__.py Triton (triton_na3d) and eager SDPA fallback registrations
ltx_core/model/video_vae/transformer/fallback_na/triton_na3d.py Triton JIT implementation of 3D neighborhood attention

Summary

  • NATTEN (NAttentionKind.NATTEN) is the fastest default backend, requiring the natten library installed.
  • Triton (NAttentionKind.TRITON) provides a JIT-compiled fallback when natten is unavailable, located in fallback_na/triton_na3d.py.
  • Eager SDPA (NAttentionKind.EAGER_SDPA) is the pure-PyTorch fallback that works on any device but trades speed for compatibility.
  • Blackwell DSL (NAttentionKind.BLACKWELL_DSL) is a specialized CuTe DSL kernel for Blackwell GPUs, enabled explicitly via DiffVAEMode.BLACKWELL_DSL or enable_blackwell_dsl().
  • Backend selection flows through DiffVAEMode.resolve() → configure_natten_backend() → runtime probing in NeighborhoodAttention3D.attention_function.

Frequently Asked Questions

How do I check which DiffVAE backend is currently active?

Inspect the natten_backend attribute on any NeighborhoodAttention3D layer after a forward pass. If None, the layer hasn't probed yet—run a dummy input through it to trigger auto-detection. The attention.py module logs the resolved backend at debug level.

Can I mix different backends in the same DiffVAE model?

Not through standard APIs. configure_natten_backend() applies one backend to the entire module subtree. For surgical backend assignment, you would need to manually set layer.natten_backend on individual NeighborhoodAttention3D instances before their first forward pass.

Why does Blackwell DSL require explicit opt-in rather than auto-detection?

The Blackwell DSL kernel (ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py) uses a fundamentally different code path that replaces NA modules entirely rather than just swapping a function pointer. This architectural change—plus the experimental nature of the DSL—makes automatic selection risky. Users must affirmatively choose DiffVAEMode.BLACKWELL_DSL.

What happens if I request NATTEN but the library is missing?

Runtime probing in attention.py catches the ImportError and automatically cascades to Triton, then Eager SDPA. No exception is raised; the model degrades gracefully. Watch logs for warnings about backend fallback to confirm this behavior during model initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →