DiffVAE Backend Options in LTX-2: NATTEN, Triton, Eager SDPA & Blackwell DSL Explained
LTX-2 supports four DiffVAE backends: NATTEN (CUDA kernel library), Triton (JIT-compiled fallback), Eager SDPA (pure PyTorch fallback), and Blackwell DSL (CuTe DSL kernel for Blackwell GPUs), selectable via the NAttentionKind enum. According to the Lightricks/LTX-2 source code, the DiffVAE architecture uses a pluggable neighborhood-attention system that automatically falls back through these options based on hardware availability and user configuration.
The video-latent diffusion VAE in LTX-2 relies heavily on 3D neighborhood attention for efficient spatiotemporal processing. The backend selection mechanism—implemented across config.py, attention.py, and apply.py—lets users prioritize speed (NATTEN), compatibility (Triton/SDPA), or cutting-edge hardware acceleration (Blackwell DSL). Each backend targets different deployment scenarios, from consumer GPUs to datacenter Blackwell clusters.
The Four DiffVAE Backend Options
LTX-2 defines all backend choices in ltx_core/model/video_vae/transformer/config.py through the NAttentionKind enum. Here's how each option works:
NATTEN: The Primary CUDA Kernel Backend
NAttentionKind.NATTEN is the default high-performance path when the natten Python package is installed. This backend calls into optimized CUDA kernels, with optional kernel-level selection (e.g., "cutlass-fna") via the backend argument of NattenAttention.
- Source location:
ltx_core/model/video_vae/transformer/config.py, lines 15-21 - Best for: Production inference on CUDA GPUs with
nattenwheels available - Performance: Fastest general-purpose option; kernel auto-tuned for spatial and temporal dimensions
The NATTEN backend is what most users will run in practice, provided they install the underlying natten dependency.
Triton: The JIT-Compiled Fallback
NAttentionKind.TRITON activates when natten is unavailable but Triton is present. LTX-2's implementation in fallback_na/triton_na3d.py exposes the same Python API as NATTEN, just compiled through the Triton JIT compiler at runtime.
- Source location:
ltx_core/model/video_vae/transformer/fallback_na/__init__.py, line 77 - Best for: Development environments, custom CUDA versions, or platforms without prebuilt
nattenbinaries - Performance: Moderate—slower than NATTEN but faster than eager PyTorch
This fallback ensures the DiffVAE model remains functional across diverse hardware without requiring recompilation.
Eager SDPA: The Device-Agnostic Fallback
NAttentionKind.EAGER_SDPA is the last-resort fallback using PyTorch's native scaled_dot_product_attention. It requires no external dependencies beyond PyTorch itself.
- Source location:
ltx_core/model/video_vae/transformer/fallback_na/__init__.py, lines 80-81 - Best for: CPU inference, debugging, or environments where both NATTEN and Triton fail
- Performance: Slowest and most memory-intensive; pure Python/Torch implementation
This backend guarantees the DiffVAE model will run anywhere PyTorch does, though latencies may be unacceptable for real-time video generation.
Blackwell DSL: The Specialized Datacenter Kernel
NAttentionKind.BLACKWELL_DSL is an experimental, hardware-specific backend for NVIDIA Blackwell (B100/B200) GPUs. It compiles a fused CuTe DSL kernel that handles the complete neighborhood-attention block, including a deferred stage-5 computation unique to the DiffVAE architecture.
- Source location:
ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py, line 25 - Best for: Large-scale inference on Blackwell hardware
- Performance: Potentially highest throughput on target hardware; requires explicit opt-in via
DiffVAEMode.BLACKWELL_DSL
This backend is not selected automatically—users must request it through the preset system.
How LTX-2 Selects Your DiffVAE Backend
The backend selection pipeline involves three stages, each implemented in specific source files:
1. User Preset Resolution (DiffVAEMode)
Users start with high-level presets defined in config.py. Calling DiffVAEMode.resolve() (lines 70-101) expands these into concrete DiffVAEConfig objects:
| Preset | Backend Behavior |
|---|---|
CHUNKED_EAGER |
Chunked processing, no compilation, NATTEN if available |
CHUNKED_COMPILE |
Chunked processing with torch.compile, NATTEN preferred |
COMBINED_COMPILE |
Combined spatial-temporal compilation, NATTEN preferred |
BLACKWELL_DSL |
Forces Blackwell DSL kernel, ignores NATTEN availability |
from ltx_core.model.video_vae.transformer.config import DiffVAEMode
# Select Blackwell DSL path explicitly
vae_cfg = DiffVAEMode.BLACKWELL_DSL.resolve()
print(vae_cfg)
# DiffVAEConfig(
# block=DiffVAEBlockKind.BLACKWELL_DSL,
# w_chunks=1,
# natten_backend=None,
# attention=NAttentionKind.BLACKWELL_DSL,
# compile_blocks=True,
# compile_det_stages=True
# )
2. Backend Injection (configure_natten_backend)
The function configure_natten_backend(module_root, backend) in ltx_core/model/video_vae/transformer/apply.py (line 108) walks the module tree and writes the selected backend into each NeighborhoodAttention3D instance:
from ltx_core.model.video_vae.transformer.attention import NAttentionKind
from ltx_core.model.video_vae.transformer.apply import configure_natten_backend
# Force Triton fallback across all NA layers
configure_natten_backend(decoder, backend=NAttentionKind.TRITON.value)
For BLACKWELL_DSL requests, enable_blackwell_dsl replaces the NA modules entirely with DSL implementations rather than just setting a backend flag.
3. Runtime Auto-Detection
Each NeighborhoodAttention3D layer stores its backend in self.natten_backend. When this is None, the attention_function (defined around line 85 in attention.py) performs runtime probing:
- Call
natten_available()to check for thenattenlibrary - If missing, attempt to import Triton fallback
- If Triton unavailable, fall back to
EAGER_SDPA
from ltx_core.model.video_vae.transformer.attention import NeighborhoodAttention3D
na_layer = NeighborhoodAttention3D(dim=256, kernel_size=(3,3,3))
print(f"Backend selected: {na_layer.natten_backend}")
# None → triggers auto-detection at first forward pass
Practical Backend Configuration Examples
Force Eager SDPA for Debugging
from ltx_core.model.video_vae.transformer.attention import NAttentionKind
from ltx_core.model.video_vae.transformer.apply import configure_natten_backend
configure_natten_backend(decoder, backend=NAttentionKind.EAGER_SDPA.value)
# All NA layers now use pure PyTorch—slower but fully inspectable
Verify Actual Runtime Backend
import torch
from ltx_core.model.video_vae.transformer.attention import NeighborhoodAttention3D
# Create layer and run dummy forward to trigger backend selection
na = NeighborhoodAttention3D(dim=512, kernel_size=(3, 3, 3), num_heads=8).cuda()
dummy = torch.randn(1, 512, 8, 8, 8).cuda()
_ = na(dummy) # Probing happens here
print(f"Resolved backend: {na.natten_backend}") # e.g., "natten", "triton", or "eager_sdpa"
Blackwell DSL Activation Checklist
from ltx_core.model.video_vae.transformer.config import DiffVAEMode
from ltx_core.model.video_vae.transformer.dsl_kernels.apply_dsl import enable_blackwell_dsl
# 1. Verify Blackwell architecture
assert torch.cuda.get_device_capability() == (10, 0) # SM100 = Blackwell
# 2. Use the preset (includes DSL enablement)
config = DiffVAEMode.BLACKWELL_DSL.resolve()
# 3. Or manually enable on existing model
enable_blackwell_dsl(model) # Swaps NA modules for DSL kernels
Key Source Files for Backend Implementation
| File | Responsibility |
|---|---|
ltx_core/model/video_vae/transformer/config.py |
NAttentionKind enum, DiffVAEMode presets, DiffVAEConfig dataclass |
ltx_core/model/video_vae/transformer/attention.py |
NeighborhoodAttention3D module, NattenAttention callable, runtime probing |
ltx_core/model/video_vae/transformer/apply.py |
configure_natten_backend(), model-wide backend injection |
ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py |
enable_blackwell_dsl(), Blackwell kernel swap logic |
ltx_core/model/video_vae/transformer/fallback_na/__init__.py |
Triton (triton_na3d) and eager SDPA fallback registrations |
ltx_core/model/video_vae/transformer/fallback_na/triton_na3d.py |
Triton JIT implementation of 3D neighborhood attention |
Summary
- NATTEN (
NAttentionKind.NATTEN) is the fastest default backend, requiring thenattenlibrary installed. - Triton (
NAttentionKind.TRITON) provides a JIT-compiled fallback whennattenis unavailable, located infallback_na/triton_na3d.py. - Eager SDPA (
NAttentionKind.EAGER_SDPA) is the pure-PyTorch fallback that works on any device but trades speed for compatibility. - Blackwell DSL (
NAttentionKind.BLACKWELL_DSL) is a specialized CuTe DSL kernel for Blackwell GPUs, enabled explicitly viaDiffVAEMode.BLACKWELL_DSLorenable_blackwell_dsl(). - Backend selection flows through
DiffVAEMode.resolve()→configure_natten_backend()→ runtime probing inNeighborhoodAttention3D.attention_function.
Frequently Asked Questions
How do I check which DiffVAE backend is currently active?
Inspect the natten_backend attribute on any NeighborhoodAttention3D layer after a forward pass. If None, the layer hasn't probed yet—run a dummy input through it to trigger auto-detection. The attention.py module logs the resolved backend at debug level.
Can I mix different backends in the same DiffVAE model?
Not through standard APIs. configure_natten_backend() applies one backend to the entire module subtree. For surgical backend assignment, you would need to manually set layer.natten_backend on individual NeighborhoodAttention3D instances before their first forward pass.
Why does Blackwell DSL require explicit opt-in rather than auto-detection?
The Blackwell DSL kernel (ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py) uses a fundamentally different code path that replaces NA modules entirely rather than just swapping a function pointer. This architectural change—plus the experimental nature of the DSL—makes automatic selection risky. Users must affirmatively choose DiffVAEMode.BLACKWELL_DSL.
What happens if I request NATTEN but the library is missing?
Runtime probing in attention.py catches the ImportError and automatically cascades to Triton, then Eager SDPA. No exception is raised; the model degrades gracefully. Watch logs for warnings about backend fallback to confirm this behavior during model initialization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →