# DiffVAE Backend Options in LTX-2: NATTEN, Triton, Eager SDPA & Blackwell DSL Explained

> Explore DiffVAE backend options in LTX-2: NATTEN, Triton, Eager SDPA, and Blackwell DSL. Understand which to use for your hardware and configuration.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-20

---

**LTX-2 supports four DiffVAE backends: NATTEN (CUDA kernel library), Triton (JIT-compiled fallback), Eager SDPA (pure PyTorch fallback), and Blackwell DSL (CuTe DSL kernel for Blackwell GPUs), selectable via the `NAttentionKind` enum.** According to the Lightricks/LTX-2 source code, the DiffVAE architecture uses a pluggable neighborhood-attention system that automatically falls back through these options based on hardware availability and user configuration.

The video-latent diffusion VAE in LTX-2 relies heavily on **3D neighborhood attention** for efficient spatiotemporal processing. The backend selection mechanism—implemented across [`config.py`](https://github.com/Lightricks/LTX-2/blob/main/config.py), [`attention.py`](https://github.com/Lightricks/LTX-2/blob/main/attention.py), and [`apply.py`](https://github.com/Lightricks/LTX-2/blob/main/apply.py)—lets users prioritize speed (NATTEN), compatibility (Triton/SDPA), or cutting-edge hardware acceleration (Blackwell DSL). Each backend targets different deployment scenarios, from consumer GPUs to datacenter Blackwell clusters.

## The Four DiffVAE Backend Options

LTX-2 defines all backend choices in [`ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/config.py) through the `NAttentionKind` enum. Here's how each option works:

### NATTEN: The Primary CUDA Kernel Backend

`NAttentionKind.NATTEN` is the **default high-performance path** when the `natten` Python package is installed. This backend calls into optimized CUDA kernels, with optional kernel-level selection (e.g., `"cutlass-fna"`) via the `backend` argument of `NattenAttention`.

- **Source location**: [`ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/config.py), lines 15-21
- **Best for**: Production inference on CUDA GPUs with `natten` wheels available
- **Performance**: Fastest general-purpose option; kernel auto-tuned for spatial and temporal dimensions

The NATTEN backend is what most users will run in practice, provided they install the underlying `natten` dependency.

### Triton: The JIT-Compiled Fallback

`NAttentionKind.TRITON` activates when `natten` is unavailable but Triton is present. LTX-2's implementation in [`fallback_na/triton_na3d.py`](https://github.com/Lightricks/LTX-2/blob/main/fallback_na/triton_na3d.py) exposes the **same Python API as NATTEN**, just compiled through the Triton JIT compiler at runtime.

- **Source location**: [`ltx_core/model/video_vae/transformer/fallback_na/__init__.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/fallback_na/__init__.py), line 77
- **Best for**: Development environments, custom CUDA versions, or platforms without prebuilt `natten` binaries
- **Performance**: Moderate—slower than NATTEN but faster than eager PyTorch

This fallback ensures the DiffVAE model remains functional across diverse hardware without requiring recompilation.

### Eager SDPA: The Device-Agnostic Fallback

`NAttentionKind.EAGER_SDPA` is the **last-resort fallback** using PyTorch's native `scaled_dot_product_attention`. It requires no external dependencies beyond PyTorch itself.

- **Source location**: [`ltx_core/model/video_vae/transformer/fallback_na/__init__.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/fallback_na/__init__.py), lines 80-81
- **Best for**: CPU inference, debugging, or environments where both NATTEN and Triton fail
- **Performance**: Slowest and most memory-intensive; pure Python/Torch implementation

This backend guarantees the DiffVAE model will run anywhere PyTorch does, though latencies may be unacceptable for real-time video generation.

### Blackwell DSL: The Specialized Datacenter Kernel

`NAttentionKind.BLACKWELL_DSL` is an **experimental, hardware-specific backend** for NVIDIA Blackwell (B100/B200) GPUs. It compiles a fused CuTe DSL kernel that handles the complete neighborhood-attention block, including a deferred stage-5 computation unique to the DiffVAE architecture.

- **Source location**: [`ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py), line 25
- **Best for**: Large-scale inference on Blackwell hardware
- **Performance**: Potentially highest throughput on target hardware; requires explicit opt-in via `DiffVAEMode.BLACKWELL_DSL`

This backend is **not selected automatically**—users must request it through the preset system.

## How LTX-2 Selects Your DiffVAE Backend

The backend selection pipeline involves three stages, each implemented in specific source files:

### 1. User Preset Resolution (`DiffVAEMode`)

Users start with high-level presets defined in [`config.py`](https://github.com/Lightricks/LTX-2/blob/main/config.py). Calling `DiffVAEMode.resolve()` (lines 70-101) expands these into concrete `DiffVAEConfig` objects:

| Preset | Backend Behavior |
|--------|----------------|
| `CHUNKED_EAGER` | Chunked processing, no compilation, NATTEN if available |
| `CHUNKED_COMPILE` | Chunked processing with `torch.compile`, NATTEN preferred |
| `COMBINED_COMPILE` | Combined spatial-temporal compilation, NATTEN preferred |
| `BLACKWELL_DSL` | Forces Blackwell DSL kernel, ignores NATTEN availability |

```python
from ltx_core.model.video_vae.transformer.config import DiffVAEMode

# Select Blackwell DSL path explicitly

vae_cfg = DiffVAEMode.BLACKWELL_DSL.resolve()
print(vae_cfg)

# DiffVAEConfig(

#     block=DiffVAEBlockKind.BLACKWELL_DSL,

#     w_chunks=1,

#     natten_backend=None,

#     attention=NAttentionKind.BLACKWELL_DSL,

#     compile_blocks=True,

#     compile_det_stages=True

# )

```

### 2. Backend Injection (`configure_natten_backend`)

The function `configure_natten_backend(module_root, backend)` in [`ltx_core/model/video_vae/transformer/apply.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/apply.py) (line 108) walks the module tree and **writes the selected backend into each `NeighborhoodAttention3D` instance**:

```python
from ltx_core.model.video_vae.transformer.attention import NAttentionKind
from ltx_core.model.video_vae.transformer.apply import configure_natten_backend

# Force Triton fallback across all NA layers

configure_natten_backend(decoder, backend=NAttentionKind.TRITON.value)

```

For `BLACKWELL_DSL` requests, `enable_blackwell_dsl` **replaces the NA modules entirely** with DSL implementations rather than just setting a backend flag.

### 3. Runtime Auto-Detection

Each `NeighborhoodAttention3D` layer stores its backend in `self.natten_backend`. When this is `None`, the `attention_function` (defined around line 85 in [`attention.py`](https://github.com/Lightricks/LTX-2/blob/main/attention.py)) performs **runtime probing**:

1. Call `natten_available()` to check for the `natten` library
2. If missing, attempt to import Triton fallback
3. If Triton unavailable, fall back to `EAGER_SDPA`

```python
from ltx_core.model.video_vae.transformer.attention import NeighborhoodAttention3D

na_layer = NeighborhoodAttention3D(dim=256, kernel_size=(3,3,3))
print(f"Backend selected: {na_layer.natten_backend}")

# None → triggers auto-detection at first forward pass

```

## Practical Backend Configuration Examples

### Force Eager SDPA for Debugging

```python
from ltx_core.model.video_vae.transformer.attention import NAttentionKind
from ltx_core.model.video_vae.transformer.apply import configure_natten_backend

configure_natten_backend(decoder, backend=NAttentionKind.EAGER_SDPA.value)

# All NA layers now use pure PyTorch—slower but fully inspectable

```

### Verify Actual Runtime Backend

```python
import torch
from ltx_core.model.video_vae.transformer.attention import NeighborhoodAttention3D

# Create layer and run dummy forward to trigger backend selection

na = NeighborhoodAttention3D(dim=512, kernel_size=(3, 3, 3), num_heads=8).cuda()
dummy = torch.randn(1, 512, 8, 8, 8).cuda()
_ = na(dummy)  # Probing happens here

print(f"Resolved backend: {na.natten_backend}")  # e.g., "natten", "triton", or "eager_sdpa"

```

### Blackwell DSL Activation Checklist

```python
from ltx_core.model.video_vae.transformer.config import DiffVAEMode
from ltx_core.model.video_vae.transformer.dsl_kernels.apply_dsl import enable_blackwell_dsl

# 1. Verify Blackwell architecture

assert torch.cuda.get_device_capability() == (10, 0)  # SM100 = Blackwell

# 2. Use the preset (includes DSL enablement)

config = DiffVAEMode.BLACKWELL_DSL.resolve()

# 3. Or manually enable on existing model

enable_blackwell_dsl(model)  # Swaps NA modules for DSL kernels

```

## Key Source Files for Backend Implementation

| File | Responsibility |
|------|--------------|
| [`ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/config.py) | `NAttentionKind` enum, `DiffVAEMode` presets, `DiffVAEConfig` dataclass |
| [`ltx_core/model/video_vae/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/attention.py) | `NeighborhoodAttention3D` module, `NattenAttention` callable, runtime probing |
| [`ltx_core/model/video_vae/transformer/apply.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/apply.py) | `configure_natten_backend()`, model-wide backend injection |
| [`ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py) | `enable_blackwell_dsl()`, Blackwell kernel swap logic |
| [`ltx_core/model/video_vae/transformer/fallback_na/__init__.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/fallback_na/__init__.py) | Triton (`triton_na3d`) and eager SDPA fallback registrations |
| [`ltx_core/model/video_vae/transformer/fallback_na/triton_na3d.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/fallback_na/triton_na3d.py) | Triton JIT implementation of 3D neighborhood attention |

## Summary

- **NATTEN** (`NAttentionKind.NATTEN`) is the fastest default backend, requiring the `natten` library installed.
- **Triton** (`NAttentionKind.TRITON`) provides a JIT-compiled fallback when `natten` is unavailable, located in [`fallback_na/triton_na3d.py`](https://github.com/Lightricks/LTX-2/blob/main/fallback_na/triton_na3d.py).
- **Eager SDPA** (`NAttentionKind.EAGER_SDPA`) is the pure-PyTorch fallback that works on any device but trades speed for compatibility.
- **Blackwell DSL** (`NAttentionKind.BLACKWELL_DSL`) is a specialized CuTe DSL kernel for Blackwell GPUs, enabled explicitly via `DiffVAEMode.BLACKWELL_DSL` or `enable_blackwell_dsl()`.
- Backend selection flows through `DiffVAEMode.resolve()` → `configure_natten_backend()` → runtime probing in `NeighborhoodAttention3D.attention_function`.

## Frequently Asked Questions

### How do I check which DiffVAE backend is currently active?

Inspect the `natten_backend` attribute on any `NeighborhoodAttention3D` layer after a forward pass. If `None`, the layer hasn't probed yet—run a dummy input through it to trigger auto-detection. The [`attention.py`](https://github.com/Lightricks/LTX-2/blob/main/attention.py) module logs the resolved backend at debug level.

### Can I mix different backends in the same DiffVAE model?

Not through standard APIs. `configure_natten_backend()` applies one backend to the entire module subtree. For surgical backend assignment, you would need to manually set `layer.natten_backend` on individual `NeighborhoodAttention3D` instances before their first forward pass.

### Why does Blackwell DSL require explicit opt-in rather than auto-detection?

The Blackwell DSL kernel ([`ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/dsl_kernels/apply_dsl.py)) uses a fundamentally different code path that **replaces NA modules entirely** rather than just swapping a function pointer. This architectural change—plus the experimental nature of the DSL—makes automatic selection risky. Users must affirmatively choose `DiffVAEMode.BLACKWELL_DSL`.

### What happens if I request NATTEN but the library is missing?

Runtime probing in [`attention.py`](https://github.com/Lightricks/LTX-2/blob/main/attention.py) catches the `ImportError` and automatically cascades to Triton, then Eager SDPA. No exception is raised; the model degrades gracefully. Watch logs for warnings about backend fallback to confirm this behavior during model initialization.