# DiffVAE Modes in LTX-2: CHUNKED_EAGER, NADiffusionDecoder, and ConvDecoder Explained

> Explore LTX-2's DiffVAE modes: CHUNKED_EAGER, NADiffusionDecoder, and ConvDecoder. Understand their trade-offs in compilation, speed, and VRAM for optimal video generation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-18

---

**LTX-2 offers four DiffVAE decode presets—CHUNKED_EAGER, CHUNKED_COMPILE, COMBINED_COMPILE, and BLACKWELL_DSL—that trade off compilation time, inference speed, and VRAM usage for video generation.**

LTX-2's **DiffVAE modes** control how the video VAE decodes latent representations back to pixel space. Each mode configures different attention backends, compilation strategies, and memory tiling patterns. Understanding these modes is essential for optimizing inference on different GPU hardware, from consumer cards to datacenter Blackwell systems.

---

## What Are DiffVAE Modes in LTX-2?

DiffVAE modes are **preset configurations** defined in the `DiffVAEMode` enum at [`packages/ltx-core/src/ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae/transformer/config.py). When you load a checkpoint, you can select a mode that determines:

- Whether to use **chunked attention** (width-split) or **full-volume attention**
- Whether to **JIT-compile** diffusion blocks and deterministic stages
- Which **attention backend** (natten, Triton/Eager, or Blackwell DSL) to invoke

The `resolve()` method in `DiffVAEMode` expands each preset into a concrete `DiffVAEConfig` with specific values for `w_chunks`, `compile_blocks`, `compile_det_stages`, and `attention`.

---

## The Four DiffVAE Decode Presets

### CHUNKED_EAGER: Minimal VRAM, No Compilation

**`CHUNKED_EAGER`** uses a width-chunked diffusion block (`w_chunks=4`) with **no JIT compilation**. It defers stage-4 context computation and runs entirely in eager mode or Triton fallback.

- **Compile time**: None (instant startup)
- **Warm runtime**: Slowest (~2–2.5× slower than `COMBINED_COMPILE`)
- **Peak VRAM**: ~50% of `COMBINED_COMPILE`
- **Hardware**: Works on any GPU; falls back to Triton/Eager if **natten** is unavailable

Use `CHUNKED_EAGER` when GPU memory is severely constrained (≤12 GB) or when you need immediate startup without compilation overhead.

### CHUNKED_COMPILE: Balanced Speed and Memory

**`CHUNKED_COMPILE`** keeps the same chunked layout as `CHUNKED_EAGER` but **compiles the non-deterministic blocks** (`compile_blocks=True`). Deterministic stages remain uncompiled to reduce compile time.

- **Compile time**: ~2× faster than `COMBINED_COMPILE`
- **Warm runtime**: ~1.4× slower than `COMBINED_COMPILE`
- **Peak VRAM**: ~50% of `COMBINED_COMPILE`
- **Hardware**: Requires **natten** backend or falls back to Triton/Eager

Use `CHUNKED_COMPILE` with ~16 GB VRAM when you can afford a one-time compilation to gain moderate speed improvements without doubling memory usage.

### COMBINED_COMPILE: Maximum Throughput

**`COMBINED_COMPILE`** uses a **combined diffusion block** (`w_chunks=1`) with full-volume attention and **compiles both blocks and deterministic stages**.

- **Compile time**: Longest (full compilation of all stages)
- **Warm runtime**: Fastest among DiffVAE presets
- **Peak VRAM**: Highest (~2× `CHUNKED_*` modes)
- **Hardware**: **Requires** natten backend

Use `COMBINED_COMPILE` with ≥24 GB VRAM for large-scale batch inference where maximum throughput outweighs compilation and memory costs.

### BLACKWELL_DSL: Datacenter-Grade Performance

**`BLACKWELL_DSL`** targets NVIDIA Blackwell GPUs with a proprietary **CuTe DSL kernel** that fuses stage-5 operations and performs in-kernel upsampling plus context projection.

- **Compile time**: Comparable to `COMBINED_COMPILE`
- **Warm runtime**: Same as `COMBINED_COMPILE` (fast)
- **Peak VRAM**: Same as `COMBINED_COMPILE`
- **Hardware**: Requires **Blackwell-DSL** backend installation

Use `BLACKWELL_DSL` only in datacenter environments with the Blackwell CuTe runtime installed.

---

## The Convolutional Decoder Alternative

LTX-2 also provides **`ConvVideoDecoder`** at [`packages/ltx-core/src/ltx_core/model/video_vae/conv_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae/conv_video_decoder.py), a separate architecture entirely.

When a checkpoint's metadata specifies `decoder_kind="conv"` (as noted in [`packages/ltx-trainer/src/ltx_trainer/validation_runner.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/validation_runner.py) at line 81), DiffVAE presets are **ignored** and the conv decoder runs instead. The conv decoder executes **single-pass 3D convolutions**—fast and memory-efficient with no tiling configuration required.

**NADiffusionDecoder** (referenced in the question title) appears to be an internal alias or documentation reference to the **natten-based diffusion decoder**, which corresponds to the `COMBINED_COMPILE` and `BLACKWELL_DSL` modes that require the natten backend. The chunked modes use alternative attention implementations when natten is unavailable.

---

## When to Use Each DiffVAE Mode

| Hardware Scenario | Recommended Mode | Rationale |
|-------------------|------------------|-----------|
| ≤12 GB VRAM, consumer GPU | `CHUNKED_EAGER` | Minimal memory, no compilation |
| ~16 GB VRAM, moderate batch | `CHUNKED_COMPILE` | Compile once, gain 40% speed vs. eager |
| ≥24 GB VRAM, production inference | `COMBINED_COMPILE` | Maximum throughput with natten |
| Blackwell datacenter GPU | `BLACKWELL_DSL` | Fused kernels for large-scale serving |
| Checkpoint has `decoder_kind="conv"` | `ConvVideoDecoder` | Fastest single-pass, ignore DiffVAE presets |

---

## Code Examples

### Selecting a DiffVAE Mode in Pipeline Construction

```python
from ltx_core.model.video_vae.transformer import DiffVAEMode
from ltx_pipelines.utils.blocks import DiffusionStage

stage = DiffusionStage.from_checkpoint(
    checkpoint_path="checkpoints/ltx2_19b.safetensors",
    diffvae_optimization=DiffVAEMode.COMBINED_COMPILE,
)

```

### Runtime Mode Override with Tiled Decoding

```python
from ltx_core.model.video_vae.transformer import DiffVAEMode
from ltx_core.model.video_vae.video_vae import recommended_decode_tiling_config

tiling_cfg = recommended_decode_tiling_config(
    height=latent.shape[-2] * 32,
    width=latent.shape[-1] * 32,
    num_frames=(latent.shape[2] - 1) * 8 + 1,
    mode=DiffVAEMode.CHUNKED_COMPILE,
    free_bytes=10 * 1024**3,  # 10 GB free budget

)
chunks = list(decoder.tiled_decode(latent, tiling_config=tiling_cfg))

```

### Inspecting Resolved Configuration

```python
from ltx_core.model.video_vae.transformer import DiffVAEMode

cfg = DiffVAEMode.CHUNKED_COMPILE.resolve()
print(cfg)

# DiffVAEConfig(

#   block=DiffVAEBlockKind.CHUNKED,

#   w_chunks=4,

#   natten_backend=None,

#   attention=NAttentionKind.NATTEN,

#   compile_blocks=True,

#   compile_det_stages=False,

# )

```

### Using the Convolutional Decoder

```python
from ltx_core.model.video_vae.video_vae import ConvVideoDecoder

decoder = ConvVideoDecoder.from_checkpoint(
    checkpoint_path="checkpoints/ltx2_19b.safetensors"
)
decoded = decoder(latent_tensor)  # No tiling config needed

```

---

## Key Source Files

| File Path | Purpose |
|-----------|---------|
| [`packages/ltx-core/src/ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae/transformer/config.py) | Defines `DiffVAEMode`, `DiffVAEConfig`, and `resolve()` logic |
| [`packages/ltx-core/src/ltx_core/model/video_vae/diffusion_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae/diffusion_video_decoder.py) | Implements diffusion-based decoding with mode-aware tiling |
| [`packages/ltx-core/src/ltx_core/model/video_vae/conv_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae/conv_video_decoder.py) | Single-pass convolutional decoder |
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) | `DiffusionStage` wrapper accepting `DiffVAEMode` |
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) | `recommended_decode_tiling_config` helper |

---

## Summary

- **DiffVAE modes** select between chunked and full-volume attention, compilation strategies, and hardware backends
- **CHUNKED_EAGER** sacrifices speed for minimal VRAM and zero compilation
- **CHUNKED_COMPILE** offers middle-ground performance with moderate compile cost
- **COMBINED_COMPILE** delivers maximum throughput on high-VRAM systems with natten
- **BLACKWELL_DSL** optimizes for datacenter Blackwell deployments
- **ConvVideoDecoder** bypasses DiffVAE entirely when checkpoints specify convolutional decoding
- The mode flows from `DiffVAEMode` → `DiffVAEConfig` → tiling configuration → decoder implementation

---

## Frequently Asked Questions

### What happens if natten is not installed?

LTX-2 falls back to Triton/Eager attention backends. Chunked modes (`CHUNKED_EAGER`, `CHUNKED_COMPILE`) degrade gracefully, but `COMBINED_COMPILE` and `BLACKWELL_DSL` will fail or fall back to slower implementations depending on your error handling.

### Can I switch DiffVAE modes after loading a checkpoint?

Yes. The mode is passed to `recommended_decode_tiling_config()` at decode time, not cached at checkpoint load. Override `mode` in your tiling configuration call to change behavior without reloading weights.

### Why does LTX-2 have both diffusion and convolutional decoders?

The **diffusion decoder** offers higher reconstruction quality through iterative refinement, configurable via DiffVAE modes for speed/quality tradeoffs. The **convolutional decoder** provides faster, deterministic single-pass decoding when quality requirements allow it. Checkpoint metadata selects which architecture to instantiate.