# How to Configure the Video VAE Decoder in LTX-2: NADiffusion Decoder vs. Convolutional

> Configure LTX-2's video VAE decoder: choose NADiffusionDecoder for quality or convolutional for lower VRAM. Learn control methods like DiffVAEMode and CLI flags.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-14

---

**The LTX-2 video generation framework provides two VAE decoder options—the NADiffusion Decoder for maximum quality and a convolutional decoder for lower VRAM usage—controlled via checkpoint contents, the `DiffVAEMode` enum, and CLI flags.**

LTX-2, developed by Lightricks, ships with dual video VAE decoder architectures. Understanding how to configure the video VAE decoder in LTX-2 is essential for optimizing quality, inference speed, and memory consumption. This guide explains the decision logic in the source code and provides exact commands to switch between implementations.

## Decoder Options Overview

LTX-2 supports two decoder types:

| Decoder | Implementation | Selection Criteria |
|--------|----------------|-------------------|
| **NADiffusion Decoder** | `ltx_core.model.video_vae.diffusion_video_decoder.DiffusionVideoDecoder` | Checkpoint contains diffusion-decoder weights **and** `DiffVAEMode` is non-default |
| **Convolutional Decoder** | `ltx_core.model.video_vae` (classic convolutional) | Checkpoint lacks diffusion weights **or** `DiffVAEMode` left at default behavior |

The NADiffusion Decoder implements a "Minimal port of the reference `NADiffusionDecoder`" as noted at line 64 of [`diffusion_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/diffusion_video_decoder.py). This decoder generally produces higher-quality reconstructions at the cost of increased computation and memory usage.

## Three Factors That Determine Decoder Selection

The framework evaluates three conditions in order to decide which decoder to instantiate:

1. **Checkpoint contents** — If `video_vae_path` contains `diffusion_video_decoder` weights, the framework *can* use the NADiffusion Decoder.

2. **`DiffVAEMode` enum** — Defined in [`ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/config.py). Non-default values (`CHUNKED_EAGER`, `CHUNKED_COMPILE`, `COMBINED_COMPILE`, `BLACKWELL_DSL`) trigger diffusion decoder construction.

3. **CLI/API configuration** — Flags `--diffvae-optimization` (sets the enum) and `--with-video-vae-decoder` (enables/disables decoder loading entirely).

## How the Code Chooses the Decoder

The selection logic lives in [`ltx_trainer/model_loader.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/model_loader.py) within the `load_video_vae_decoder()` function (lines 194–226):

```python
def load_video_vae_decoder(video_vae_path, device, dtype, diffusion_vae=True, ...):
    if diffusion_vae:
        # Build NADiffusion Decoder with CUTLASS-FNA optimized operations

        decoder = build_cutlass_fna_diffusion_decoder_op(...)
        return DiffusionVideoDecoder(decoder, ...)
    else:
        # Fall back to convolutional decoder with default tiling

        return StaticConvolutionalDecoder(...)

```

During validation, `ValidationRunner._load_decoder_components` (lines 481–502 in [`ltx_trainer/validation_runner.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/validation_runner.py)) inspects the returned object:

```python
self._vae_decoder = load_video_vae_decoder(...)

if isinstance(self._vae_decoder, DiffusionVideoDecoder):
    # Add diffusion decoder weights to memory budget calculation

    memory_estimate += self._vae_decoder.estimated_memory_usage()
else:
    # Treat as convolutional decoder with standard memory profile

    pass

```

## CLI and Python API Configuration

### Use the NADiffusion Decoder (Highest Quality)

Requires a checkpoint with diffusion weights and a non-default `DiffVAEMode`:

```bash
python -m ltx_pipelines.ti2vid_two_stages \
    --model_path /path/to/checkpoint.safetensors \
    --diffvae-optimization CHUNKED_EAGER

```

Equivalent Python:

```python
from ltx_pipelines.ti2vid_two_stages import Ti2VidTwoStages
from ltx_core.model.video_vae.transformer.config import DiffVAEMode

pipeline = Ti2VidTwoStages(
    model_path="/path/to/checkpoint.safetensors",
    diffvae_optimization=DiffVAEMode.CHUNKED_EAGER,
)

```

### Force the Convolutional Decoder (Lower VRAM)

Either omit diffusion weights from the checkpoint path, or explicitly disable:

```bash
python -m ltx_pipelines.ti2vid_two_stages \
    --model_path /path/to/checkpoint.safetensors \
    --with-video-vae-decoder=True \
    --diffusion-vae=False

```

Direct Python control via `load_video_vae_decoder`:

```python
from ltx_core.model.video_vae.model_configurator import load_video_vae_decoder
import torch

decoder = load_video_vae_decoder(
    video_vae_path="/path/to/checkpoint.safetensors",
    device="cuda",
    dtype=torch.bfloat16,
    diffusion_vae=False,  # Forces convolutional decoder

)

```

### Skip Video VAE Decoder Entirely

For audio-only validation or v2a (video-to-audio) workflows:

```bash
python -m ltx_trainer.validate \
    --model_path /path/to/checkpoint.safetensors \
    --with-video-vae-decoder=False

```

Python:

```python
pipeline = Ti2VidTwoStages(
    model_path="/path/to/checkpoint.safetensors",
    with_video_vae_decoder=False,
)

```

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/config.py) | `DiffVAEMode` enum definition and defaults |
| [`ltx_core/model/video_vae/diffusion_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/diffusion_video_decoder.py) | NADiffusion Decoder implementation |
| [`ltx_trainer/model_loader.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/model_loader.py) | `load_video_vae_decoder()` selection logic (lines 194–226) |
| [`ltx_trainer/validation_runner.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/validation_runner.py) | Runtime decoder type detection (lines 481–502) |
| [`ltx_trainer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/config.py) | CLI argument definitions: `--diffvae-optimization` (line 405), `--with-video-vae-decoder` (line 568) |
| [`ltx_pipelines/ti2vid_two_stages_mgpu.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/ti2vid_two_stages_mgpu.py) | Example pipeline propagating `diffvae_optimization` (line 66) |

## Performance and Quality Trade-offs

When you configure the video VAE decoder in LTX-2, consider these characteristics:

- **NADiffusion Decoder**: Superior reconstruction fidelity, especially for fine details and temporal consistency. Requires CUTLASS-FNA kernel support. Higher VRAM footprint and longer inference time.

- **Convolutional Decoder**: Faster tile-based decoding, lower memory requirements. Suitable for rapid prototyping, batch processing, or hardware-constrained deployments.

The `DiffVAEMode` values beyond `CHUNKED_EAGER` enable additional optimizations: `CHUNKED_COMPILE` and `COMBINED_COMPILE` use torch.compile for kernel fusion, while `BLACKWELL_DSL` targets NVIDIA Blackwell architecture-specific optimizations.

## Summary

- **NADiffusion Decoder** activates automatically with diffusion-enabled checkpoints and non-default `DiffVAEMode` settings
- **Convolutional Decoder** serves as the fallback when diffusion weights are absent or explicitly disabled
- Control via `--diffvae-optimization` (enum selection) and `--with-video-vae-decoder` (decoder loading toggle)
- Direct Python API access through `load_video_vae_decoder()` with `diffusion_vae=True/False`
- Core logic spans `ltx_core` model definitions and `ltx_trainer` loading infrastructure

## Frequently Asked Questions

### What happens if I set `--diffvae-optimization` but my checkpoint lacks diffusion weights?

The framework attempts to load diffusion weights and raises an error or warning depending on the pipeline. To avoid this, either use a checkpoint with diffusion decoder weights, or explicitly pass `diffusion_vae=False` to `load_video_vae_decoder()`.

### Can I switch decoders without changing checkpoints?

Not directly for the NADiffusion Decoder—it requires the corresponding weights in the checkpoint. However, you can force the convolutional decoder on any checkpoint by setting `diffusion_vae=False`. The convolutional decoder weights are standard VAE components present in all LTX-2 checkpoints.

### Does `--diffvae-optimization BLACKWELL_DSL` require specific hardware?

Yes. The `BLACKWELL_DSL` mode targets NVIDIA Blackwell architecture GPUs and requires compatible CUDA/cuDNN versions with CUTLASS-FNA support. For older GPUs, use `CHUNKED_EAGER` or `CHUNKED_COMPILE` instead.

### How do I verify which decoder is actually loaded at runtime?

Inspect `ValidationRunner._vae_decoder` or check the type after calling `load_video_vae_decoder()`:

```python
from ltx_core.model.video_vae.diffusion_video_decoder import DiffusionVideoDecoder

decoder = load_video_vae_decoder(...)
print("Using NADiffusion Decoder:", isinstance(decoder, DiffusionVideoDecoder))

```