How to Configure the Video VAE Decoder in LTX-2: NADiffusion Decoder vs. Convolutional

The LTX-2 video generation framework provides two VAE decoder options—the NADiffusion Decoder for maximum quality and a convolutional decoder for lower VRAM usage—controlled via checkpoint contents, the DiffVAEMode enum, and CLI flags.

LTX-2, developed by Lightricks, ships with dual video VAE decoder architectures. Understanding how to configure the video VAE decoder in LTX-2 is essential for optimizing quality, inference speed, and memory consumption. This guide explains the decision logic in the source code and provides exact commands to switch between implementations.

Decoder Options Overview

LTX-2 supports two decoder types:

Decoder Implementation Selection Criteria
NADiffusion Decoder ltx_core.model.video_vae.diffusion_video_decoder.DiffusionVideoDecoder Checkpoint contains diffusion-decoder weights and DiffVAEMode is non-default
Convolutional Decoder ltx_core.model.video_vae (classic convolutional) Checkpoint lacks diffusion weights or DiffVAEMode left at default behavior

The NADiffusion Decoder implements a "Minimal port of the reference NADiffusionDecoder" as noted at line 64 of diffusion_video_decoder.py. This decoder generally produces higher-quality reconstructions at the cost of increased computation and memory usage.

Three Factors That Determine Decoder Selection

The framework evaluates three conditions in order to decide which decoder to instantiate:

  1. Checkpoint contents — If video_vae_path contains diffusion_video_decoder weights, the framework can use the NADiffusion Decoder.

  2. DiffVAEMode enum — Defined in ltx_core/model/video_vae/transformer/config.py. Non-default values (CHUNKED_EAGER, CHUNKED_COMPILE, COMBINED_COMPILE, BLACKWELL_DSL) trigger diffusion decoder construction.

  3. CLI/API configuration — Flags --diffvae-optimization (sets the enum) and --with-video-vae-decoder (enables/disables decoder loading entirely).

How the Code Chooses the Decoder

The selection logic lives in ltx_trainer/model_loader.py within the load_video_vae_decoder() function (lines 194–226):

def load_video_vae_decoder(video_vae_path, device, dtype, diffusion_vae=True, ...):
    if diffusion_vae:
        # Build NADiffusion Decoder with CUTLASS-FNA optimized operations

        decoder = build_cutlass_fna_diffusion_decoder_op(...)
        return DiffusionVideoDecoder(decoder, ...)
    else:
        # Fall back to convolutional decoder with default tiling

        return StaticConvolutionalDecoder(...)

During validation, ValidationRunner._load_decoder_components (lines 481–502 in ltx_trainer/validation_runner.py) inspects the returned object:

self._vae_decoder = load_video_vae_decoder(...)

if isinstance(self._vae_decoder, DiffusionVideoDecoder):
    # Add diffusion decoder weights to memory budget calculation

    memory_estimate += self._vae_decoder.estimated_memory_usage()
else:
    # Treat as convolutional decoder with standard memory profile

    pass

CLI and Python API Configuration

Use the NADiffusion Decoder (Highest Quality)

Requires a checkpoint with diffusion weights and a non-default DiffVAEMode:

python -m ltx_pipelines.ti2vid_two_stages \
    --model_path /path/to/checkpoint.safetensors \
    --diffvae-optimization CHUNKED_EAGER

Equivalent Python:

from ltx_pipelines.ti2vid_two_stages import Ti2VidTwoStages
from ltx_core.model.video_vae.transformer.config import DiffVAEMode

pipeline = Ti2VidTwoStages(
    model_path="/path/to/checkpoint.safetensors",
    diffvae_optimization=DiffVAEMode.CHUNKED_EAGER,
)

Force the Convolutional Decoder (Lower VRAM)

Either omit diffusion weights from the checkpoint path, or explicitly disable:

python -m ltx_pipelines.ti2vid_two_stages \
    --model_path /path/to/checkpoint.safetensors \
    --with-video-vae-decoder=True \
    --diffusion-vae=False

Direct Python control via load_video_vae_decoder:

from ltx_core.model.video_vae.model_configurator import load_video_vae_decoder
import torch

decoder = load_video_vae_decoder(
    video_vae_path="/path/to/checkpoint.safetensors",
    device="cuda",
    dtype=torch.bfloat16,
    diffusion_vae=False,  # Forces convolutional decoder

)

Skip Video VAE Decoder Entirely

For audio-only validation or v2a (video-to-audio) workflows:

python -m ltx_trainer.validate \
    --model_path /path/to/checkpoint.safetensors \
    --with-video-vae-decoder=False

Python:

pipeline = Ti2VidTwoStages(
    model_path="/path/to/checkpoint.safetensors",
    with_video_vae_decoder=False,
)

Key Source Files Reference

File Purpose
ltx_core/model/video_vae/transformer/config.py DiffVAEMode enum definition and defaults
ltx_core/model/video_vae/diffusion_video_decoder.py NADiffusion Decoder implementation
ltx_trainer/model_loader.py load_video_vae_decoder() selection logic (lines 194–226)
ltx_trainer/validation_runner.py Runtime decoder type detection (lines 481–502)
ltx_trainer/config.py CLI argument definitions: --diffvae-optimization (line 405), --with-video-vae-decoder (line 568)
ltx_pipelines/ti2vid_two_stages_mgpu.py Example pipeline propagating diffvae_optimization (line 66)

Performance and Quality Trade-offs

When you configure the video VAE decoder in LTX-2, consider these characteristics:

  • NADiffusion Decoder: Superior reconstruction fidelity, especially for fine details and temporal consistency. Requires CUTLASS-FNA kernel support. Higher VRAM footprint and longer inference time.

  • Convolutional Decoder: Faster tile-based decoding, lower memory requirements. Suitable for rapid prototyping, batch processing, or hardware-constrained deployments.

The DiffVAEMode values beyond CHUNKED_EAGER enable additional optimizations: CHUNKED_COMPILE and COMBINED_COMPILE use torch.compile for kernel fusion, while BLACKWELL_DSL targets NVIDIA Blackwell architecture-specific optimizations.

Summary

  • NADiffusion Decoder activates automatically with diffusion-enabled checkpoints and non-default DiffVAEMode settings
  • Convolutional Decoder serves as the fallback when diffusion weights are absent or explicitly disabled
  • Control via --diffvae-optimization (enum selection) and --with-video-vae-decoder (decoder loading toggle)
  • Direct Python API access through load_video_vae_decoder() with diffusion_vae=True/False
  • Core logic spans ltx_core model definitions and ltx_trainer loading infrastructure

Frequently Asked Questions

What happens if I set --diffvae-optimization but my checkpoint lacks diffusion weights?

The framework attempts to load diffusion weights and raises an error or warning depending on the pipeline. To avoid this, either use a checkpoint with diffusion decoder weights, or explicitly pass diffusion_vae=False to load_video_vae_decoder().

Can I switch decoders without changing checkpoints?

Not directly for the NADiffusion Decoder—it requires the corresponding weights in the checkpoint. However, you can force the convolutional decoder on any checkpoint by setting diffusion_vae=False. The convolutional decoder weights are standard VAE components present in all LTX-2 checkpoints.

Does --diffvae-optimization BLACKWELL_DSL require specific hardware?

Yes. The BLACKWELL_DSL mode targets NVIDIA Blackwell architecture GPUs and requires compatible CUDA/cuDNN versions with CUTLASS-FNA support. For older GPUs, use CHUNKED_EAGER or CHUNKED_COMPILE instead.

How do I verify which decoder is actually loaded at runtime?

Inspect ValidationRunner._vae_decoder or check the type after calling load_video_vae_decoder():

from ltx_core.model.video_vae.diffusion_video_decoder import DiffusionVideoDecoder

decoder = load_video_vae_decoder(...)
print("Using NADiffusion Decoder:", isinstance(decoder, DiffusionVideoDecoder))

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →