How to Configure the Video VAE Decoder in LTX-2: NADiffusion Decoder vs. Convolutional
The LTX-2 video generation framework provides two VAE decoder options—the NADiffusion Decoder for maximum quality and a convolutional decoder for lower VRAM usage—controlled via checkpoint contents, the DiffVAEMode enum, and CLI flags.
LTX-2, developed by Lightricks, ships with dual video VAE decoder architectures. Understanding how to configure the video VAE decoder in LTX-2 is essential for optimizing quality, inference speed, and memory consumption. This guide explains the decision logic in the source code and provides exact commands to switch between implementations.
Decoder Options Overview
LTX-2 supports two decoder types:
| Decoder | Implementation | Selection Criteria |
|---|---|---|
| NADiffusion Decoder | ltx_core.model.video_vae.diffusion_video_decoder.DiffusionVideoDecoder |
Checkpoint contains diffusion-decoder weights and DiffVAEMode is non-default |
| Convolutional Decoder | ltx_core.model.video_vae (classic convolutional) |
Checkpoint lacks diffusion weights or DiffVAEMode left at default behavior |
The NADiffusion Decoder implements a "Minimal port of the reference NADiffusionDecoder" as noted at line 64 of diffusion_video_decoder.py. This decoder generally produces higher-quality reconstructions at the cost of increased computation and memory usage.
Three Factors That Determine Decoder Selection
The framework evaluates three conditions in order to decide which decoder to instantiate:
-
Checkpoint contents — If
video_vae_pathcontainsdiffusion_video_decoderweights, the framework can use the NADiffusion Decoder. -
DiffVAEModeenum — Defined inltx_core/model/video_vae/transformer/config.py. Non-default values (CHUNKED_EAGER,CHUNKED_COMPILE,COMBINED_COMPILE,BLACKWELL_DSL) trigger diffusion decoder construction. -
CLI/API configuration — Flags
--diffvae-optimization(sets the enum) and--with-video-vae-decoder(enables/disables decoder loading entirely).
How the Code Chooses the Decoder
The selection logic lives in ltx_trainer/model_loader.py within the load_video_vae_decoder() function (lines 194–226):
def load_video_vae_decoder(video_vae_path, device, dtype, diffusion_vae=True, ...):
if diffusion_vae:
# Build NADiffusion Decoder with CUTLASS-FNA optimized operations
decoder = build_cutlass_fna_diffusion_decoder_op(...)
return DiffusionVideoDecoder(decoder, ...)
else:
# Fall back to convolutional decoder with default tiling
return StaticConvolutionalDecoder(...)
During validation, ValidationRunner._load_decoder_components (lines 481–502 in ltx_trainer/validation_runner.py) inspects the returned object:
self._vae_decoder = load_video_vae_decoder(...)
if isinstance(self._vae_decoder, DiffusionVideoDecoder):
# Add diffusion decoder weights to memory budget calculation
memory_estimate += self._vae_decoder.estimated_memory_usage()
else:
# Treat as convolutional decoder with standard memory profile
pass
CLI and Python API Configuration
Use the NADiffusion Decoder (Highest Quality)
Requires a checkpoint with diffusion weights and a non-default DiffVAEMode:
python -m ltx_pipelines.ti2vid_two_stages \
--model_path /path/to/checkpoint.safetensors \
--diffvae-optimization CHUNKED_EAGER
Equivalent Python:
from ltx_pipelines.ti2vid_two_stages import Ti2VidTwoStages
from ltx_core.model.video_vae.transformer.config import DiffVAEMode
pipeline = Ti2VidTwoStages(
model_path="/path/to/checkpoint.safetensors",
diffvae_optimization=DiffVAEMode.CHUNKED_EAGER,
)
Force the Convolutional Decoder (Lower VRAM)
Either omit diffusion weights from the checkpoint path, or explicitly disable:
python -m ltx_pipelines.ti2vid_two_stages \
--model_path /path/to/checkpoint.safetensors \
--with-video-vae-decoder=True \
--diffusion-vae=False
Direct Python control via load_video_vae_decoder:
from ltx_core.model.video_vae.model_configurator import load_video_vae_decoder
import torch
decoder = load_video_vae_decoder(
video_vae_path="/path/to/checkpoint.safetensors",
device="cuda",
dtype=torch.bfloat16,
diffusion_vae=False, # Forces convolutional decoder
)
Skip Video VAE Decoder Entirely
For audio-only validation or v2a (video-to-audio) workflows:
python -m ltx_trainer.validate \
--model_path /path/to/checkpoint.safetensors \
--with-video-vae-decoder=False
Python:
pipeline = Ti2VidTwoStages(
model_path="/path/to/checkpoint.safetensors",
with_video_vae_decoder=False,
)
Key Source Files Reference
| File | Purpose |
|---|---|
ltx_core/model/video_vae/transformer/config.py |
DiffVAEMode enum definition and defaults |
ltx_core/model/video_vae/diffusion_video_decoder.py |
NADiffusion Decoder implementation |
ltx_trainer/model_loader.py |
load_video_vae_decoder() selection logic (lines 194–226) |
ltx_trainer/validation_runner.py |
Runtime decoder type detection (lines 481–502) |
ltx_trainer/config.py |
CLI argument definitions: --diffvae-optimization (line 405), --with-video-vae-decoder (line 568) |
ltx_pipelines/ti2vid_two_stages_mgpu.py |
Example pipeline propagating diffvae_optimization (line 66) |
Performance and Quality Trade-offs
When you configure the video VAE decoder in LTX-2, consider these characteristics:
-
NADiffusion Decoder: Superior reconstruction fidelity, especially for fine details and temporal consistency. Requires CUTLASS-FNA kernel support. Higher VRAM footprint and longer inference time.
-
Convolutional Decoder: Faster tile-based decoding, lower memory requirements. Suitable for rapid prototyping, batch processing, or hardware-constrained deployments.
The DiffVAEMode values beyond CHUNKED_EAGER enable additional optimizations: CHUNKED_COMPILE and COMBINED_COMPILE use torch.compile for kernel fusion, while BLACKWELL_DSL targets NVIDIA Blackwell architecture-specific optimizations.
Summary
- NADiffusion Decoder activates automatically with diffusion-enabled checkpoints and non-default
DiffVAEModesettings - Convolutional Decoder serves as the fallback when diffusion weights are absent or explicitly disabled
- Control via
--diffvae-optimization(enum selection) and--with-video-vae-decoder(decoder loading toggle) - Direct Python API access through
load_video_vae_decoder()withdiffusion_vae=True/False - Core logic spans
ltx_coremodel definitions andltx_trainerloading infrastructure
Frequently Asked Questions
What happens if I set --diffvae-optimization but my checkpoint lacks diffusion weights?
The framework attempts to load diffusion weights and raises an error or warning depending on the pipeline. To avoid this, either use a checkpoint with diffusion decoder weights, or explicitly pass diffusion_vae=False to load_video_vae_decoder().
Can I switch decoders without changing checkpoints?
Not directly for the NADiffusion Decoder—it requires the corresponding weights in the checkpoint. However, you can force the convolutional decoder on any checkpoint by setting diffusion_vae=False. The convolutional decoder weights are standard VAE components present in all LTX-2 checkpoints.
Does --diffvae-optimization BLACKWELL_DSL require specific hardware?
Yes. The BLACKWELL_DSL mode targets NVIDIA Blackwell architecture GPUs and requires compatible CUDA/cuDNN versions with CUTLASS-FNA support. For older GPUs, use CHUNKED_EAGER or CHUNKED_COMPILE instead.
How do I verify which decoder is actually loaded at runtime?
Inspect ValidationRunner._vae_decoder or check the type after calling load_video_vae_decoder():
from ltx_core.model.video_vae.diffusion_video_decoder import DiffusionVideoDecoder
decoder = load_video_vae_decoder(...)
print("Using NADiffusion Decoder:", isinstance(decoder, DiffusionVideoDecoder))
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →