DiffVAE Modes in LTX-2: CHUNKED_EAGER, NADiffusionDecoder, and ConvDecoder Explained
LTX-2 offers four DiffVAE decode presets—CHUNKED_EAGER, CHUNKED_COMPILE, COMBINED_COMPILE, and BLACKWELL_DSL—that trade off compilation time, inference speed, and VRAM usage for video generation.
LTX-2's DiffVAE modes control how the video VAE decodes latent representations back to pixel space. Each mode configures different attention backends, compilation strategies, and memory tiling patterns. Understanding these modes is essential for optimizing inference on different GPU hardware, from consumer cards to datacenter Blackwell systems.
What Are DiffVAE Modes in LTX-2?
DiffVAE modes are preset configurations defined in the DiffVAEMode enum at packages/ltx-core/src/ltx_core/model/video_vae/transformer/config.py. When you load a checkpoint, you can select a mode that determines:
- Whether to use chunked attention (width-split) or full-volume attention
- Whether to JIT-compile diffusion blocks and deterministic stages
- Which attention backend (natten, Triton/Eager, or Blackwell DSL) to invoke
The resolve() method in DiffVAEMode expands each preset into a concrete DiffVAEConfig with specific values for w_chunks, compile_blocks, compile_det_stages, and attention.
The Four DiffVAE Decode Presets
CHUNKED_EAGER: Minimal VRAM, No Compilation
CHUNKED_EAGER uses a width-chunked diffusion block (w_chunks=4) with no JIT compilation. It defers stage-4 context computation and runs entirely in eager mode or Triton fallback.
- Compile time: None (instant startup)
- Warm runtime: Slowest (~2–2.5× slower than
COMBINED_COMPILE) - Peak VRAM: ~50% of
COMBINED_COMPILE - Hardware: Works on any GPU; falls back to Triton/Eager if natten is unavailable
Use CHUNKED_EAGER when GPU memory is severely constrained (≤12 GB) or when you need immediate startup without compilation overhead.
CHUNKED_COMPILE: Balanced Speed and Memory
CHUNKED_COMPILE keeps the same chunked layout as CHUNKED_EAGER but compiles the non-deterministic blocks (compile_blocks=True). Deterministic stages remain uncompiled to reduce compile time.
- Compile time: ~2× faster than
COMBINED_COMPILE - Warm runtime: ~1.4× slower than
COMBINED_COMPILE - Peak VRAM: ~50% of
COMBINED_COMPILE - Hardware: Requires natten backend or falls back to Triton/Eager
Use CHUNKED_COMPILE with ~16 GB VRAM when you can afford a one-time compilation to gain moderate speed improvements without doubling memory usage.
COMBINED_COMPILE: Maximum Throughput
COMBINED_COMPILE uses a combined diffusion block (w_chunks=1) with full-volume attention and compiles both blocks and deterministic stages.
- Compile time: Longest (full compilation of all stages)
- Warm runtime: Fastest among DiffVAE presets
- Peak VRAM: Highest (~2×
CHUNKED_*modes) - Hardware: Requires natten backend
Use COMBINED_COMPILE with ≥24 GB VRAM for large-scale batch inference where maximum throughput outweighs compilation and memory costs.
BLACKWELL_DSL: Datacenter-Grade Performance
BLACKWELL_DSL targets NVIDIA Blackwell GPUs with a proprietary CuTe DSL kernel that fuses stage-5 operations and performs in-kernel upsampling plus context projection.
- Compile time: Comparable to
COMBINED_COMPILE - Warm runtime: Same as
COMBINED_COMPILE(fast) - Peak VRAM: Same as
COMBINED_COMPILE - Hardware: Requires Blackwell-DSL backend installation
Use BLACKWELL_DSL only in datacenter environments with the Blackwell CuTe runtime installed.
The Convolutional Decoder Alternative
LTX-2 also provides ConvVideoDecoder at packages/ltx-core/src/ltx_core/model/video_vae/conv_video_decoder.py, a separate architecture entirely.
When a checkpoint's metadata specifies decoder_kind="conv" (as noted in packages/ltx-trainer/src/ltx_trainer/validation_runner.py at line 81), DiffVAE presets are ignored and the conv decoder runs instead. The conv decoder executes single-pass 3D convolutions—fast and memory-efficient with no tiling configuration required.
NADiffusionDecoder (referenced in the question title) appears to be an internal alias or documentation reference to the natten-based diffusion decoder, which corresponds to the COMBINED_COMPILE and BLACKWELL_DSL modes that require the natten backend. The chunked modes use alternative attention implementations when natten is unavailable.
When to Use Each DiffVAE Mode
| Hardware Scenario | Recommended Mode | Rationale |
|---|---|---|
| ≤12 GB VRAM, consumer GPU | CHUNKED_EAGER |
Minimal memory, no compilation |
| ~16 GB VRAM, moderate batch | CHUNKED_COMPILE |
Compile once, gain 40% speed vs. eager |
| ≥24 GB VRAM, production inference | COMBINED_COMPILE |
Maximum throughput with natten |
| Blackwell datacenter GPU | BLACKWELL_DSL |
Fused kernels for large-scale serving |
Checkpoint has decoder_kind="conv" |
ConvVideoDecoder |
Fastest single-pass, ignore DiffVAE presets |
Code Examples
Selecting a DiffVAE Mode in Pipeline Construction
from ltx_core.model.video_vae.transformer import DiffVAEMode
from ltx_pipelines.utils.blocks import DiffusionStage
stage = DiffusionStage.from_checkpoint(
checkpoint_path="checkpoints/ltx2_19b.safetensors",
diffvae_optimization=DiffVAEMode.COMBINED_COMPILE,
)
Runtime Mode Override with Tiled Decoding
from ltx_core.model.video_vae.transformer import DiffVAEMode
from ltx_core.model.video_vae.video_vae import recommended_decode_tiling_config
tiling_cfg = recommended_decode_tiling_config(
height=latent.shape[-2] * 32,
width=latent.shape[-1] * 32,
num_frames=(latent.shape[2] - 1) * 8 + 1,
mode=DiffVAEMode.CHUNKED_COMPILE,
free_bytes=10 * 1024**3, # 10 GB free budget
)
chunks = list(decoder.tiled_decode(latent, tiling_config=tiling_cfg))
Inspecting Resolved Configuration
from ltx_core.model.video_vae.transformer import DiffVAEMode
cfg = DiffVAEMode.CHUNKED_COMPILE.resolve()
print(cfg)
# DiffVAEConfig(
# block=DiffVAEBlockKind.CHUNKED,
# w_chunks=4,
# natten_backend=None,
# attention=NAttentionKind.NATTEN,
# compile_blocks=True,
# compile_det_stages=False,
# )
Using the Convolutional Decoder
from ltx_core.model.video_vae.video_vae import ConvVideoDecoder
decoder = ConvVideoDecoder.from_checkpoint(
checkpoint_path="checkpoints/ltx2_19b.safetensors"
)
decoded = decoder(latent_tensor) # No tiling config needed
Key Source Files
| File Path | Purpose |
|---|---|
packages/ltx-core/src/ltx_core/model/video_vae/transformer/config.py |
Defines DiffVAEMode, DiffVAEConfig, and resolve() logic |
packages/ltx-core/src/ltx_core/model/video_vae/diffusion_video_decoder.py |
Implements diffusion-based decoding with mode-aware tiling |
packages/ltx-core/src/ltx_core/model/video_vae/conv_video_decoder.py |
Single-pass convolutional decoder |
packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py |
DiffusionStage wrapper accepting DiffVAEMode |
packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py |
recommended_decode_tiling_config helper |
Summary
- DiffVAE modes select between chunked and full-volume attention, compilation strategies, and hardware backends
- CHUNKED_EAGER sacrifices speed for minimal VRAM and zero compilation
- CHUNKED_COMPILE offers middle-ground performance with moderate compile cost
- COMBINED_COMPILE delivers maximum throughput on high-VRAM systems with natten
- BLACKWELL_DSL optimizes for datacenter Blackwell deployments
- ConvVideoDecoder bypasses DiffVAE entirely when checkpoints specify convolutional decoding
- The mode flows from
DiffVAEMode→DiffVAEConfig→ tiling configuration → decoder implementation
Frequently Asked Questions
What happens if natten is not installed?
LTX-2 falls back to Triton/Eager attention backends. Chunked modes (CHUNKED_EAGER, CHUNKED_COMPILE) degrade gracefully, but COMBINED_COMPILE and BLACKWELL_DSL will fail or fall back to slower implementations depending on your error handling.
Can I switch DiffVAE modes after loading a checkpoint?
Yes. The mode is passed to recommended_decode_tiling_config() at decode time, not cached at checkpoint load. Override mode in your tiling configuration call to change behavior without reloading weights.
Why does LTX-2 have both diffusion and convolutional decoders?
The diffusion decoder offers higher reconstruction quality through iterative refinement, configurable via DiffVAE modes for speed/quality tradeoffs. The convolutional decoder provides faster, deterministic single-pass decoding when quality requirements allow it. Checkpoint metadata selects which architecture to instantiate.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →