DiffVAE Modes in LTX-2: CHUNKED_EAGER, NADiffusionDecoder, and ConvDecoder Explained

LTX-2 offers four DiffVAE decode presets—CHUNKED_EAGER, CHUNKED_COMPILE, COMBINED_COMPILE, and BLACKWELL_DSL—that trade off compilation time, inference speed, and VRAM usage for video generation.

LTX-2's DiffVAE modes control how the video VAE decodes latent representations back to pixel space. Each mode configures different attention backends, compilation strategies, and memory tiling patterns. Understanding these modes is essential for optimizing inference on different GPU hardware, from consumer cards to datacenter Blackwell systems.


What Are DiffVAE Modes in LTX-2?

DiffVAE modes are preset configurations defined in the DiffVAEMode enum at packages/ltx-core/src/ltx_core/model/video_vae/transformer/config.py. When you load a checkpoint, you can select a mode that determines:

  • Whether to use chunked attention (width-split) or full-volume attention
  • Whether to JIT-compile diffusion blocks and deterministic stages
  • Which attention backend (natten, Triton/Eager, or Blackwell DSL) to invoke

The resolve() method in DiffVAEMode expands each preset into a concrete DiffVAEConfig with specific values for w_chunks, compile_blocks, compile_det_stages, and attention.


The Four DiffVAE Decode Presets

CHUNKED_EAGER: Minimal VRAM, No Compilation

CHUNKED_EAGER uses a width-chunked diffusion block (w_chunks=4) with no JIT compilation. It defers stage-4 context computation and runs entirely in eager mode or Triton fallback.

  • Compile time: None (instant startup)
  • Warm runtime: Slowest (~2–2.5× slower than COMBINED_COMPILE)
  • Peak VRAM: ~50% of COMBINED_COMPILE
  • Hardware: Works on any GPU; falls back to Triton/Eager if natten is unavailable

Use CHUNKED_EAGER when GPU memory is severely constrained (≤12 GB) or when you need immediate startup without compilation overhead.

CHUNKED_COMPILE: Balanced Speed and Memory

CHUNKED_COMPILE keeps the same chunked layout as CHUNKED_EAGER but compiles the non-deterministic blocks (compile_blocks=True). Deterministic stages remain uncompiled to reduce compile time.

  • Compile time: ~2× faster than COMBINED_COMPILE
  • Warm runtime: ~1.4× slower than COMBINED_COMPILE
  • Peak VRAM: ~50% of COMBINED_COMPILE
  • Hardware: Requires natten backend or falls back to Triton/Eager

Use CHUNKED_COMPILE with ~16 GB VRAM when you can afford a one-time compilation to gain moderate speed improvements without doubling memory usage.

COMBINED_COMPILE: Maximum Throughput

COMBINED_COMPILE uses a combined diffusion block (w_chunks=1) with full-volume attention and compiles both blocks and deterministic stages.

  • Compile time: Longest (full compilation of all stages)
  • Warm runtime: Fastest among DiffVAE presets
  • Peak VRAM: Highest (~2× CHUNKED_* modes)
  • Hardware: Requires natten backend

Use COMBINED_COMPILE with ≥24 GB VRAM for large-scale batch inference where maximum throughput outweighs compilation and memory costs.

BLACKWELL_DSL: Datacenter-Grade Performance

BLACKWELL_DSL targets NVIDIA Blackwell GPUs with a proprietary CuTe DSL kernel that fuses stage-5 operations and performs in-kernel upsampling plus context projection.

  • Compile time: Comparable to COMBINED_COMPILE
  • Warm runtime: Same as COMBINED_COMPILE (fast)
  • Peak VRAM: Same as COMBINED_COMPILE
  • Hardware: Requires Blackwell-DSL backend installation

Use BLACKWELL_DSL only in datacenter environments with the Blackwell CuTe runtime installed.


The Convolutional Decoder Alternative

LTX-2 also provides ConvVideoDecoder at packages/ltx-core/src/ltx_core/model/video_vae/conv_video_decoder.py, a separate architecture entirely.

When a checkpoint's metadata specifies decoder_kind="conv" (as noted in packages/ltx-trainer/src/ltx_trainer/validation_runner.py at line 81), DiffVAE presets are ignored and the conv decoder runs instead. The conv decoder executes single-pass 3D convolutions—fast and memory-efficient with no tiling configuration required.

NADiffusionDecoder (referenced in the question title) appears to be an internal alias or documentation reference to the natten-based diffusion decoder, which corresponds to the COMBINED_COMPILE and BLACKWELL_DSL modes that require the natten backend. The chunked modes use alternative attention implementations when natten is unavailable.


When to Use Each DiffVAE Mode

Hardware Scenario Recommended Mode Rationale
≤12 GB VRAM, consumer GPU CHUNKED_EAGER Minimal memory, no compilation
~16 GB VRAM, moderate batch CHUNKED_COMPILE Compile once, gain 40% speed vs. eager
≥24 GB VRAM, production inference COMBINED_COMPILE Maximum throughput with natten
Blackwell datacenter GPU BLACKWELL_DSL Fused kernels for large-scale serving
Checkpoint has decoder_kind="conv" ConvVideoDecoder Fastest single-pass, ignore DiffVAE presets

Code Examples

Selecting a DiffVAE Mode in Pipeline Construction

from ltx_core.model.video_vae.transformer import DiffVAEMode
from ltx_pipelines.utils.blocks import DiffusionStage

stage = DiffusionStage.from_checkpoint(
    checkpoint_path="checkpoints/ltx2_19b.safetensors",
    diffvae_optimization=DiffVAEMode.COMBINED_COMPILE,
)

Runtime Mode Override with Tiled Decoding

from ltx_core.model.video_vae.transformer import DiffVAEMode
from ltx_core.model.video_vae.video_vae import recommended_decode_tiling_config

tiling_cfg = recommended_decode_tiling_config(
    height=latent.shape[-2] * 32,
    width=latent.shape[-1] * 32,
    num_frames=(latent.shape[2] - 1) * 8 + 1,
    mode=DiffVAEMode.CHUNKED_COMPILE,
    free_bytes=10 * 1024**3,  # 10 GB free budget

)
chunks = list(decoder.tiled_decode(latent, tiling_config=tiling_cfg))

Inspecting Resolved Configuration

from ltx_core.model.video_vae.transformer import DiffVAEMode

cfg = DiffVAEMode.CHUNKED_COMPILE.resolve()
print(cfg)

# DiffVAEConfig(

#   block=DiffVAEBlockKind.CHUNKED,

#   w_chunks=4,

#   natten_backend=None,

#   attention=NAttentionKind.NATTEN,

#   compile_blocks=True,

#   compile_det_stages=False,

# )

Using the Convolutional Decoder

from ltx_core.model.video_vae.video_vae import ConvVideoDecoder

decoder = ConvVideoDecoder.from_checkpoint(
    checkpoint_path="checkpoints/ltx2_19b.safetensors"
)
decoded = decoder(latent_tensor)  # No tiling config needed

Key Source Files

File Path Purpose
packages/ltx-core/src/ltx_core/model/video_vae/transformer/config.py Defines DiffVAEMode, DiffVAEConfig, and resolve() logic
packages/ltx-core/src/ltx_core/model/video_vae/diffusion_video_decoder.py Implements diffusion-based decoding with mode-aware tiling
packages/ltx-core/src/ltx_core/model/video_vae/conv_video_decoder.py Single-pass convolutional decoder
packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py DiffusionStage wrapper accepting DiffVAEMode
packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py recommended_decode_tiling_config helper

Summary

  • DiffVAE modes select between chunked and full-volume attention, compilation strategies, and hardware backends
  • CHUNKED_EAGER sacrifices speed for minimal VRAM and zero compilation
  • CHUNKED_COMPILE offers middle-ground performance with moderate compile cost
  • COMBINED_COMPILE delivers maximum throughput on high-VRAM systems with natten
  • BLACKWELL_DSL optimizes for datacenter Blackwell deployments
  • ConvVideoDecoder bypasses DiffVAE entirely when checkpoints specify convolutional decoding
  • The mode flows from DiffVAEMode → DiffVAEConfig → tiling configuration → decoder implementation

Frequently Asked Questions

What happens if natten is not installed?

LTX-2 falls back to Triton/Eager attention backends. Chunked modes (CHUNKED_EAGER, CHUNKED_COMPILE) degrade gracefully, but COMBINED_COMPILE and BLACKWELL_DSL will fail or fall back to slower implementations depending on your error handling.

Can I switch DiffVAE modes after loading a checkpoint?

Yes. The mode is passed to recommended_decode_tiling_config() at decode time, not cached at checkpoint load. Override mode in your tiling configuration call to change behavior without reloading weights.

Why does LTX-2 have both diffusion and convolutional decoders?

The diffusion decoder offers higher reconstruction quality through iterative refinement, configurable via DiffVAE modes for speed/quality tradeoffs. The convolutional decoder provides faster, deterministic single-pass decoding when quality requirements allow it. Checkpoint metadata selects which architecture to instantiate.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →