# Diffusion VAE Decoder Optimization Modes in LTX-2: Complete Guide to `DiffVAEMode`

> Explore the four DiffVAEMode presets in LTX-2: CHUNKED_EAGER, CHUNKED_COMPILE, COMBINED_COMPILE, and BLACKWELL_DSL. Optimize your Diffusion VAE decoder for speed, compilation time, and memory.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-20

---

**LTX-2 provides four `DiffVAEMode` presets—`CHUNKED_EAGER`, `CHUNKED_COMPILE`, `COMBINED_COMPILE`, and `BLACKWELL_DSL`—that trade off between compilation time, inference speed, and GPU memory consumption when running the Diffusion VAE decoder.**

The **Diffusion VAE (DiffVAE)** decoder in LTX-2 is a transformer-based video generation component that can be optimized through several compilation and execution strategies. These modes are defined in [`ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/config.py) and control how diffusion blocks are built, whether torch.compile is applied, and which attention backends are used.

## The Four Diffusion VAE Decoder Optimization Modes

LTX-2 implements its optimization presets through the `DiffVAEMode` enum, which expands into concrete configurations via the `resolve()` method. Each mode selects a specific combination of block type, attention backend, deferred up-sampling behavior, and compilation flags.

### CHUNKED_EAGER: Lowest Memory, No Compilation

`CHUNKED_EAGER` uses **deferred stage-4 up-sampling** with a configurable number of **W-chunks** (default: 4). When the `natten` backend is unavailable, it automatically falls back to Triton or eager PyTorch kernels.

- **Compile time:** None (fully eager execution)
- **Inference speed:** ≈2–2.5× slower than `COMBINED_COMPILE`
- **Peak VRAM:** ≈50% of `COMBINED_COMPILE` (lowest overall)

This mode is ideal for rapid prototyping, debugging, or deployments where compilation overhead is unacceptable.

### CHUNKED_COMPILE: Balanced Compile Speed and Memory Efficiency

`CHUNKED_COMPILE` shares the same chunked architecture as `CHUNKED_EAGER` but applies **torch.compile** to the diffusion blocks only. This reduces compilation time compared to fully combined modes while retaining most memory benefits.

- **Compile time:** ≈2× faster than `COMBINED_COMPILE`
- **Inference speed:** ≈1.4× slower than `COMBINED_COMPILE`
- **Peak VRAM:** Same low usage as `CHUNKED_EAGER`

According to the LTX-2 source code, this mode represents a pragmatic middle ground for production deployments where memory constraints matter but some compilation overhead is tolerable.

### COMBINED_COMPILE: Maximum Speed, Highest VRAM

`COMBINED_COMPILE` runs diffusion blocks with **combined context** (full-volume attention) and **no deferred up-sampling**. This mode requires the `natten` backend and delivers the fastest raw inference throughput.

- **Compile time:** Slowest (both deterministic stages and diffusion blocks compiled)
- **Inference speed:** Baseline (fastest)
- **Peak VRAM:** Highest usage

Use this mode when GPU memory is abundant and minimum latency is the priority.

### BLACKWELL_DSL: Datacenter-Grade Kernel Fusion

`BLACKWELL_DSL` leverages NVIDIA's **Blackwell CuTe DSL** to fuse stage-4 up-sampling and context projection into a single kernel. It uses deferred stage-4 inputs like the chunked modes but achieves superior VRAM efficiency through custom DSL compilation.

- **Compile time:** Compiled (custom DSL kernel)
- **Inference speed:** Comparable to other compiled modes
- **Peak VRAM:** ≈2.5× lower than `COMBINED_COMPILE` (best efficiency)

This mode targets datacenter deployments with Blackwell-class GPUs where memory efficiency and throughput must be simultaneously optimized.

## Using DiffVAE Optimization Modes in LTX-2

### Setting a Mode in a VideoPipeline

The most common way to configure optimization is through the `video_vae_path` initialization:

```python
from ltx_core.model.video_vae.transformer import DiffVAEMode

# Use chunked-eager mode (default behavior)

pipeline = VideoPipeline(
    video_vae_path=model_paths.video_vae(),
    diffvae_optimization=DiffVAEMode.CHUNKED_EAGER,
)

```

### Applying a Mode to an Existing Decoder

For manual decoder manipulation, use `apply_diffvae_mode()` from [`ltx_core/model/video_vae/transformer/apply.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/apply.py):

```python
from ltx_core.model.video_vae.transformer.apply import apply_diffvae_mode
from ltx_core.model.video_vae.diffusion_video_decoder import DiffusionVideoDecoder
from ltx_core.model.video_vae.transformer.config import DiffVAEMode

# Instantiate the decoder

decoder = DiffusionVideoDecoder()

# Apply full compilation with combined context

apply_diffvae_mode(decoder, mode=DiffVAEMode.COMBINED_COMPILE)

```

This function performs **in-place module mutation** to reconfigure the decoder according to the selected mode's specifications.

### Inspecting Resolved Configurations

The `resolve()` method exposes the concrete parameters for each mode:

```python
cfg = DiffVAEMode.BLACKWELL_DSL.resolve()
print(cfg)

# Output:

# DiffVAEConfig(

#   block=DiffVAEBlockKind.BLACKWELL_DSL,

#   w_chunks=1,

#   natten_backend=None,

#   attention=NAttentionKind.BLACKWELL_DSL,

#   compile_blocks=True,

#   compile_det_stages=True,

# )

```

This allows programmatic inspection of which block kind, attention mechanism, and compilation flags a given mode selects.

## Key Source Files and Architecture

| File | Role |
|------|------|
| [`ltx_core/model/video_vae/transformer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/config.py) | Defines `DiffVAEMode` enum and `resolve()` logic |
| [`ltx_core/model/video_vae/diffusion_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/diffusion_video_decoder.py) | Implements the diffusion-based VAE decoder |
| [`ltx_core/model/video_vae/diffusion_tiling.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/diffusion_tiling.py) | Computes memory budgets and tile sizes per mode |
| [`ltx_core/model/video_vae/transformer/apply.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/apply.py) | Applies `DiffVAEMode` to decoder instances |

The configuration resolution in [`config.py`](https://github.com/Lightricks/LTX-2/blob/main/config.py) (lines 48-70) is the authoritative source for mode definitions, as implemented in Lightricks/LTX-2.

## Summary

- **LTX-2 offers four `DiffVAEMode` presets** that control how the Diffusion VAE decoder executes: `CHUNKED_EAGER`, `CHUNKED_COMPILE`, `COMBINED_COMPILE`, and `BLACKWELL_DSL`.
- **Memory vs. speed trade-offs** are explicit: chunked modes minimize VRAM, combined mode maximizes throughput, and Blackwell DSL optimizes both for datacenter GPUs.
- **No compilation is required** for `CHUNKED_EAGER`, making it the default for development and debugging workflows.
- **torch.compile integration** is granular: `CHUNKED_COMPILE` compiles only diffusion blocks, while `COMBINED_COMPILE` compiles both stages and blocks.
- **Blackwell DSL** provides the most memory-efficient path through custom kernel fusion, requiring compatible hardware.

## Frequently Asked Questions

### What is the default Diffusion VAE decoder optimization mode in LTX-2?

`CHUNKED_EAGER` is the default mode. It requires no compilation, uses approximately half the VRAM of `COMBINED_COMPILE`, and automatically falls back to standard PyTorch kernels when optimized backends are unavailable. This makes it suitable for development environments and diverse hardware configurations.

### How do I reduce GPU memory usage when decoding videos with LTX-2?

Select either `CHUNKED_EAGER` or `CHUNKED_COMPILE` for 50% lower peak VRAM versus `COMBINED_COMPILE`. For maximum efficiency on Blackwell GPUs, use `BLACKWELL_DSL`, which achieves approximately 2.5× lower memory consumption than the combined mode through fused kernel execution.

### Why is my LTX-2 pipeline taking so long to start?

Long startup times indicate that torch.compile is active. `COMBINED_COMPILE` has the slowest compilation because it compiles both deterministic stages and diffusion blocks. Switch to `CHUNKED_COMPILE` for 2× faster compilation with similar runtime performance, or use `CHUNKED_EAGER` to eliminate compilation entirely.

### Can I change the optimization mode after creating a decoder?

Yes. The `apply_diffvae_mode()` function in [`ltx_core/model/video_vae/transformer/apply.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer/apply.py) reconfigures an existing `DiffusionVideoDecoder` instance in place. This allows dynamic mode switching without rebuilding the pipeline, though compilation overhead will still apply if the new mode requires it.