# Convolutional vs Diffusion Video VAE in LTX‑2: Architecture, Speed, and Use‑Case Comparison

> Compare LTX-2 video VAEs: explore ConvVAE vs DiffVAE architectures, speed, and use cases. Understand their encoder and decoder differences for optimal performance.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: comparison
- Published: 2026-08-20

---

**LTX‑2 provides two video variational auto‑encoder (VAE) variants that share a common encoder but split into deterministic convolutional decoding (ConvVAE) or stochastic diffusion decoding (DiffVAE).**

Both implementations live in Lightricks' LTX‑2 repository under `packages/ltx-core/src/ltx_core/model/video_vae/`. They encode video into a compressed latent space with identical `VideoEncoder` logic, then diverge in how they reconstruct pixels from those latents. This article breaks down the architectural differences, performance trade‑offs, and when to use each decoder based on the source code.

---

## Core Architectural Difference

The fundamental split occurs in the decoder design:

- **ConvVAE** (`ConvVideoDecoder`) — deterministic feed‑forward reconstruction through stacked convolutional upsampling blocks
- **DiffVAE** (`DiffusionVideoDecoder`) — iterative denoising process using neighborhood‑attention (NATTEN) blocks with a diffusion schedule

Both conform to the `VideoDecoder` protocol defined in [`video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/video_vae.py), enabling drop‑in replacement across pipelines.

---

## Decoder Implementation Details

### Convolutional Video VAE (ConvVAE)

The convolutional decoder implements handcrafted upsampling via `DepthToSpaceUpsample` blocks in [`conv_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/conv_video_decoder.py). The architecture progresses through three compression modes:

- `compress_time` — temporal upsampling
- `compress_space` — spatial upsampling  
- `compress_all` — combined spatiotemporal upsampling

Each block doubles resolution through depth‑to‑space rearrangement followed by residual processing. Temporal handling follows a specific formula: after temporal upsampling, the first frame is removed, yielding `F = 8·(F'‑1)+1` output frames from `F'` latent frames.

Optional **timestep conditioning** (`timestep_conditioning=True`) injects learned noise with scale/shift modulation before the upsampling pipeline, though the core path remains deterministic.

```python
from ltx_core.model.video_vae.model_configurator import VideoDecoderConfigurator
from ltx_core.loader.sft_loader import SafetensorsModelStateDictLoader

checkpoint_path = "path/to/conv_vae.safetensors"
metadata = SafetensorsModelStateDictLoader().metadata(checkpoint_path)

conv_decoder = VideoDecoderConfigurator.from_metadata(metadata)
latent = torch.randn(1, 128, 5, 16, 16)
video = conv_decoder(latent)  # shape (1, 3, 33, 512, 512)

```

### Diffusion Video VAE (DiffVAE)

The diffusion decoder in [`diffusion_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/diffusion_video_decoder.py) splits reconstruction into deterministic and stochastic phases:

**Deterministic stages (1‑4):** NATTEN‑based neighborhood attention upsampling to build context volume

**Diffusion stage (5):** Iterative denoising loop operating on noisy pixel field `x_t`

Key components include:

- `t_embedder` — timestep embedding projection
- Model output type configuration (`"v"` or `"x0"`)
- `default_inference_timesteps` — diffusion step count (default=2)

Temporal handling inserts NATTEN padding to prevent border artifacts, then executes the diffusion schedule. Unlike ConvVAE, each inference uses stochastic noise sampling (or fixed generator seed).

```python
from ltx_core.model.video_vae.model_configurator import VideoDecoderConfigurator
from ltx_core.loader.sft_loader import SafetensorsModelStateDictLoader

checkpoint_path = "path/to/diff_vae.safetensors"
metadata = SafetensorsModelStateDictLoader().metadata(checkpoint_path)

diff_decoder = VideoDecoderConfigurator.from_metadata(metadata)
latent = torch.randn(1, 128, 5, 16, 16)
video = next(diff_decoder._decode_pixels(latent))  # shape (1, 3, 33, 512, 512)

```

---

## Performance and Computational Trade‑offs

| Factor | ConvVAE | DiffVAE |
|--------|---------|---------|
| **Forward passes** | Single pass through convolution stack | Multiple diffusion iterations (default 2+ steps) |
| **Memory footprint** | Lower — no intermediate context retention | Higher — maintains context tensors across steps |
| **Inference speed** | Faster — deterministic GPU utilization | Slower — iterative sampling overhead |
| **Output quality** | Standard fidelity | Higher fidelity for complex motion and texture |
| **Determinism** | Fully deterministic (without timestep conditioning) | Stochastic by design |

---

## When to Use Each Video VAE in LTX‑2

**Choose ConvVAE when:**
- Real‑time or latency‑constrained generation is required
- Memory budgets are tight (edge deployment, batch processing)
- Standard video reconstruction quality suffices
- Deterministic, reproducible outputs are mandatory

**Choose DiffVAE when:**
- Maximum visual fidelity is prioritized over speed
- Complex motion patterns or fine textures need preservation
- Synthesis quality improvements justify computational cost
- The downstream pipeline can absorb variable inference times

---

## Shared Infrastructure and Factory Loading

Both decoders instantiate through `VideoDecoderConfigurator.from_metadata()` in [`model_configurator.py`](https://github.com/Lightricks/LTX-2/blob/main/model_configurator.py), which inspects checkpoint metadata to dispatch the correct class. The shared `VideoDecoder` protocol ensures compatibility with:

- Tiling utilities for memory‑efficient large‑video decoding
- `decode_video()` interface for pipeline integration
- Per‑channel normalization from [`ops.py`](https://github.com/Lightricks/LTX-2/blob/main/ops.py)

---

## Summary

- **ConvVAE** (`ConvVideoDecoder`) uses deterministic `DepthToSpaceUpsample` blocks for single‑pass pixel reconstruction in [`conv_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/conv_video_decoder.py)
- **DiffVAE** (`DiffusionVideoDecoder`) applies NATTEN attention upsampling followed by iterative diffusion denoising in [`diffusion_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/diffusion_video_decoder.py)
- Both share the `VideoEncoder` and `VideoDecoder` protocol from [`video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/video_vae.py), enabling unified pipeline integration
- ConvVAE prioritizes speed and determinism; DiffVAE trades computational overhead for synthesis quality
- The `VideoDecoderConfigurator` factory selects implementations automatically from checkpoint metadata

---

## Frequently Asked Questions

### How does LTX‑2 choose between ConvVAE and DiffVAE at runtime?

The `VideoDecoderConfigurator.from_metadata()` method inspects the Safetensors checkpoint metadata and instantiates either `ConvVideoDecoder` or `DiffusionVideoDecoder` based on stored configuration keys. Both classes implement the identical `VideoDecoder` protocol, so calling code requires no modification when switching between variants.

### What causes the different output frame formulas between the two decoders?

ConvVAE removes the first frame after temporal upsampling to achieve `F = 8·(F'‑1)+1`, while DiffVAE preserves frames through its padded NATTEN stages and diffusion loop. This architectural choice in [`diffusion_video_decoder.py`](https://github.com/Lightricks/LTX-2/blob/main/diffusion_video_decoder.py) eliminates border artifacts that would otherwise appear from asymmetric temporal padding during denoising.

### Can DiffVAE run with fewer diffusion steps for faster inference?

Yes. The `default_inference_timesteps` parameter controls step count (default=2). Reducing steps accelerates decoding but may degrade output quality. The source supports values as low as 1, though this approaches deterministic behavior without full diffusion benefits.