Convolutional vs Diffusion Video VAE in LTX‑2: Architecture, Speed, and Use‑Case Comparison

LTX‑2 provides two video variational auto‑encoder (VAE) variants that share a common encoder but split into deterministic convolutional decoding (ConvVAE) or stochastic diffusion decoding (DiffVAE).

Both implementations live in Lightricks' LTX‑2 repository under packages/ltx-core/src/ltx_core/model/video_vae/. They encode video into a compressed latent space with identical VideoEncoder logic, then diverge in how they reconstruct pixels from those latents. This article breaks down the architectural differences, performance trade‑offs, and when to use each decoder based on the source code.


Core Architectural Difference

The fundamental split occurs in the decoder design:

  • ConvVAE (ConvVideoDecoder) — deterministic feed‑forward reconstruction through stacked convolutional upsampling blocks
  • DiffVAE (DiffusionVideoDecoder) — iterative denoising process using neighborhood‑attention (NATTEN) blocks with a diffusion schedule

Both conform to the VideoDecoder protocol defined in video_vae.py, enabling drop‑in replacement across pipelines.


Decoder Implementation Details

Convolutional Video VAE (ConvVAE)

The convolutional decoder implements handcrafted upsampling via DepthToSpaceUpsample blocks in conv_video_decoder.py. The architecture progresses through three compression modes:

  • compress_time — temporal upsampling
  • compress_space — spatial upsampling
  • compress_all — combined spatiotemporal upsampling

Each block doubles resolution through depth‑to‑space rearrangement followed by residual processing. Temporal handling follows a specific formula: after temporal upsampling, the first frame is removed, yielding F = 8·(F'‑1)+1 output frames from F' latent frames.

Optional timestep conditioning (timestep_conditioning=True) injects learned noise with scale/shift modulation before the upsampling pipeline, though the core path remains deterministic.

from ltx_core.model.video_vae.model_configurator import VideoDecoderConfigurator
from ltx_core.loader.sft_loader import SafetensorsModelStateDictLoader

checkpoint_path = "path/to/conv_vae.safetensors"
metadata = SafetensorsModelStateDictLoader().metadata(checkpoint_path)

conv_decoder = VideoDecoderConfigurator.from_metadata(metadata)
latent = torch.randn(1, 128, 5, 16, 16)
video = conv_decoder(latent)  # shape (1, 3, 33, 512, 512)

Diffusion Video VAE (DiffVAE)

The diffusion decoder in diffusion_video_decoder.py splits reconstruction into deterministic and stochastic phases:

Deterministic stages (1‑4): NATTEN‑based neighborhood attention upsampling to build context volume

Diffusion stage (5): Iterative denoising loop operating on noisy pixel field x_t

Key components include:

  • t_embedder — timestep embedding projection
  • Model output type configuration ("v" or "x0")
  • default_inference_timesteps — diffusion step count (default=2)

Temporal handling inserts NATTEN padding to prevent border artifacts, then executes the diffusion schedule. Unlike ConvVAE, each inference uses stochastic noise sampling (or fixed generator seed).

from ltx_core.model.video_vae.model_configurator import VideoDecoderConfigurator
from ltx_core.loader.sft_loader import SafetensorsModelStateDictLoader

checkpoint_path = "path/to/diff_vae.safetensors"
metadata = SafetensorsModelStateDictLoader().metadata(checkpoint_path)

diff_decoder = VideoDecoderConfigurator.from_metadata(metadata)
latent = torch.randn(1, 128, 5, 16, 16)
video = next(diff_decoder._decode_pixels(latent))  # shape (1, 3, 33, 512, 512)

Performance and Computational Trade‑offs

Factor ConvVAE DiffVAE
Forward passes Single pass through convolution stack Multiple diffusion iterations (default 2+ steps)
Memory footprint Lower — no intermediate context retention Higher — maintains context tensors across steps
Inference speed Faster — deterministic GPU utilization Slower — iterative sampling overhead
Output quality Standard fidelity Higher fidelity for complex motion and texture
Determinism Fully deterministic (without timestep conditioning) Stochastic by design

When to Use Each Video VAE in LTX‑2

Choose ConvVAE when:

  • Real‑time or latency‑constrained generation is required
  • Memory budgets are tight (edge deployment, batch processing)
  • Standard video reconstruction quality suffices
  • Deterministic, reproducible outputs are mandatory

Choose DiffVAE when:

  • Maximum visual fidelity is prioritized over speed
  • Complex motion patterns or fine textures need preservation
  • Synthesis quality improvements justify computational cost
  • The downstream pipeline can absorb variable inference times

Shared Infrastructure and Factory Loading

Both decoders instantiate through VideoDecoderConfigurator.from_metadata() in model_configurator.py, which inspects checkpoint metadata to dispatch the correct class. The shared VideoDecoder protocol ensures compatibility with:

  • Tiling utilities for memory‑efficient large‑video decoding
  • decode_video() interface for pipeline integration
  • Per‑channel normalization from ops.py

Summary

  • ConvVAE (ConvVideoDecoder) uses deterministic DepthToSpaceUpsample blocks for single‑pass pixel reconstruction in conv_video_decoder.py
  • DiffVAE (DiffusionVideoDecoder) applies NATTEN attention upsampling followed by iterative diffusion denoising in diffusion_video_decoder.py
  • Both share the VideoEncoder and VideoDecoder protocol from video_vae.py, enabling unified pipeline integration
  • ConvVAE prioritizes speed and determinism; DiffVAE trades computational overhead for synthesis quality
  • The VideoDecoderConfigurator factory selects implementations automatically from checkpoint metadata

Frequently Asked Questions

How does LTX‑2 choose between ConvVAE and DiffVAE at runtime?

The VideoDecoderConfigurator.from_metadata() method inspects the Safetensors checkpoint metadata and instantiates either ConvVideoDecoder or DiffusionVideoDecoder based on stored configuration keys. Both classes implement the identical VideoDecoder protocol, so calling code requires no modification when switching between variants.

What causes the different output frame formulas between the two decoders?

ConvVAE removes the first frame after temporal upsampling to achieve F = 8·(F'‑1)+1, while DiffVAE preserves frames through its padded NATTEN stages and diffusion loop. This architectural choice in diffusion_video_decoder.py eliminates border artifacts that would otherwise appear from asymmetric temporal padding during denoising.

Can DiffVAE run with fewer diffusion steps for faster inference?

Yes. The default_inference_timesteps parameter controls step count (default=2). Reducing steps accelerates decoding but may degrade output quality. The source supports values as low as 1, though this approaches deterministic behavior without full diffusion benefits.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →