Convolutional vs Diffusion Video VAE in LTX‑2: Architecture, Speed, and Use‑Case Comparison
LTX‑2 provides two video variational auto‑encoder (VAE) variants that share a common encoder but split into deterministic convolutional decoding (ConvVAE) or stochastic diffusion decoding (DiffVAE).
Both implementations live in Lightricks' LTX‑2 repository under packages/ltx-core/src/ltx_core/model/video_vae/. They encode video into a compressed latent space with identical VideoEncoder logic, then diverge in how they reconstruct pixels from those latents. This article breaks down the architectural differences, performance trade‑offs, and when to use each decoder based on the source code.
Core Architectural Difference
The fundamental split occurs in the decoder design:
- ConvVAE (
ConvVideoDecoder) — deterministic feed‑forward reconstruction through stacked convolutional upsampling blocks - DiffVAE (
DiffusionVideoDecoder) — iterative denoising process using neighborhood‑attention (NATTEN) blocks with a diffusion schedule
Both conform to the VideoDecoder protocol defined in video_vae.py, enabling drop‑in replacement across pipelines.
Decoder Implementation Details
Convolutional Video VAE (ConvVAE)
The convolutional decoder implements handcrafted upsampling via DepthToSpaceUpsample blocks in conv_video_decoder.py. The architecture progresses through three compression modes:
compress_time— temporal upsamplingcompress_space— spatial upsamplingcompress_all— combined spatiotemporal upsampling
Each block doubles resolution through depth‑to‑space rearrangement followed by residual processing. Temporal handling follows a specific formula: after temporal upsampling, the first frame is removed, yielding F = 8·(F'‑1)+1 output frames from F' latent frames.
Optional timestep conditioning (timestep_conditioning=True) injects learned noise with scale/shift modulation before the upsampling pipeline, though the core path remains deterministic.
from ltx_core.model.video_vae.model_configurator import VideoDecoderConfigurator
from ltx_core.loader.sft_loader import SafetensorsModelStateDictLoader
checkpoint_path = "path/to/conv_vae.safetensors"
metadata = SafetensorsModelStateDictLoader().metadata(checkpoint_path)
conv_decoder = VideoDecoderConfigurator.from_metadata(metadata)
latent = torch.randn(1, 128, 5, 16, 16)
video = conv_decoder(latent) # shape (1, 3, 33, 512, 512)
Diffusion Video VAE (DiffVAE)
The diffusion decoder in diffusion_video_decoder.py splits reconstruction into deterministic and stochastic phases:
Deterministic stages (1‑4): NATTEN‑based neighborhood attention upsampling to build context volume
Diffusion stage (5): Iterative denoising loop operating on noisy pixel field x_t
Key components include:
t_embedder— timestep embedding projection- Model output type configuration (
"v"or"x0") default_inference_timesteps— diffusion step count (default=2)
Temporal handling inserts NATTEN padding to prevent border artifacts, then executes the diffusion schedule. Unlike ConvVAE, each inference uses stochastic noise sampling (or fixed generator seed).
from ltx_core.model.video_vae.model_configurator import VideoDecoderConfigurator
from ltx_core.loader.sft_loader import SafetensorsModelStateDictLoader
checkpoint_path = "path/to/diff_vae.safetensors"
metadata = SafetensorsModelStateDictLoader().metadata(checkpoint_path)
diff_decoder = VideoDecoderConfigurator.from_metadata(metadata)
latent = torch.randn(1, 128, 5, 16, 16)
video = next(diff_decoder._decode_pixels(latent)) # shape (1, 3, 33, 512, 512)
Performance and Computational Trade‑offs
| Factor | ConvVAE | DiffVAE |
|---|---|---|
| Forward passes | Single pass through convolution stack | Multiple diffusion iterations (default 2+ steps) |
| Memory footprint | Lower — no intermediate context retention | Higher — maintains context tensors across steps |
| Inference speed | Faster — deterministic GPU utilization | Slower — iterative sampling overhead |
| Output quality | Standard fidelity | Higher fidelity for complex motion and texture |
| Determinism | Fully deterministic (without timestep conditioning) | Stochastic by design |
When to Use Each Video VAE in LTX‑2
Choose ConvVAE when:
- Real‑time or latency‑constrained generation is required
- Memory budgets are tight (edge deployment, batch processing)
- Standard video reconstruction quality suffices
- Deterministic, reproducible outputs are mandatory
Choose DiffVAE when:
- Maximum visual fidelity is prioritized over speed
- Complex motion patterns or fine textures need preservation
- Synthesis quality improvements justify computational cost
- The downstream pipeline can absorb variable inference times
Shared Infrastructure and Factory Loading
Both decoders instantiate through VideoDecoderConfigurator.from_metadata() in model_configurator.py, which inspects checkpoint metadata to dispatch the correct class. The shared VideoDecoder protocol ensures compatibility with:
- Tiling utilities for memory‑efficient large‑video decoding
decode_video()interface for pipeline integration- Per‑channel normalization from
ops.py
Summary
- ConvVAE (
ConvVideoDecoder) uses deterministicDepthToSpaceUpsampleblocks for single‑pass pixel reconstruction inconv_video_decoder.py - DiffVAE (
DiffusionVideoDecoder) applies NATTEN attention upsampling followed by iterative diffusion denoising indiffusion_video_decoder.py - Both share the
VideoEncoderandVideoDecoderprotocol fromvideo_vae.py, enabling unified pipeline integration - ConvVAE prioritizes speed and determinism; DiffVAE trades computational overhead for synthesis quality
- The
VideoDecoderConfiguratorfactory selects implementations automatically from checkpoint metadata
Frequently Asked Questions
How does LTX‑2 choose between ConvVAE and DiffVAE at runtime?
The VideoDecoderConfigurator.from_metadata() method inspects the Safetensors checkpoint metadata and instantiates either ConvVideoDecoder or DiffusionVideoDecoder based on stored configuration keys. Both classes implement the identical VideoDecoder protocol, so calling code requires no modification when switching between variants.
What causes the different output frame formulas between the two decoders?
ConvVAE removes the first frame after temporal upsampling to achieve F = 8·(F'‑1)+1, while DiffVAE preserves frames through its padded NATTEN stages and diffusion loop. This architectural choice in diffusion_video_decoder.py eliminates border artifacts that would otherwise appear from asymmetric temporal padding during denoising.
Can DiffVAE run with fewer diffusion steps for faster inference?
Yes. The default_inference_timesteps parameter controls step count (default=2). Reducing steps accelerates decoding but may degrade output quality. The source supports values as low as 1, though this approaches deterministic behavior without full diffusion benefits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →