How to Convert Video Pixels to Latent Space Using LTX-2 Video VAE Encoding
You convert video pixels to latent space in LTX-2 by preprocessing raw video into a [C, F, H, W] tensor, applying temporal subsampling aligned to the VAE's causal factor of 8, and encoding through the video VAE which compresses spatial dimensions by 32× and temporal dimensions by 8×.
LTX-2 by Lightricks relies on a video variational autoencoder (VAE) to compress high-dimensional video data into efficient latent representations for diffusion training. Converting raw pixels to this latent space requires understanding the VAE's architectural constraints and following a specific preprocessing pipeline implemented in the official source code. This guide walks through the complete conversion process based on the actual implementation in the Lightricks/LTX-2 repository.
Understanding the LTX-2 Video VAE Compression Factors
The video VAE encoder defined in ltx_core/model/video_vae.py performs aggressive spatio-temporal compression to create compact latent representations that the diffusion model processes.
Spatial Compression (32×)
The VAE applies a spatial factor of 32 (VAE_SPATIAL_FACTOR = 32), meaning each 32×32 pixel block in the original frame compresses to a single latent pixel. This reduces the spatial resolution significantly while preserving semantic information.
Temporal Compression (8×)
The VAE uses a causal temporal receptive field with a factor of 8 (VAE_TEMPORAL_FACTOR = 8). Every group of 8 original frames maps to a single latent frame, with the exception of the first frame which is encoded independently as a standalone latent.
After encoding, the latent dimensions follow these formulas implemented in _tiled_encode_video (lines 701-705 of process_videos.py):
latent_height = pixel_height // 32
latent_width = pixel_width // 32
latent_frames = 1 + (pixel_frames - 1) // 8
The Four-Stage Video-to-Latent Conversion Pipeline
The conversion pipeline in packages/ltx-trainer/scripts/process_videos.py consists of four logical stages that transform raw video bytes into storable latent tensors.
Stage 1: Load and Preprocess Raw Media
The pipeline begins in MediaDataset._preprocess_video (lines 77-102), which handles the initial tensor conversion:
- Reads the video file into a tensor of shape [C, F, H, W] (channels, frames, height, width)
- Resizes and crops to a bucket-matched resolution based on predefined buckets
- Applies per-pixel normalization to the ±1 range
- Ensures the frame dimension exists for VAE compatibility
For video loading, the code uses read_video from ltx_trainer.video_utils, which returns the tensor along with the original FPS.
Stage 2: VAE-Aligned Temporal Subsampling
Before encoding, the pipeline applies temporal subsampling in _compute_temporal_subsample_indices (lines 83-92) to align with the VAE's temporal factor:
- Always keeps frame 0 (encoded independently)
- Samples every N-th frame based on
temporal_subsample_factor - Ensures that after subsampling, groups of 8 frames will map to single latent frames
This subsampling occurs in _preprocess_video (lines 92-96) where the video tensor is indexed with the computed indices.
Stage 3: Encode with the Video VAE
The core encoding happens in _encode_video (lines 606-632), which interfaces with the VAE model:
- Expands the input tensor from
[C, F, H, W]to [B, C, F, H, W] (adding batch dimension) - Moves the tensor to the VAE's device and dtype (typically
bfloat16) - Calls the VAE encoder directly or uses
_tiled_encode_video(lines 673-718) for large videos
The VAE encoder outputs a dictionary containing the latent tensor with shape [B, C, F', H', W'], where spatial dimensions are divided by 32 and the temporal dimension follows the compression formula above.
Stage 4: Post-Process and Store Latents
Finally, compute_latents (lines 544-564) handles persistence:
- Detaches the latent tensor and moves it to CPU
- Saves the
.ptfile containing the latent tensor and metadata - Stores
num_frames,height,width, and the effective FPS (original FPS divided bytemporal_subsample_factor)
Code Implementation Examples
Minimal Script for Single Video Conversion
For processing individual videos, use the core utilities from process_videos.py and model_loader.py:
import torch
from pathlib import Path
from ltx_trainer.model_loader import load_video_vae_encoder
from ltx_trainer.video_utils import read_video
from ltx_trainer.scripts.process_videos import (
_encode_video,
VAE_SPATIAL_FACTOR,
VAE_TEMPORAL_FACTOR,
)
def video_to_latent(
video_path: Path,
model_path: str,
device: str = "cuda",
temporal_subsample_factor: int = 1,
) -> dict[str, torch.Tensor]:
# 1. Load raw frames [C, F, H, W]
video_tensor, fps = read_video(video_path, max_frames=9999)
# 2. Temporal subsampling (keep frame 0 + every N-th frame)
if temporal_subsample_factor > 1:
indices = [0] + list(range(1, video_tensor.shape[1], temporal_subsample_factor))
video_tensor = video_tensor[:, indices, :, :]
# 3. Load VAE encoder
vae = load_video_vae_encoder(
checkpoint_path=model_path,
device=torch.device(device),
dtype=torch.bfloat16,
)
# 4. Encode (no tiling for small videos)
latent_dict = _encode_video(vae=vae, video=video_tensor, use_tiling=False)
# 5. Attach effective FPS
latent_dict["fps"] = fps / temporal_subsample_factor
return latent_dict
# Usage
if __name__ == "__main__":
latent = video_to_latent(
video_path=Path("example.mp4"),
model_path="ltx2.safetensors",
temporal_subsample_factor=2,
)
print("Latent shape:", latent["latents"].shape) # [C, F', H', W']
print("Effective FPS:", latent["fps"])
Batch Processing with the compute_latents CLI
For dataset-scale processing, use the high-level entry point:
python -m ltx_trainer.scripts.process_videos \
dataset.csv \
--video-column video_path \
--resolution-buckets "25,768,768" "50,1024,1024" \
--output-dir ./latents \
--model-source ./ltx2.safetensors \
--temporal-subsample-factor 2 \
--batch-size 4 \
--device cuda
This command builds a MediaDataset, applies the full preprocessing pipeline, and writes per-sample latent files under ./latents with associated metadata.
Tiling for High-Resolution Videos
When processing high-resolution videos that exceed GPU memory, enable spatial tiling:
latent = _encode_video(
vae=vae,
video=video_tensor,
use_tiling=True, # Enable spatial tiling
tile_size=1024, # Must be multiple of 32
tile_overlap=256, # Must be multiple of 32
)
The _tiled_encode_video function splits frames into overlapping tiles, encodes each independently, and blends results using a linear feather mask (lines 674-718). This keeps peak memory usage proportional to tile_size² rather than full frame resolution.
Summary
- Load and preprocess video using
read_videoto create[C, F, H, W]tensors with normalization to ±1 range. - Apply temporal subsampling aligned to
VAE_TEMPORAL_FACTOR = 8, always preserving frame 0 as an independent latent. - Encode using
_encode_videowhich wraps the VAE fromltx_core/model/video_vae.py, compressing spatially by 32× and temporally by 8×. - Handle large videos by enabling tiling in
_tiled_encode_videowith parameters that are multiples of 32. - Store results via
compute_latentswhich saves latent tensors with metadata including effective FPS (original FPS divided by subsample factor).
Frequently Asked Questions
What tensor shape does the LTX-2 Video VAE expect as input?
The VAE encoder expects input tensors in NCDHW format: [B, C, F, H, W] where B is batch size, C is channels (typically 3), F is frames, H is height, and W is width. The raw preprocessing pipeline produces [C, F, H, W] tensors, and _encode_video automatically expands the batch dimension before passing to the model.
Why does temporal subsampling always start with frame 0?
The LTX-2 Video VAE uses a causal temporal receptive field where the first frame is encoded as a standalone latent independent of subsequent frames. After frame 0, the VAE processes groups of 8 frames to produce single latent frames. Starting subsampling at frame 0 ensures this causal structure is preserved while reducing temporal redundancy.
How do I handle videos that exceed available GPU memory?
Use the tiling implementation in _tiled_encode_video (lines 673-718 of process_videos.py) by setting use_tiling=True and specifying tile_size and tile_overlap (both must be multiples of 32). This splits large spatial dimensions into overlapping patches, encodes them sequentially, and blends the outputs with a linear feather mask to avoid visible seams.
Where are the core VAE encoder architecture and loading utilities located?
The VAE architecture definition resides in packages/ltx-core/src/ltx_core/model/video_vae.py, which defines the spatial and temporal compression factors. The loading utilities are in packages/ltx-trainer/src/ltx_trainer/model_loader.py, specifically the load_video_vae_encoder function that instantiates the encoder from a checkpoint file (.safetensors or .pt).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →