# How to Convert Video Pixels to Latent Space Using LTX-2 Video VAE Encoding

> Learn how to convert video pixels to latent space with LTX-2 Video VAE encoding. Compress spatial and temporal dimensions efficiently for advanced video processing with this powerful technique.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: tutorial
- Published: 2026-06-20

---

**You convert video pixels to latent space in LTX-2 by preprocessing raw video into a `[C, F, H, W]` tensor, applying temporal subsampling aligned to the VAE's causal factor of 8, and encoding through the video VAE which compresses spatial dimensions by 32× and temporal dimensions by 8×.**

LTX-2 by Lightricks relies on a video variational autoencoder (VAE) to compress high-dimensional video data into efficient latent representations for diffusion training. Converting raw pixels to this latent space requires understanding the VAE's architectural constraints and following a specific preprocessing pipeline implemented in the official source code. This guide walks through the complete conversion process based on the actual implementation in the `Lightricks/LTX-2` repository.

## Understanding the LTX-2 Video VAE Compression Factors

The video VAE encoder defined in [`ltx_core/model/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae.py) performs aggressive spatio-temporal compression to create compact latent representations that the diffusion model processes.

### Spatial Compression (32×)

The VAE applies a **spatial factor of 32** (`VAE_SPATIAL_FACTOR = 32`), meaning each 32×32 pixel block in the original frame compresses to a single latent pixel. This reduces the spatial resolution significantly while preserving semantic information.

### Temporal Compression (8×)

The VAE uses a **causal temporal receptive field** with a factor of 8 (`VAE_TEMPORAL_FACTOR = 8`). Every group of 8 original frames maps to a single latent frame, with the exception of the first frame which is encoded independently as a standalone latent.

After encoding, the latent dimensions follow these formulas implemented in `_tiled_encode_video` (lines 701-705 of [`process_videos.py`](https://github.com/Lightricks/LTX-2/blob/main/process_videos.py)):

```text
latent_height = pixel_height // 32
latent_width  = pixel_width // 32
latent_frames = 1 + (pixel_frames - 1) // 8

```

## The Four-Stage Video-to-Latent Conversion Pipeline

The conversion pipeline in [`packages/ltx-trainer/scripts/process_videos.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/scripts/process_videos.py) consists of four logical stages that transform raw video bytes into storable latent tensors.

### Stage 1: Load and Preprocess Raw Media

The pipeline begins in `MediaDataset._preprocess_video` (lines 77-102), which handles the initial tensor conversion:

- Reads the video file into a tensor of shape **[C, F, H, W]** (channels, frames, height, width)
- Resizes and crops to a bucket-matched resolution based on predefined buckets
- Applies per-pixel normalization to the ±1 range
- Ensures the frame dimension exists for VAE compatibility

For video loading, the code uses `read_video` from `ltx_trainer.video_utils`, which returns the tensor along with the original FPS.

### Stage 2: VAE-Aligned Temporal Subsampling

Before encoding, the pipeline applies temporal subsampling in `_compute_temporal_subsample_indices` (lines 83-92) to align with the VAE's temporal factor:

- Always keeps frame 0 (encoded independently)
- Samples every N-th frame based on `temporal_subsample_factor`
- Ensures that after subsampling, groups of 8 frames will map to single latent frames

This subsampling occurs in `_preprocess_video` (lines 92-96) where the video tensor is indexed with the computed indices.

### Stage 3: Encode with the Video VAE

The core encoding happens in `_encode_video` (lines 606-632), which interfaces with the VAE model:

- Expands the input tensor from `[C, F, H, W]` to **[B, C, F, H, W]** (adding batch dimension)
- Moves the tensor to the VAE's device and dtype (typically `bfloat16`)
- Calls the VAE encoder directly or uses `_tiled_encode_video` (lines 673-718) for large videos

The VAE encoder outputs a dictionary containing the latent tensor with shape **[B, C, F', H', W']**, where spatial dimensions are divided by 32 and the temporal dimension follows the compression formula above.

### Stage 4: Post-Process and Store Latents

Finally, `compute_latents` (lines 544-564) handles persistence:

- Detaches the latent tensor and moves it to CPU
- Saves the `.pt` file containing the latent tensor and metadata
- Stores `num_frames`, `height`, `width`, and the **effective FPS** (original FPS divided by `temporal_subsample_factor`)

## Code Implementation Examples

### Minimal Script for Single Video Conversion

For processing individual videos, use the core utilities from [`process_videos.py`](https://github.com/Lightricks/LTX-2/blob/main/process_videos.py) and [`model_loader.py`](https://github.com/Lightricks/LTX-2/blob/main/model_loader.py):

```python
import torch
from pathlib import Path
from ltx_trainer.model_loader import load_video_vae_encoder
from ltx_trainer.video_utils import read_video
from ltx_trainer.scripts.process_videos import (
    _encode_video,
    VAE_SPATIAL_FACTOR,
    VAE_TEMPORAL_FACTOR,
)

def video_to_latent(
    video_path: Path,
    model_path: str,
    device: str = "cuda",
    temporal_subsample_factor: int = 1,
) -> dict[str, torch.Tensor]:
    # 1. Load raw frames [C, F, H, W]

    video_tensor, fps = read_video(video_path, max_frames=9999)
    
    # 2. Temporal subsampling (keep frame 0 + every N-th frame)

    if temporal_subsample_factor > 1:
        indices = [0] + list(range(1, video_tensor.shape[1], temporal_subsample_factor))
        video_tensor = video_tensor[:, indices, :, :]
    
    # 3. Load VAE encoder

    vae = load_video_vae_encoder(
        checkpoint_path=model_path,
        device=torch.device(device),
        dtype=torch.bfloat16,
    )
    
    # 4. Encode (no tiling for small videos)

    latent_dict = _encode_video(vae=vae, video=video_tensor, use_tiling=False)
    
    # 5. Attach effective FPS

    latent_dict["fps"] = fps / temporal_subsample_factor
    return latent_dict

# Usage

if __name__ == "__main__":
    latent = video_to_latent(
        video_path=Path("example.mp4"),
        model_path="ltx2.safetensors",
        temporal_subsample_factor=2,
    )
    print("Latent shape:", latent["latents"].shape)  # [C, F', H', W']

    print("Effective FPS:", latent["fps"])

```

### Batch Processing with the compute_latents CLI

For dataset-scale processing, use the high-level entry point:

```bash
python -m ltx_trainer.scripts.process_videos \
    dataset.csv \
    --video-column video_path \
    --resolution-buckets "25,768,768" "50,1024,1024" \
    --output-dir ./latents \
    --model-source ./ltx2.safetensors \
    --temporal-subsample-factor 2 \
    --batch-size 4 \
    --device cuda

```

This command builds a `MediaDataset`, applies the full preprocessing pipeline, and writes per-sample latent files under `./latents` with associated metadata.

### Tiling for High-Resolution Videos

When processing high-resolution videos that exceed GPU memory, enable spatial tiling:

```python
latent = _encode_video(
    vae=vae,
    video=video_tensor,
    use_tiling=True,          # Enable spatial tiling

    tile_size=1024,           # Must be multiple of 32

    tile_overlap=256,         # Must be multiple of 32

)

```

The `_tiled_encode_video` function splits frames into overlapping tiles, encodes each independently, and blends results using a linear feather mask (lines 674-718). This keeps peak memory usage proportional to `tile_size²` rather than full frame resolution.

## Summary

- **Load and preprocess** video using `read_video` to create `[C, F, H, W]` tensors with normalization to ±1 range.
- **Apply temporal subsampling** aligned to `VAE_TEMPORAL_FACTOR = 8`, always preserving frame 0 as an independent latent.
- **Encode** using `_encode_video` which wraps the VAE from [`ltx_core/model/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae.py), compressing spatially by 32× and temporally by 8×.
- **Handle large videos** by enabling tiling in `_tiled_encode_video` with parameters that are multiples of 32.
- **Store results** via `compute_latents` which saves latent tensors with metadata including effective FPS (original FPS divided by subsample factor).

## Frequently Asked Questions

### What tensor shape does the LTX-2 Video VAE expect as input?

The VAE encoder expects input tensors in **NCDHW format**: `[B, C, F, H, W]` where B is batch size, C is channels (typically 3), F is frames, H is height, and W is width. The raw preprocessing pipeline produces `[C, F, H, W]` tensors, and `_encode_video` automatically expands the batch dimension before passing to the model.

### Why does temporal subsampling always start with frame 0?

The LTX-2 Video VAE uses a causal temporal receptive field where the first frame is encoded as a standalone latent independent of subsequent frames. After frame 0, the VAE processes groups of 8 frames to produce single latent frames. Starting subsampling at frame 0 ensures this causal structure is preserved while reducing temporal redundancy.

### How do I handle videos that exceed available GPU memory?

Use the tiling implementation in `_tiled_encode_video` (lines 673-718 of [`process_videos.py`](https://github.com/Lightricks/LTX-2/blob/main/process_videos.py)) by setting `use_tiling=True` and specifying `tile_size` and `tile_overlap` (both must be multiples of 32). This splits large spatial dimensions into overlapping patches, encodes them sequentially, and blends the outputs with a linear feather mask to avoid visible seams.

### Where are the core VAE encoder architecture and loading utilities located?

The VAE architecture definition resides in [`packages/ltx-core/src/ltx_core/model/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae.py), which defines the spatial and temporal compression factors. The loading utilities are in [`packages/ltx-trainer/src/ltx_trainer/model_loader.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/model_loader.py), specifically the `load_video_vae_encoder` function that instantiates the encoder from a checkpoint file (`.safetensors` or `.pt`).