How to Perform Video-to-Video Transformations with LTX-2: The Complete IC-LoRA Guide

LTX-2 enables video-to-video transformations by training an IC-LoRA (In-Context LoRA) adapter that conditions on reference video latents, allowing you to generate new videos that preserve the motion and structure of a source video while applying novel text-guided changes.

The Lightricks/LTX-2 repository provides a specialized workflow for video-to-video generation that goes beyond simple frame-to-frame editing. By leveraging the VideoToVideoStrategy class and a two-stage inference pipeline, you can encode reference videos into latent tensors, train a motion-preserving LoRA, and generate transformed videos that maintain temporal consistency with the original source.

Step 1: Encode Reference and Target Videos into Latents

Video-to-video transformations in LTX-2 operate on latent tensors rather than raw pixels. You must first encode both your reference videos (the source motion/style) and target videos (the desired outputs) using the Video VAE encoder.

Loading and Encoding Videos

The video_utils.py module provides utilities for reading video files and the ltx_core.loader module exposes the VAE encoder. According to the source code, the encoder outputs tensors of shape [C, F, H, W] (channels, frames, height, width).

from ltx_trainer.video_utils import read_video
from ltx_core.loader import load_video_vae_encoder
import torch

# Load raw frames (tensor shape [F, C, H, W]) and fps

raw_video, fps = read_video("ref.mp4")  # Returns torch.Tensor

# Initialize the VAE encoder (CPU is recommended for preprocessing)

vae_encoder = load_video_vae_encoder(
    model_path="path/to/video_vae", 
    device="cpu", 
    dtype=torch.bfloat16
)

# Encode to latent representation (shape [C, F, H, W])

ref_latent = vae_encoder.encode(raw_video.to(torch.bfloat16))

# Save for training (e.g., in a .pt file)

torch.save(ref_latent, "reference_latents/ref_001.pt")

Repeat this process for every target video, storing them in a separate directory (e.g., latents/). The training configuration will point to both directories via the reference_latents_dir parameter.

Step 2: Train with the VideoToVideoStrategy

The core of LTX-2's video-to-video capability resides in the VideoToVideoStrategy class located in packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py. This IC-LoRA implementation concatenates clean reference latents with noised target latents and computes loss only on the target portion.

Configuring the IC-LoRA Training Strategy

You must enable LoRA training and specify the video_to_video strategy in your configuration file. The config.py module validates that LoRA mode is active when using this strategy.


# config.yaml

training:
  strategy: video_to_video              # Activates IC-LoRA

  lora: true                            # Required for video-to-video

  reference_latents_dir: reference_latents  # Path to reference latents

  # Standard training parameters (batch size, learning rate, etc.)

Launching Training

Execute the training script with your configuration:

python -m ltx_trainer.scripts.train \
    --config config.yaml \
    --checkpoint-path /path/to/model_weights \
    --output-dir /output/ltx2_video2video

During training, the prepare_training_inputs method (lines 78-124 of video_to_video.py) performs three critical operations:

  1. Loads both target and reference latents from their respective directories.
  2. Infers the reference down-scale factor using _infer_reference_downscale_factor and rescales reference positions to align with the target coordinate space.
  3. Concatenates the clean reference latents with noisy target latents, then masks the loss computation to only the target portion.

This process trains a distilled LoRA that learns to inject reference motion into the generation process.

Step 3: Run Inference with Reference Conditioning

After training, you use the TI2VidTwoStagesPipeline to generate new videos by providing reference latents as conditioning inputs. The pipeline inserts reference latents as an unconditioned prefix and generates the target video on top of it.

Loading Reference Latents and Creating the Conditioner

Load your trained distilled LoRA and reference latent, then build a conditioning wrapper using combined_image_conditionings from utils/helpers.py:

import torch
from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline, OffloadMode
from ltx_core.loader import load_video_vae_encoder
from ltx_core.components.guiders import MultiModalGuiderParams, create_multimodal_guider_factory
from ltx_pipelines.utils.helpers import combined_image_conditionings

# Load the reference latent (must match the downscale factor used during training)

ref_latent = torch.load("reference_latents/ref_001.pt")  # Shape: [C, F, H, W]

# Build the conditioning function

def reference_conditioner(encoder):
    return combined_image_conditionings(
        images=[],                                 # No image conditioning

        height=ref_latent.shape[2],
        width=ref_latent.shape[3],
        video_encoder=encoder,
        dtype=torch.bfloat16,
        device=encoder.device,
        reference_latent=ref_latent               # Inject reference video

    )

Running the Two-Stage Pipeline

Initialize the pipeline with the trained checkpoint and run generation:


# Initialize pipeline (distilled LoRA is already baked into the checkpoint)

pipeline = TI2VidTwoStagesPipeline(
    checkpoint_path="/path/to/checkpoint",
    distilled_lora=[],               # IC-LoRA weights loaded from checkpoint

    spatial_upsampler_path="/path/to/upscaler",
    gemma_root="/path/to/gemma",
    loras=[],                       # Additional LoRAs if needed

    offload_mode=OffloadMode.NONE,
)

# Define the transformation prompt

prompt = "a sunrise over mountains"
negative_prompt = ""

# Configure guidance parameters

video_guidance = MultiModalGuiderParams(
    cfg_scale=7.0, 
    stg_scale=0.0, 
    rescale_scale=0.0
)

# Generate video

generated_video, generated_audio = pipeline(
    prompt=prompt,
    negative_prompt=negative_prompt,
    seed=42,
    height=512,
    width=512,
    num_frames=64,
    frame_rate=24.0,
    num_inference_steps=50,
    video_guider_params=video_guidance,
    audio_guider_params=MultiModalGuiderParams(cfg_scale=1.0),
    images=[],                     # Not used in video-to-video mode

    tiling_config=None,
    max_batch_size=1,
    stage_1_sigmas=None,
)

The pipeline executes in two stages: first generating a low-resolution video conditioned on the reference latent, then upsampling and refining using the distilled LoRA.

Saving the Output

Convert the generated tensor to a video file using encode_video from media_io.py:

from ltx_pipelines.utils.media_io import encode_video

encode_video(
    video=generated_video,
    fps=24.0,
    audio=generated_audio,
    output_path="output/video2video_result.mp4",
    video_chunks_number=1,
)

The resulting video preserves the motion dynamics of your reference while reflecting the new text prompt.

Summary

  • Video-to-video transformations in LTX-2 require encoding videos into latent tensors using the Video VAE located in ltx_core/model/video_vae.
  • The IC-LoRA strategy (VideoToVideoStrategy) concatenates clean reference latents with noisy targets and computes loss only on the target portion during training.
  • Training requires LoRA to be enabled (lora: true) and a valid reference_latents_dir in the configuration.
  • The reference down-scale factor is automatically inferred to align reference and target coordinate spaces.
  • Inference uses the two-stage pipeline (TI2VidTwoStagesPipeline) with reference latents injected via combined_image_conditionings.
  • Final outputs are encoded using encode_video from ltx_pipelines/utils/media_io.py.

Frequently Asked Questions

What is IC-LoRA in the context of LTX-2 video-to-video transformations?

IC-LoRA (In-Context LoRA) is the training strategy implemented in video_to_video.py that enables video-to-video transformations by treating reference video latents as context. It concatenates clean reference latents with noised target latents during the diffusion process, allowing the model to learn motion transfer while only computing loss on the target portion. This approach requires LoRA to be enabled because it trains a lightweight adapter that distills the reference motion patterns without modifying the base model weights.

Why must I use LoRA for video-to-video training in LTX-2?

The video_to_video strategy enforces LoRA training because the IC-LoRA approach is designed to learn a specialized, lightweight transformation that can be applied to different reference videos at inference time. As validated in config.py, setting strategy: video_to_video without lora: true will raise a configuration error. This design keeps the base model weights frozen while learning the motion-preserving adaptation in the LoRA layers, making the training efficient and the resulting model portable.

How does the reference down-scale factor work in LTX-2?

The reference down-scale factor is automatically inferred by the _infer_reference_downscale_factor method in video_to_video.py based on the spatial dimensions of your reference and target latents. If your reference video is encoded at a lower resolution than your target generation (e.g., 4× smaller), the strategy rescales the reference positions to align with the target coordinate space before concatenation. This ensures temporal consistency between the reference motion and the generated output regardless of resolution differences.

Can I perform video-to-video transformations without training a new model?

No, LTX-2 requires training a specific IC-LoRA adapter for video-to-video transformations. Unlike image-to-image translation that might work with pretrained models, the video-to-video workflow requires the model to learn the specific motion characteristics of your reference videos through the VideoToVideoStrategy. However, once trained, this distilled LoRA can be applied to any reference video of similar characteristics at inference time without retraining.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →