# How to Perform Video-to-Video Transformations with LTX-2: The Complete IC-LoRA Guide

> Learn video-to-video transformations with LTX-2's IC-LoRA guide. Generate new videos preserving source motion and structure with text-guided changes.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-06-21

---

**LTX-2 enables video-to-video transformations by training an IC-LoRA (In-Context LoRA) adapter that conditions on reference video latents, allowing you to generate new videos that preserve the motion and structure of a source video while applying novel text-guided changes.**

The Lightricks/LTX-2 repository provides a specialized workflow for video-to-video generation that goes beyond simple frame-to-frame editing. By leveraging the `VideoToVideoStrategy` class and a two-stage inference pipeline, you can encode reference videos into latent tensors, train a motion-preserving LoRA, and generate transformed videos that maintain temporal consistency with the original source.

## Step 1: Encode Reference and Target Videos into Latents

Video-to-video transformations in LTX-2 operate on **latent tensors** rather than raw pixels. You must first encode both your reference videos (the source motion/style) and target videos (the desired outputs) using the Video VAE encoder.

### Loading and Encoding Videos

The [`video_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/video_utils.py) module provides utilities for reading video files and the `ltx_core.loader` module exposes the VAE encoder. According to the source code, the encoder outputs tensors of shape **[C, F, H, W]** (channels, frames, height, width).

```python
from ltx_trainer.video_utils import read_video
from ltx_core.loader import load_video_vae_encoder
import torch

# Load raw frames (tensor shape [F, C, H, W]) and fps

raw_video, fps = read_video("ref.mp4")  # Returns torch.Tensor

# Initialize the VAE encoder (CPU is recommended for preprocessing)

vae_encoder = load_video_vae_encoder(
    model_path="path/to/video_vae", 
    device="cpu", 
    dtype=torch.bfloat16
)

# Encode to latent representation (shape [C, F, H, W])

ref_latent = vae_encoder.encode(raw_video.to(torch.bfloat16))

# Save for training (e.g., in a .pt file)

torch.save(ref_latent, "reference_latents/ref_001.pt")

```

Repeat this process for every target video, storing them in a separate directory (e.g., `latents/`). The training configuration will point to both directories via the `reference_latents_dir` parameter.

## Step 2: Train with the VideoToVideoStrategy

The core of LTX-2's video-to-video capability resides in the `VideoToVideoStrategy` class located in [`packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py). This IC-LoRA implementation concatenates clean reference latents with noised target latents and computes loss only on the target portion.

### Configuring the IC-LoRA Training Strategy

You must enable LoRA training and specify the `video_to_video` strategy in your configuration file. The [`config.py`](https://github.com/Lightricks/LTX-2/blob/main/config.py) module validates that LoRA mode is active when using this strategy.

```yaml

# config.yaml

training:
  strategy: video_to_video              # Activates IC-LoRA

  lora: true                            # Required for video-to-video

  reference_latents_dir: reference_latents  # Path to reference latents

  # Standard training parameters (batch size, learning rate, etc.)

```

### Launching Training

Execute the training script with your configuration:

```bash
python -m ltx_trainer.scripts.train \
    --config config.yaml \
    --checkpoint-path /path/to/model_weights \
    --output-dir /output/ltx2_video2video

```

During training, the `prepare_training_inputs` method (lines 78-124 of [`video_to_video.py`](https://github.com/Lightricks/LTX-2/blob/main/video_to_video.py)) performs three critical operations:

1. **Loads** both target and reference latents from their respective directories.
2. **Infers the reference down-scale factor** using `_infer_reference_downscale_factor` and rescales reference positions to align with the target coordinate space.
3. **Concatenates** the clean reference latents with noisy target latents, then masks the loss computation to only the target portion.

This process trains a **distilled LoRA** that learns to inject reference motion into the generation process.

## Step 3: Run Inference with Reference Conditioning

After training, you use the `TI2VidTwoStagesPipeline` to generate new videos by providing reference latents as conditioning inputs. The pipeline inserts reference latents as an unconditioned prefix and generates the target video on top of it.

### Loading Reference Latents and Creating the Conditioner

Load your trained distilled LoRA and reference latent, then build a conditioning wrapper using `combined_image_conditionings` from [`utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/helpers.py):

```python
import torch
from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline, OffloadMode
from ltx_core.loader import load_video_vae_encoder
from ltx_core.components.guiders import MultiModalGuiderParams, create_multimodal_guider_factory
from ltx_pipelines.utils.helpers import combined_image_conditionings

# Load the reference latent (must match the downscale factor used during training)

ref_latent = torch.load("reference_latents/ref_001.pt")  # Shape: [C, F, H, W]

# Build the conditioning function

def reference_conditioner(encoder):
    return combined_image_conditionings(
        images=[],                                 # No image conditioning

        height=ref_latent.shape[2],
        width=ref_latent.shape[3],
        video_encoder=encoder,
        dtype=torch.bfloat16,
        device=encoder.device,
        reference_latent=ref_latent               # Inject reference video

    )

```

### Running the Two-Stage Pipeline

Initialize the pipeline with the trained checkpoint and run generation:

```python

# Initialize pipeline (distilled LoRA is already baked into the checkpoint)

pipeline = TI2VidTwoStagesPipeline(
    checkpoint_path="/path/to/checkpoint",
    distilled_lora=[],               # IC-LoRA weights loaded from checkpoint

    spatial_upsampler_path="/path/to/upscaler",
    gemma_root="/path/to/gemma",
    loras=[],                       # Additional LoRAs if needed

    offload_mode=OffloadMode.NONE,
)

# Define the transformation prompt

prompt = "a sunrise over mountains"
negative_prompt = ""

# Configure guidance parameters

video_guidance = MultiModalGuiderParams(
    cfg_scale=7.0, 
    stg_scale=0.0, 
    rescale_scale=0.0
)

# Generate video

generated_video, generated_audio = pipeline(
    prompt=prompt,
    negative_prompt=negative_prompt,
    seed=42,
    height=512,
    width=512,
    num_frames=64,
    frame_rate=24.0,
    num_inference_steps=50,
    video_guider_params=video_guidance,
    audio_guider_params=MultiModalGuiderParams(cfg_scale=1.0),
    images=[],                     # Not used in video-to-video mode

    tiling_config=None,
    max_batch_size=1,
    stage_1_sigmas=None,
)

```

The pipeline executes in two stages: first generating a low-resolution video conditioned on the reference latent, then upsampling and refining using the distilled LoRA.

### Saving the Output

Convert the generated tensor to a video file using `encode_video` from [`media_io.py`](https://github.com/Lightricks/LTX-2/blob/main/media_io.py):

```python
from ltx_pipelines.utils.media_io import encode_video

encode_video(
    video=generated_video,
    fps=24.0,
    audio=generated_audio,
    output_path="output/video2video_result.mp4",
    video_chunks_number=1,
)

```

The resulting video preserves the motion dynamics of your reference while reflecting the new text prompt.

## Summary

- **Video-to-video transformations** in LTX-2 require encoding videos into latent tensors using the Video VAE located in `ltx_core/model/video_vae`.
- The **IC-LoRA strategy** (`VideoToVideoStrategy`) concatenates clean reference latents with noisy targets and computes loss only on the target portion during training.
- Training requires **LoRA to be enabled** (`lora: true`) and a valid `reference_latents_dir` in the configuration.
- The **reference down-scale factor** is automatically inferred to align reference and target coordinate spaces.
- Inference uses the **two-stage pipeline** (`TI2VidTwoStagesPipeline`) with reference latents injected via `combined_image_conditionings`.
- Final outputs are encoded using `encode_video` from [`ltx_pipelines/utils/media_io.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/media_io.py).

## Frequently Asked Questions

### What is IC-LoRA in the context of LTX-2 video-to-video transformations?

IC-LoRA (In-Context LoRA) is the training strategy implemented in [`video_to_video.py`](https://github.com/Lightricks/LTX-2/blob/main/video_to_video.py) that enables video-to-video transformations by treating reference video latents as context. It concatenates clean reference latents with noised target latents during the diffusion process, allowing the model to learn motion transfer while only computing loss on the target portion. This approach requires LoRA to be enabled because it trains a lightweight adapter that distills the reference motion patterns without modifying the base model weights.

### Why must I use LoRA for video-to-video training in LTX-2?

The `video_to_video` strategy enforces LoRA training because the IC-LoRA approach is designed to learn a specialized, lightweight transformation that can be applied to different reference videos at inference time. As validated in [`config.py`](https://github.com/Lightricks/LTX-2/blob/main/config.py), setting `strategy: video_to_video` without `lora: true` will raise a configuration error. This design keeps the base model weights frozen while learning the motion-preserving adaptation in the LoRA layers, making the training efficient and the resulting model portable.

### How does the reference down-scale factor work in LTX-2?

The reference down-scale factor is automatically inferred by the `_infer_reference_downscale_factor` method in [`video_to_video.py`](https://github.com/Lightricks/LTX-2/blob/main/video_to_video.py) based on the spatial dimensions of your reference and target latents. If your reference video is encoded at a lower resolution than your target generation (e.g., 4× smaller), the strategy rescales the reference positions to align with the target coordinate space before concatenation. This ensures temporal consistency between the reference motion and the generated output regardless of resolution differences.

### Can I perform video-to-video transformations without training a new model?

No, LTX-2 requires training a specific IC-LoRA adapter for video-to-video transformations. Unlike image-to-image translation that might work with pretrained models, the video-to-video workflow requires the model to learn the specific motion characteristics of your reference videos through the `VideoToVideoStrategy`. However, once trained, this distilled LoRA can be applied to any reference video of similar characteristics at inference time without retraining.