# How to Use IC-LoRA for Video-to-Video Transformations in LTX-2: A Complete Guide

> Master IC-LoRA for video-to-video transformations in LTX-2. This guide details using ICLoraPipeline and VideoToVideoStrategy for seamless style and structure transfer.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-20

---

**LTX-2 implements In-Context LoRA (IC-LoRA) for video-to-video transformations through the `ICLoraPipeline` inference class and `VideoToVideoStrategy` training strategy, enabling two-stage generation where a reference video's style and structure are transferred to a newly generated target video.**

Video-to-video generation in LTX-2 relies on a specialized adaptation of Low-Rank Adaptation called **In-Context LoRA (IC-LoRA)**. This technique, implemented in the Lightricks/LTX-2 repository, allows you to condition generation on an existing reference video—transferring its visual characteristics, motion patterns, or aesthetic style to new content. Unlike standard LoRA that learns static parameter updates, IC-LoRA operates in-context: the reference video is concatenated with the noisy target latent and processed jointly by the diffusion transformer.

---

## Architecture Overview

LTX-2's IC-LoRA system spans both training and inference, with dedicated components handling each phase:

| Component | Role | Source File |
|-----------|------|-------------|
| **`VideoToVideoStrategy`** | Training logic that concatenates clean reference latents with noised target latents, builds conditioning masks, and computes loss only on the target portion | [`packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py) |
| **`ICLoraPipeline`** | Two-stage inference pipeline: Stage 1 generates low-resolution video conditioned on reference; Stage 2 upsamples spatially | [`packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py) |
| **`append_ic_lora_reference_video_conditionings`** | Converts file paths to `VideoConditionByReferenceLatent` items, handling resolution down-scaling and temporal subsampling | [`packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py) |
| **Metadata helpers** | Reads `reference_downscale_factor` and `reference_temporal_scale_factor` from LoRA safetensors to ensure training/inference consistency | [`packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py) |

---

## How IC-LoRA Training Works for Video-to-Video

The `VideoToVideoStrategy` class orchestrates the training process. Here's the operational flow:

1. **Dataset structure** — Each training sample must include a `reference_latents/` directory containing pre-encoded reference video latents. The directory structure follows the standard LTX-2 dataset conventions.

2. **Batch preparation** — In `VideoToVideoStrategy.prepare_training_inputs` (lines 51-78), the strategy:
   - Loads `ref_latents` alongside target latents
   - Computes the down-scale factor as `target_resolution / reference_resolution`
   - Concatenates latents into tensor shape `[ref, noisy_target]`

3. **Conditioning mask construction** — Reference tokens are always treated as conditioning (timestep = 0). Target tokens may receive first-frame conditioning with probability `first_frame_conditioning_p`.

4. **Loss computation** — The diffusion loss is calculated **only on the target portion**, preserving reference latents as pure conditioning signals.

The down-scale factor inferred during training is serialized into the LoRA safetensors metadata, ensuring the inference pipeline can replicate the exact spatial relationship between reference and target.

---

## How to Run IC-LoRA Inference for Video-to-Video

The `ICLoraPipeline` (lines 60-75 of [`ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/ic_lora.py)) implements a two-stage generation process specifically designed for IC-LoRA video-to-video transformations.

### Stage 1: IC-LoRA Generation

- Loads reference down-scale and temporal scale factors from LoRA metadata via `read_lora_reference_downscale_factor` and `read_lora_reference_temporal_scale_factor`
- Processes `video_conditioning` tuples of `(file_path, strength)` through `append_ic_lora_reference_video_conditionings`
- Down-samples reference video to `height // downscale_factor` and `width // downscale_factor`
- Denoises concatenated latent sequence with the distilled transformer

### Stage 2: Spatial Upsampling

- A 2× latent spatial upsampler refines the Stage 1 output to final resolution

### Critical Requirements

- **Distilled checkpoints only** — IC-LoRA requires `ltx-2.5-22b-distilled-transformer-*` models. Full (non-distilled) pipelines do not support this mechanism.
- **Matching metadata** — All LoRA files must share identical `reference_downscale_factor` and `reference_temporal_scale_factor` values; conflicts raise `ValueError`.

---

## Complete Code Example: Video-to-Video with IC-LoRA

```python
from ltx_pipelines import ICLoraPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# -------------------------------------------------

# 1. Initialize pipeline with distilled components

# -------------------------------------------------

model_paths = ModelPaths(
    transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
    video_vae="models/ltx-2.5/vae/video_vae.safetensors",
    audio_vae="models/ltx-2.5/vae/audio_vae.safetensors",
    spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)

lora_paths = [
    # IC-LoRA trained with V2V strategy (contains reference_downscale_factor in metadata)

    "models/ltx-2.5/loras/ltx-2.5-22b-ic-lora-video.safetensors"
]

pipeline = ICLoraPipeline(
    model_paths=model_paths,
    spatial_upsampler_path=model_paths.spatial_upsampler,
    loras=lora_paths,
)

# -------------------------------------------------

# 2. Configure generation parameters

# -------------------------------------------------

prompt = "A serene mountain lake at dawn, matching the reference video's cinematic color grade"
seed = 12345
height, width = 512, 768
num_frames = 121          # Must satisfy 8k+1 for LTX-2 (e.g., 121, 129, 137)

frame_rate = 30.0

# Reference video for style/structure transfer: (path, strength)

video_conditioning = [
    ("assets/cinematic_reference.mp4", 1.0)   # 1.0 = full influence

]

# -------------------------------------------------

# 3. Execute generation

# -------------------------------------------------

frames_iter, audio, tiling_cfg = pipeline(
    prompt=prompt,
    seed=seed,
    height=height,
    width=width,
    num_frames=num_frames,
    frame_rate=frame_rate,
    images=[],                     # No image conditioning for pure V2V

    video_conditioning=video_conditioning,
    enhance_prompt=False,
)

# Consume generator to obtain frames

generated_frames = list(frames_iter)   # Each: torch.Tensor [3, H, W] in [0, 1]

```

---

## Applying Spatial Masks for Selective IC-LoRA Influence

To restrict the reference video's influence to specific regions, provide a boolean attention mask. The pipeline automatically down-samples this mask to latent resolution via `downsample_mask_video_to_latent` in [`iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/iclora_utils.py).

```python
import torch

# Create mask at target resolution: True = apply reference influence

mask = torch.zeros((height, width), dtype=torch.bool)
mask[100:400, 200:600] = True   # Center region receives conditioning

frames_iter, audio, tiling_cfg = pipeline(
    prompt="Abstract flowing patterns, geometric style from reference",
    seed=45678,
    height=512,
    width=768,
    num_frames=121,
    video_conditioning=[("assets/geometric_style.mp4", 1.0)],
    conditioning_attention_mask=mask,
    conditioning_attention_strength=0.85,   # Scale mask effectiveness

    images=[],
    enhance_prompt=False,
)

```

The mask is processed by `append_ic_lora_reference_video_conditionings` (lines 93-105 of [`iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/iclora_utils.py)), which handles the coordinate transformation from pixel space to latent space.

---

## Training Your Own IC-LoRA for Video-to-Video

To train a custom V2V IC-LoRA, use the provided configuration:

```bash

# From repository root

python -m ltx_trainer fit --config packages/ltx-trainer/configs/v2v_ic_lora.yaml

```

Key configuration parameters in [`v2v_ic_lora.yaml`](https://github.com/Lightricks/LTX-2/blob/main/v2v_ic_lora.yaml):

| Parameter | Purpose |
|-----------|---------|
| `training_strategy: VideoToVideoStrategy` | Enables reference/target latent concatenation |
| `first_frame_conditioning_p` | Probability of conditioning target on first frame |
| `lora_rank` | Rank of low-rank adaptation matrices |

The trained LoRA will embed `reference_downscale_factor` and `reference_temporal_scale_factor` in its safetensors metadata, readable by `read_lora_reference_downscale_factor` during inference.

---

## Key Implementation Files

| Path | Function |
|------|----------|
| [`packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py) | `VideoToVideoStrategy` — training logic for V2V IC-LoRA |
| [`packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py) | `ICLoraPipeline` — two-stage inference pipeline |
| [`packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py) | Reference conditioning utilities and metadata readers |
| [`packages/ltx-trainer/configs/v2v_ic_lora.yaml`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/configs/v2v_ic_lora.yaml) | Example training configuration |
| [`packages/ltx-trainer/docs/training-modes.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/docs/training-modes.md) | Documentation on IC-LoRA training modes |

---

## Summary

- **IC-LoRA for video-to-video** uses concatenated reference and target latents with loss computed only on the target portion, implemented in `VideoToVideoStrategy`.
- **Inference requires `ICLoraPipeline`** with distilled transformer checkpoints and LoRA files containing reference scaling metadata.
- **Reference video conditioning** is specified via `video_conditioning=[(path, strength), ...]` and processed by `append_ic_lora_reference_video_conditionings`.
- **Spatial control** is achieved through `conditioning_attention_mask`, automatically down-sampled to latent resolution.
- **Two-stage generation** produces low-resolution draft followed by 2× spatial upsampling for final output.

---

## Frequently Asked Questions

### Can I use IC-LoRA with the full (non-distilled) LTX-2 model?

No. IC-LoRA is exclusively supported with distilled transformer checkpoints. The `ICLoraPipeline` expects the two-stage architecture (draft + upsampler) that only distilled models provide. Attempting to initialize with a full-model checkpoint will result in configuration errors or undefined behavior.

### How does the pipeline know what resolution to down-scale my reference video?

The down-scale factor is read from LoRA metadata via `read_lora_reference_downscale_factor` in [`iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/iclora_utils.py). This value was computed during training as the ratio of reference to target resolution. All LoRA files loaded together must have identical scaling factors; mismatches raise a `ValueError` to prevent inconsistent conditioning.

### Can I use IC-LoRA for image-to-video generation?

Yes. Provide a single-frame video (or convert an image to a 1-frame video) as your reference. The pipeline treats this identically: the frame is encoded to latent space, concatenated with the noisy target latent, and processed by the transformer. Set `num_frames` in the target to your desired output length while the reference remains 1 frame.

### What are the training dataset requirements for V2V IC-LoRA?

Each sample must include a `reference_latents/` directory alongside the standard LTX-2 latent structure. The reference latents are encoded with the same VAE as targets. The `VideoToVideoStrategy` automatically infers spatial and temporal scaling relationships from the latent dimensions. Consult [`packages/ltx-trainer/docs/training-modes.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/docs/training-modes.md) for directory layout specifications.