How to Use IC-LoRA for Video-to-Video Transformations in LTX-2: A Complete Guide

LTX-2 implements In-Context LoRA (IC-LoRA) for video-to-video transformations through the ICLoraPipeline inference class and VideoToVideoStrategy training strategy, enabling two-stage generation where a reference video's style and structure are transferred to a newly generated target video.

Video-to-video generation in LTX-2 relies on a specialized adaptation of Low-Rank Adaptation called In-Context LoRA (IC-LoRA). This technique, implemented in the Lightricks/LTX-2 repository, allows you to condition generation on an existing reference video—transferring its visual characteristics, motion patterns, or aesthetic style to new content. Unlike standard LoRA that learns static parameter updates, IC-LoRA operates in-context: the reference video is concatenated with the noisy target latent and processed jointly by the diffusion transformer.


Architecture Overview

LTX-2's IC-LoRA system spans both training and inference, with dedicated components handling each phase:

Component Role Source File
VideoToVideoStrategy Training logic that concatenates clean reference latents with noised target latents, builds conditioning masks, and computes loss only on the target portion packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py
ICLoraPipeline Two-stage inference pipeline: Stage 1 generates low-resolution video conditioned on reference; Stage 2 upsamples spatially packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py
append_ic_lora_reference_video_conditionings Converts file paths to VideoConditionByReferenceLatent items, handling resolution down-scaling and temporal subsampling packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py
Metadata helpers Reads reference_downscale_factor and reference_temporal_scale_factor from LoRA safetensors to ensure training/inference consistency packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py

How IC-LoRA Training Works for Video-to-Video

The VideoToVideoStrategy class orchestrates the training process. Here's the operational flow:

  1. Dataset structure — Each training sample must include a reference_latents/ directory containing pre-encoded reference video latents. The directory structure follows the standard LTX-2 dataset conventions.

  2. Batch preparation — In VideoToVideoStrategy.prepare_training_inputs (lines 51-78), the strategy:

    • Loads ref_latents alongside target latents
    • Computes the down-scale factor as target_resolution / reference_resolution
    • Concatenates latents into tensor shape [ref, noisy_target]
  3. Conditioning mask construction — Reference tokens are always treated as conditioning (timestep = 0). Target tokens may receive first-frame conditioning with probability first_frame_conditioning_p.

  4. Loss computation — The diffusion loss is calculated only on the target portion, preserving reference latents as pure conditioning signals.

The down-scale factor inferred during training is serialized into the LoRA safetensors metadata, ensuring the inference pipeline can replicate the exact spatial relationship between reference and target.


How to Run IC-LoRA Inference for Video-to-Video

The ICLoraPipeline (lines 60-75 of ic_lora.py) implements a two-stage generation process specifically designed for IC-LoRA video-to-video transformations.

Stage 1: IC-LoRA Generation

  • Loads reference down-scale and temporal scale factors from LoRA metadata via read_lora_reference_downscale_factor and read_lora_reference_temporal_scale_factor
  • Processes video_conditioning tuples of (file_path, strength) through append_ic_lora_reference_video_conditionings
  • Down-samples reference video to height // downscale_factor and width // downscale_factor
  • Denoises concatenated latent sequence with the distilled transformer

Stage 2: Spatial Upsampling

  • A 2× latent spatial upsampler refines the Stage 1 output to final resolution

Critical Requirements

  • Distilled checkpoints only — IC-LoRA requires ltx-2.5-22b-distilled-transformer-* models. Full (non-distilled) pipelines do not support this mechanism.
  • Matching metadata — All LoRA files must share identical reference_downscale_factor and reference_temporal_scale_factor values; conflicts raise ValueError.

Complete Code Example: Video-to-Video with IC-LoRA

from ltx_pipelines import ICLoraPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# -------------------------------------------------

# 1. Initialize pipeline with distilled components

# -------------------------------------------------

model_paths = ModelPaths(
    transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
    video_vae="models/ltx-2.5/vae/video_vae.safetensors",
    audio_vae="models/ltx-2.5/vae/audio_vae.safetensors",
    spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)

lora_paths = [
    # IC-LoRA trained with V2V strategy (contains reference_downscale_factor in metadata)

    "models/ltx-2.5/loras/ltx-2.5-22b-ic-lora-video.safetensors"
]

pipeline = ICLoraPipeline(
    model_paths=model_paths,
    spatial_upsampler_path=model_paths.spatial_upsampler,
    loras=lora_paths,
)

# -------------------------------------------------

# 2. Configure generation parameters

# -------------------------------------------------

prompt = "A serene mountain lake at dawn, matching the reference video's cinematic color grade"
seed = 12345
height, width = 512, 768
num_frames = 121          # Must satisfy 8k+1 for LTX-2 (e.g., 121, 129, 137)

frame_rate = 30.0

# Reference video for style/structure transfer: (path, strength)

video_conditioning = [
    ("assets/cinematic_reference.mp4", 1.0)   # 1.0 = full influence

]

# -------------------------------------------------

# 3. Execute generation

# -------------------------------------------------

frames_iter, audio, tiling_cfg = pipeline(
    prompt=prompt,
    seed=seed,
    height=height,
    width=width,
    num_frames=num_frames,
    frame_rate=frame_rate,
    images=[],                     # No image conditioning for pure V2V

    video_conditioning=video_conditioning,
    enhance_prompt=False,
)

# Consume generator to obtain frames

generated_frames = list(frames_iter)   # Each: torch.Tensor [3, H, W] in [0, 1]

Applying Spatial Masks for Selective IC-LoRA Influence

To restrict the reference video's influence to specific regions, provide a boolean attention mask. The pipeline automatically down-samples this mask to latent resolution via downsample_mask_video_to_latent in iclora_utils.py.

import torch

# Create mask at target resolution: True = apply reference influence

mask = torch.zeros((height, width), dtype=torch.bool)
mask[100:400, 200:600] = True   # Center region receives conditioning

frames_iter, audio, tiling_cfg = pipeline(
    prompt="Abstract flowing patterns, geometric style from reference",
    seed=45678,
    height=512,
    width=768,
    num_frames=121,
    video_conditioning=[("assets/geometric_style.mp4", 1.0)],
    conditioning_attention_mask=mask,
    conditioning_attention_strength=0.85,   # Scale mask effectiveness

    images=[],
    enhance_prompt=False,
)

The mask is processed by append_ic_lora_reference_video_conditionings (lines 93-105 of iclora_utils.py), which handles the coordinate transformation from pixel space to latent space.


Training Your Own IC-LoRA for Video-to-Video

To train a custom V2V IC-LoRA, use the provided configuration:


# From repository root

python -m ltx_trainer fit --config packages/ltx-trainer/configs/v2v_ic_lora.yaml

Key configuration parameters in v2v_ic_lora.yaml:

Parameter Purpose
training_strategy: VideoToVideoStrategy Enables reference/target latent concatenation
first_frame_conditioning_p Probability of conditioning target on first frame
lora_rank Rank of low-rank adaptation matrices

The trained LoRA will embed reference_downscale_factor and reference_temporal_scale_factor in its safetensors metadata, readable by read_lora_reference_downscale_factor during inference.


Key Implementation Files

Path Function
packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py VideoToVideoStrategy — training logic for V2V IC-LoRA
packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py ICLoraPipeline — two-stage inference pipeline
packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py Reference conditioning utilities and metadata readers
packages/ltx-trainer/configs/v2v_ic_lora.yaml Example training configuration
packages/ltx-trainer/docs/training-modes.md Documentation on IC-LoRA training modes

Summary

  • IC-LoRA for video-to-video uses concatenated reference and target latents with loss computed only on the target portion, implemented in VideoToVideoStrategy.
  • Inference requires ICLoraPipeline with distilled transformer checkpoints and LoRA files containing reference scaling metadata.
  • Reference video conditioning is specified via video_conditioning=[(path, strength), ...] and processed by append_ic_lora_reference_video_conditionings.
  • Spatial control is achieved through conditioning_attention_mask, automatically down-sampled to latent resolution.
  • Two-stage generation produces low-resolution draft followed by 2× spatial upsampling for final output.

Frequently Asked Questions

Can I use IC-LoRA with the full (non-distilled) LTX-2 model?

No. IC-LoRA is exclusively supported with distilled transformer checkpoints. The ICLoraPipeline expects the two-stage architecture (draft + upsampler) that only distilled models provide. Attempting to initialize with a full-model checkpoint will result in configuration errors or undefined behavior.

How does the pipeline know what resolution to down-scale my reference video?

The down-scale factor is read from LoRA metadata via read_lora_reference_downscale_factor in iclora_utils.py. This value was computed during training as the ratio of reference to target resolution. All LoRA files loaded together must have identical scaling factors; mismatches raise a ValueError to prevent inconsistent conditioning.

Can I use IC-LoRA for image-to-video generation?

Yes. Provide a single-frame video (or convert an image to a 1-frame video) as your reference. The pipeline treats this identically: the frame is encoded to latent space, concatenated with the noisy target latent, and processed by the transformer. Set num_frames in the target to your desired output length while the reference remains 1 frame.

What are the training dataset requirements for V2V IC-LoRA?

Each sample must include a reference_latents/ directory alongside the standard LTX-2 latent structure. The reference latents are encoded with the same VAE as targets. The VideoToVideoStrategy automatically infers spatial and temporal scaling relationships from the latent dimensions. Consult packages/ltx-trainer/docs/training-modes.md for directory layout specifications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →