How to Use IC-LoRA for Video-to-Video Transformations in LTX-2: A Complete Guide
LTX-2 implements In-Context LoRA (IC-LoRA) for video-to-video transformations through the ICLoraPipeline inference class and VideoToVideoStrategy training strategy, enabling two-stage generation where a reference video's style and structure are transferred to a newly generated target video.
Video-to-video generation in LTX-2 relies on a specialized adaptation of Low-Rank Adaptation called In-Context LoRA (IC-LoRA). This technique, implemented in the Lightricks/LTX-2 repository, allows you to condition generation on an existing reference video—transferring its visual characteristics, motion patterns, or aesthetic style to new content. Unlike standard LoRA that learns static parameter updates, IC-LoRA operates in-context: the reference video is concatenated with the noisy target latent and processed jointly by the diffusion transformer.
Architecture Overview
LTX-2's IC-LoRA system spans both training and inference, with dedicated components handling each phase:
| Component | Role | Source File |
|---|---|---|
VideoToVideoStrategy |
Training logic that concatenates clean reference latents with noised target latents, builds conditioning masks, and computes loss only on the target portion | packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py |
ICLoraPipeline |
Two-stage inference pipeline: Stage 1 generates low-resolution video conditioned on reference; Stage 2 upsamples spatially | packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py |
append_ic_lora_reference_video_conditionings |
Converts file paths to VideoConditionByReferenceLatent items, handling resolution down-scaling and temporal subsampling |
packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py |
| Metadata helpers | Reads reference_downscale_factor and reference_temporal_scale_factor from LoRA safetensors to ensure training/inference consistency |
packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py |
How IC-LoRA Training Works for Video-to-Video
The VideoToVideoStrategy class orchestrates the training process. Here's the operational flow:
-
Dataset structure — Each training sample must include a
reference_latents/directory containing pre-encoded reference video latents. The directory structure follows the standard LTX-2 dataset conventions. -
Batch preparation — In
VideoToVideoStrategy.prepare_training_inputs(lines 51-78), the strategy:- Loads
ref_latentsalongside target latents - Computes the down-scale factor as
target_resolution / reference_resolution - Concatenates latents into tensor shape
[ref, noisy_target]
- Loads
-
Conditioning mask construction — Reference tokens are always treated as conditioning (timestep = 0). Target tokens may receive first-frame conditioning with probability
first_frame_conditioning_p. -
Loss computation — The diffusion loss is calculated only on the target portion, preserving reference latents as pure conditioning signals.
The down-scale factor inferred during training is serialized into the LoRA safetensors metadata, ensuring the inference pipeline can replicate the exact spatial relationship between reference and target.
How to Run IC-LoRA Inference for Video-to-Video
The ICLoraPipeline (lines 60-75 of ic_lora.py) implements a two-stage generation process specifically designed for IC-LoRA video-to-video transformations.
Stage 1: IC-LoRA Generation
- Loads reference down-scale and temporal scale factors from LoRA metadata via
read_lora_reference_downscale_factorandread_lora_reference_temporal_scale_factor - Processes
video_conditioningtuples of(file_path, strength)throughappend_ic_lora_reference_video_conditionings - Down-samples reference video to
height // downscale_factorandwidth // downscale_factor - Denoises concatenated latent sequence with the distilled transformer
Stage 2: Spatial Upsampling
- A 2× latent spatial upsampler refines the Stage 1 output to final resolution
Critical Requirements
- Distilled checkpoints only — IC-LoRA requires
ltx-2.5-22b-distilled-transformer-*models. Full (non-distilled) pipelines do not support this mechanism. - Matching metadata — All LoRA files must share identical
reference_downscale_factorandreference_temporal_scale_factorvalues; conflicts raiseValueError.
Complete Code Example: Video-to-Video with IC-LoRA
from ltx_pipelines import ICLoraPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
# -------------------------------------------------
# 1. Initialize pipeline with distilled components
# -------------------------------------------------
model_paths = ModelPaths(
transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
video_vae="models/ltx-2.5/vae/video_vae.safetensors",
audio_vae="models/ltx-2.5/vae/audio_vae.safetensors",
spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)
lora_paths = [
# IC-LoRA trained with V2V strategy (contains reference_downscale_factor in metadata)
"models/ltx-2.5/loras/ltx-2.5-22b-ic-lora-video.safetensors"
]
pipeline = ICLoraPipeline(
model_paths=model_paths,
spatial_upsampler_path=model_paths.spatial_upsampler,
loras=lora_paths,
)
# -------------------------------------------------
# 2. Configure generation parameters
# -------------------------------------------------
prompt = "A serene mountain lake at dawn, matching the reference video's cinematic color grade"
seed = 12345
height, width = 512, 768
num_frames = 121 # Must satisfy 8k+1 for LTX-2 (e.g., 121, 129, 137)
frame_rate = 30.0
# Reference video for style/structure transfer: (path, strength)
video_conditioning = [
("assets/cinematic_reference.mp4", 1.0) # 1.0 = full influence
]
# -------------------------------------------------
# 3. Execute generation
# -------------------------------------------------
frames_iter, audio, tiling_cfg = pipeline(
prompt=prompt,
seed=seed,
height=height,
width=width,
num_frames=num_frames,
frame_rate=frame_rate,
images=[], # No image conditioning for pure V2V
video_conditioning=video_conditioning,
enhance_prompt=False,
)
# Consume generator to obtain frames
generated_frames = list(frames_iter) # Each: torch.Tensor [3, H, W] in [0, 1]
Applying Spatial Masks for Selective IC-LoRA Influence
To restrict the reference video's influence to specific regions, provide a boolean attention mask. The pipeline automatically down-samples this mask to latent resolution via downsample_mask_video_to_latent in iclora_utils.py.
import torch
# Create mask at target resolution: True = apply reference influence
mask = torch.zeros((height, width), dtype=torch.bool)
mask[100:400, 200:600] = True # Center region receives conditioning
frames_iter, audio, tiling_cfg = pipeline(
prompt="Abstract flowing patterns, geometric style from reference",
seed=45678,
height=512,
width=768,
num_frames=121,
video_conditioning=[("assets/geometric_style.mp4", 1.0)],
conditioning_attention_mask=mask,
conditioning_attention_strength=0.85, # Scale mask effectiveness
images=[],
enhance_prompt=False,
)
The mask is processed by append_ic_lora_reference_video_conditionings (lines 93-105 of iclora_utils.py), which handles the coordinate transformation from pixel space to latent space.
Training Your Own IC-LoRA for Video-to-Video
To train a custom V2V IC-LoRA, use the provided configuration:
# From repository root
python -m ltx_trainer fit --config packages/ltx-trainer/configs/v2v_ic_lora.yaml
Key configuration parameters in v2v_ic_lora.yaml:
| Parameter | Purpose |
|---|---|
training_strategy: VideoToVideoStrategy |
Enables reference/target latent concatenation |
first_frame_conditioning_p |
Probability of conditioning target on first frame |
lora_rank |
Rank of low-rank adaptation matrices |
The trained LoRA will embed reference_downscale_factor and reference_temporal_scale_factor in its safetensors metadata, readable by read_lora_reference_downscale_factor during inference.
Key Implementation Files
| Path | Function |
|---|---|
packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py |
VideoToVideoStrategy — training logic for V2V IC-LoRA |
packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py |
ICLoraPipeline — two-stage inference pipeline |
packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py |
Reference conditioning utilities and metadata readers |
packages/ltx-trainer/configs/v2v_ic_lora.yaml |
Example training configuration |
packages/ltx-trainer/docs/training-modes.md |
Documentation on IC-LoRA training modes |
Summary
- IC-LoRA for video-to-video uses concatenated reference and target latents with loss computed only on the target portion, implemented in
VideoToVideoStrategy. - Inference requires
ICLoraPipelinewith distilled transformer checkpoints and LoRA files containing reference scaling metadata. - Reference video conditioning is specified via
video_conditioning=[(path, strength), ...]and processed byappend_ic_lora_reference_video_conditionings. - Spatial control is achieved through
conditioning_attention_mask, automatically down-sampled to latent resolution. - Two-stage generation produces low-resolution draft followed by 2× spatial upsampling for final output.
Frequently Asked Questions
Can I use IC-LoRA with the full (non-distilled) LTX-2 model?
No. IC-LoRA is exclusively supported with distilled transformer checkpoints. The ICLoraPipeline expects the two-stage architecture (draft + upsampler) that only distilled models provide. Attempting to initialize with a full-model checkpoint will result in configuration errors or undefined behavior.
How does the pipeline know what resolution to down-scale my reference video?
The down-scale factor is read from LoRA metadata via read_lora_reference_downscale_factor in iclora_utils.py. This value was computed during training as the ratio of reference to target resolution. All LoRA files loaded together must have identical scaling factors; mismatches raise a ValueError to prevent inconsistent conditioning.
Can I use IC-LoRA for image-to-video generation?
Yes. Provide a single-frame video (or convert an image to a 1-frame video) as your reference. The pipeline treats this identically: the frame is encoded to latent space, concatenated with the noisy target latent, and processed by the transformer. Set num_frames in the target to your desired output length while the reference remains 1 frame.
What are the training dataset requirements for V2V IC-LoRA?
Each sample must include a reference_latents/ directory alongside the standard LTX-2 latent structure. The reference latents are encoded with the same VAE as targets. The VideoToVideoStrategy automatically infers spatial and temporal scaling relationships from the latent dimensions. Consult packages/ltx-trainer/docs/training-modes.md for directory layout specifications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →