How to Apply IC-LoRA for Video-to-Video Transformations Using LTX-2
IC-LoRA for video-to-video transformations in LTX-2 requires training a LoRA with the VideoToVideoStrategy and running inference through the ICLoraPipeline with reference video conditioning.
LTX-2 implements In-Context LoRA (IC-LoRA) through a specialized two-stage architecture that enables powerful video-to-video (V2V) transformations. This guide walks through the complete workflow—from training configuration to inference—based on the official Lightricks/LTX-2 source code.
Architecture Overview
IC-LoRA in LTX-2 splits responsibilities between training and inference components:
| Component | Purpose | Source Location |
|---|---|---|
VideoToVideoStrategy |
Training-side batch processing, reference/target latent concatenation, and conditioning mask construction | packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py |
ICLoraPipeline |
Inference-side two-stage generation with distilled transformer | packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py |
append_ic_lora_reference_video_conditionings() |
Converts video paths into structured conditioning items with automatic down-sampling | packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py |
read_lora_reference_downscale_factor() / read_lora_reference_temporal_scale_factor() |
Extract scaling metadata from LoRA safetensors for consistent resizing | packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py |
Training IC-LoRA for Video-to-Video
Dataset Requirements
Each training sample must contain a reference video latent directory (reference_latents/). The dataset layout follows the standard LTX-2 structure with this additional subdirectory.
Training Flow
The VideoToVideoStrategy.prepare_training_inputs method handles five critical operations:
- Loads reference latents (
ref_latents) alongside target latents - Infers down-scale factor—the ratio of reference resolution to target resolution
- Builds concatenated latent tensor
[ref, noisy_target] - Constructs conditioning masks—reference tokens always use timestep 0; target tokens may condition on first frame with probability
first_frame_conditioning_p - Computes loss only on target portion—reference latents remain untouched
This design forces the model to learn style/structure transfer rather than reconstruction.
Configuration
Use the example training config at packages/ltx-trainer/configs/v2v_ic_lora.yaml. Key parameters include:
first_frame_conditioning_p: Probability of conditioning target on reference first frame- Reference latent paths resolved relative to each sample
Inference with IC-LoRA for Video-to-Video
IC-LoRA inference requires distilled model checkpoints. Standard full-model pipelines do not support this feature.
Setup and Initialization
from ltx_pipelines import ICLoraPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
# Initialize with distilled transformer and upsampler
model_paths = ModelPaths(
transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
video_vae="models/ltx-2.5/vae/video_vae.safetensors",
audio_vae="models/ltx-2.5/vae/audio_vae.safetensors",
spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)
# LoRA must contain reference_downscale_factor in metadata
lora_paths = [
"models/ltx-2.5/loras/ltx-2.5-22b-ic-lora-video.safetensors"
]
pipeline = ICLoraPipeline(
model_paths=model_paths,
spatial_upsampler_path=model_paths.spatial_upsampler,
loras=lora_paths,
)
The ICLoraPipeline.__init__ automatically calls read_lora_reference_downscale_factor() and read_lora_reference_temporal_scale_factor() to extract scaling metadata. Conflicting factors across multiple LoRAs raise an error.
Generation with Reference Video
# Reference video provides style/structure to transfer
video_conditioning = [
("assets/reference_style.mp4", 1.0) # (path, strength)
]
frames_iter, audio, tiling_cfg = pipeline(
prompt="A sunrise over a futuristic city, in the style of the reference video",
seed=12345,
height=512,
width=768,
num_frames=121, # Must satisfy 8k+1 convention
frame_rate=30.0,
images=[], # No image conditioning
video_conditioning=video_conditioning,
enhance_prompt=False,
)
generated_frames = list(frames_iter)
The pipeline executes:
- Stage 1: Distilled transformer denoises concatenated
[ref, noisy_target]latents usingVideoConditionByReferenceLatentitems - Stage 2: Spatial upsampler (2×) refines to final resolution
The append_ic_lora_reference_video_conditionings helper in iclora_utils.py automatically down-samples the reference video to match target dimensions using the stored reference_downscale_factor.
Spatial Masking for Selective Application
Control where IC-LoRA influence applies using attention masks:
import torch
# Binary mask: True = apply reference conditioning
mask = torch.zeros((height, width), dtype=torch.bool)
mask[100:400, 200:600] = True # Center region only
frames_iter, audio, tiling_cfg = pipeline(
prompt="Futuristic city with selective style transfer",
seed=12345,
height=512,
width=768,
num_frames=121,
video_conditioning=[("assets/reference_style.mp4", 1.0)],
conditioning_attention_mask=mask,
conditioning_attention_strength=0.8, # Scale mask influence
)
The downsample_mask_video_to_latent utility (also in iclora_utils.py) handles resolution matching automatically.
Key Implementation Files
| Path | Description |
|---|---|
packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py |
VideoToVideoStrategy class with prepare_training_inputs() method |
packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py |
ICLoraPipeline two-stage inference implementation |
packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py |
Metadata readers, mask utilities, conditioning helpers |
packages/ltx-trainer/configs/v2v_ic_lora.yaml |
Example training configuration |
packages/ltx-trainer/docs/training-modes.md |
IC-LoRA training mode documentation |
Summary
- IC-LoRA enables video-to-video style transfer by conditioning generation on clean reference latents
- Training requires
VideoToVideoStrategywith reference latent directories and proper metadata storage - Inference requires
ICLoraPipelinewith distilled models and LoRAs containingreference_downscale_factor - Reference scaling is automatic—the pipeline reads metadata and applies consistent spatial/temporal subsampling
- Spatial control available via
conditioning_attention_maskfor selective region application
Frequently Asked Questions
What model type does IC-LoRA require?
IC-LoRA requires distilled transformer checkpoints only. The two-stage ICLoraPipeline is incompatible with standard full-model pipelines. Use paths like ltx-2.5-22b-distilled-transformer-bf16.safetensors.
How is the reference video resolution matched to the target?
The read_lora_reference_downscale_factor() function extracts the ratio from LoRA safetensors metadata. The reference is down-sampled to height // downscale_factor by width // downscale_factor before latent encoding.
Can I use multiple reference videos?
Yes. Pass multiple tuples to video_conditioning: [("ref1.mp4", 0.8), ("ref2.mp4", 0.5)]. Each becomes a separate VideoConditionByReferenceLatent. Ensure all referenced LoRAs share identical down-scale and temporal scale factors.
What frame count should I use for generation?
LTX-2 requires num_frames = 8k + 1 (e.g., 121, 129, 137). The temporal pipeline architecture enforces this constraint. Frame rates are flexible but 24-30 fps works best with most trained LoRAs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →