How to Use Image-Conditioned LoRA (IC-LoRA) for Video Transformations in LTX-2
LTX-2 provides a two-stage diffusion pipeline that uses Image-Conditioned LoRA (IC-LoRA) to transform videos by conditioning generation on reference videos, with automatic handling of spatial scaling, temporal subsampling, and optional attention masks.
LTX-2 by Lightricks introduces Image-Conditioned LoRA (IC-LoRA), a powerful mechanism for guiding video generation through reference videos and visual signals such as depth maps or pose sequences. This article explores how to leverage the ICLoraPipeline class and associated utilities in ltx_pipelines/ic_lora.py to perform video transformations using LoRA adapters with video conditioning.
The Two-Stage IC-LoRA Architecture
The ICLoraPipeline orchestrates a two-stage diffusion process that balances quality and computational efficiency. Understanding this architecture is essential for configuring your video transformations correctly.
Stage 1: Low-Resolution Generation with LoRA
Stage 1 generates a low-resolution video at half the target resolution (e.g., 256×256 for a 512×512 output) using a distilled diffusion model. This stage applies your LoRA adapters to condition the generation on reference videos.
In ltx_pipelines/ic_lora.py, the pipeline initializes Stage 1 by passing LoRA weights to the DiffusionStage class:
self.stage_1 = DiffusionStage(
# ... other parameters ...
loras=tuple(loras)
)
The LoRA files may embed reference_downscale_factor and reference_temporal_scale_factor metadata in their safetensors headers. The pipeline automatically reads these values via read_lora_reference_downscale_factor and read_lora_reference_temporal_scale_factor in ltx_pipelines/iclora_utils.py to resize and subsample reference videos accordingly. If you load multiple LoRAs with conflicting scale factors, the pipeline raises a clear error.
Stage 2: Upscaling and Refinement
Stage 2 upscales the Stage 1 output to the target resolution and refines it using another distilled model. Critically, this stage runs without LoRA adapters (pure distilled model) to ensure high-quality upsampling:
self.stage_2 = DiffusionStage(
# ... other parameters ...
loras=() # Empty tuple - no LoRA applied
)
You can skip Stage 2 using --skip-stage-2 for rapid prototyping or memory-constrained environments, though this yields half-resolution output.
Preparing Video Conditionings
The append_ic_lora_reference_video_conditionings function in ltx_pipelines/iclora_utils.py handles the conversion of reference videos into latent conditioning signals.
Reference Video Scaling
When you provide a reference video path, the system creates a VideoConditionByReferenceLatent object that undergoes automatic spatial downscaling and temporal subsampling based on the LoRA metadata. The utility functions read these parameters directly from the safetensors file:
read_lora_reference_downscale_factor: Determines spatial resize ratiosread_lora_reference_temporal_scale_factor: Controls frame subsampling rates
Spatial Attention Masks
For spatially varying influence, you can supply a mask video via --conditioning-attention-mask. The downsample_mask_video_to_latent function processes this mask by:
- Loading the video via
_load_mask_videoinic_lora.py - Normalizing values to
[0, 1] - Downsampling to latent token dimensions
- Multiplying by
conditioning_attention_strength
The resulting mask modulates the reference video's influence per pixel, where 0 ignores the reference and 1 applies full conditioning.
Command-Line Usage
The CLI implementation in ic_lora.py provides direct access to the pipeline through the main() function, which parses arguments including --video-conditioning, --conditioning-attention-mask, and --skip-stage-2.
Basic Video Transformation
python -m ltx_pipelines.ic_lora \
--distilled-checkpoint-path models/distilled.pt \
--spatial-upsampler-path models/upsampler.pt \
--gemma-root /path/to/gemma \
--lora path/to/ic_lora.safetensors 1.0 \
--video-conditioning path/to/reference.mp4 0.8 \
--prompt "A dancer performing in rain" \
--seed 42 \
--height 512 \
--width 512 \
--num-frames 24 \
--frame-rate 24 \
--output-path results/dance.mp4
With Spatial Masking
python -m ltx_pipelines.ic_lora \
--distilled-checkpoint-path models/distilled.pt \
--spatial-upsampler-path models/upsampler.pt \
--gemma-root /path/to/gemma \
--lora path/to/ic_lora.safetensors 1.0 \
--video-conditioning path/to/reference.mp4 0.8 \
--conditioning-attention-mask path/to/mask.mp4 0.5 \
--prompt "A dancer performing in rain" \
--height 512 \
--width 512 \
--num-frames 24 \
--output-path results/dance_masked.mp4
Quick Preview (Skip Stage 2)
python -m ltx_pipelines.ic_lora \
--distilled-checkpoint-path models/distilled.pt \
--spatial-upsampler-path models/upsampler.pt \
--gemma-root /path/to/gemma \
--lora path/to/ic_lora.safetensors 1.0 \
--video-conditioning path/to/reference.mp4 1.0 \
--skip-stage-2 \
--output-path quick_preview.mp4
Python API Implementation
For programmatic control, import ICLoraPipeline from ltx_pipelines.ic_lora and configure it with LoraPathStrengthAndSDOps objects.
Basic Pipeline Setup
from ltx_pipelines.ic_lora import ICLoraPipeline
from ltx_core.loader import LoraPathStrengthAndSDOps
from pathlib import Path
# Configure LoRA adapter
lora = LoraPathStrengthAndSDOps(
path=Path("models/ic_lora.safetensors"),
strength=1.0,
# The ops field is populated automatically by the loader
)
# Initialize pipeline
pipeline = ICLoraPipeline(
distilled_checkpoint_path="models/distilled.pt",
spatial_upsampler_path="models/upsampler.pt",
gemma_root="/path/to/gemma",
loras=[lora],
)
# Generate video
video, audio = pipeline(
prompt="A futuristic city skyline at sunset",
seed=123,
height=512,
width=512,
num_frames=30,
frame_rate=30,
images=[], # No image conditioning
video_conditioning=[("data/ref_video.mp4", 0.9)],
conditioning_attention_strength=0.8,
skip_stage_2=False,
)
# Encode and save
from ltx_pipelines.utils.media_io import encode_video
encode_video(video, fps=30, audio=audio, output_path="outputs/result.mp4")
Adding Spatial Masks Manually
To apply spatially varying attention weights, load a mask video using the internal _load_mask_video helper:
from ltx_pipelines.ic_lora import _load_mask_video
# Load mask at Stage 1 resolution (half target)
mask_tensor = _load_mask_video(
mask_path="data/mask.mp4",
height=256, # 512 // 2
width=256,
num_frames=24,
)
# Generate with mask
video, audio = pipeline(
prompt="A cat dancing on a beach",
seed=777,
height=512,
width=512,
num_frames=24,
frame_rate=24,
images=[],
video_conditioning=[("data/cat_ref.mp4", 1.0)],
conditioning_attention_strength=0.7,
conditioning_attention_mask=mask_tensor,
)
The mask tensor is normalized to [0, 1], downsampled to latent space via downsample_mask_video_to_latent, and multiplied by the conditioning_attention_strength scalar.
Multiple LoRA Adapters
You can pass multiple LoRA files to combine different conditioning signals:
lora1 = LoraPathStrengthAndSDOps(path=Path("models/motion_lora.safetensors"), strength=0.8)
lora2 = LoraPathStrengthAndSDOps(path=Path("models/style_lora.safetensors"), strength=0.6)
pipeline = ICLoraPipeline(
# ... paths ...
loras=[lora1, lora2],
)
Ensure that all LoRAs share compatible reference_downscale_factor and reference_temporal_scale_factor metadata, as conflicting values will raise an error during initialization.
Summary
- IC-LoRA enables video generation conditioned on reference videos through a specialized two-stage pipeline in
ltx_pipelines/ic_lora.py. - Stage 1 generates low-resolution output at half target resolution using LoRA adapters and automatic scaling via metadata in
iclora_utils.py. - Stage 2 upscales results using a distilled model without LoRA weights; skip this stage with
--skip-stage-2for faster iteration. - Video conditioning supports optional spatial masks via
downsample_mask_video_to_latentfor per-pixel control over reference influence. - Metadata automation eliminates manual calibration by reading
reference_downscale_factorandreference_temporal_scale_factordirectly from safetensors files.
Frequently Asked Questions
What is the difference between Stage 1 and Stage 2 in IC-LoRA?
Stage 1 generates a low-resolution video (half the target dimensions) using the distilled model with your LoRA adapters attached, processing reference videos according to their embedded metadata. Stage 2 upscales this result to full resolution using a separate distilled model without any LoRA weights, providing quality refinement. Stage 2 can be skipped using --skip-stage-2 to produce half-resolution outputs quickly.
How does LTX-2 handle resolution differences between reference and target videos?
The pipeline automatically reads reference_downscale_factor and reference_temporal_scale_factor from each LoRA's safetensors metadata via read_lora_reference_downscale_factor and read_lora_reference_temporal_scale_factor in iclora_utils.py. These values determine how the reference video is spatially resized and temporally subsampled to match the generation process. If multiple LoRAs specify conflicting scale factors, the pipeline raises an error requesting compatible adapters.
Can I use multiple LoRA adapters simultaneously with IC-LoRA?
Yes, you can pass multiple --lora arguments via CLI or a list of LoraPathStrengthAndSDOps objects via the Python API. However, all LoRAs must share identical reference_downscale_factor and reference_temporal_scale_factor metadata, as the pipeline applies a single scaling configuration per generation. Conflicting metadata triggers a validation error to prevent inconsistent conditioning.
How do I create a spatial attention mask for video conditioning?
Supply a grayscale video mask via --conditioning-attention-mask (CLI) or the conditioning_attention_mask parameter (Python API). The mask is loaded via _load_mask_video, normalized to [0, 1], and downsampled to latent space using downsample_mask_video_to_latent. Values of 0 completely block the reference video's influence at that pixel, while 1 applies full strength multiplied by your conditioning_attention_strength parameter.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →