How to Use Image-Conditioned LoRA (IC-LoRA) for Video Transformations in LTX-2

LTX-2 provides a two-stage diffusion pipeline that uses Image-Conditioned LoRA (IC-LoRA) to transform videos by conditioning generation on reference videos, with automatic handling of spatial scaling, temporal subsampling, and optional attention masks.

LTX-2 by Lightricks introduces Image-Conditioned LoRA (IC-LoRA), a powerful mechanism for guiding video generation through reference videos and visual signals such as depth maps or pose sequences. This article explores how to leverage the ICLoraPipeline class and associated utilities in ltx_pipelines/ic_lora.py to perform video transformations using LoRA adapters with video conditioning.

The Two-Stage IC-LoRA Architecture

The ICLoraPipeline orchestrates a two-stage diffusion process that balances quality and computational efficiency. Understanding this architecture is essential for configuring your video transformations correctly.

Stage 1: Low-Resolution Generation with LoRA

Stage 1 generates a low-resolution video at half the target resolution (e.g., 256×256 for a 512×512 output) using a distilled diffusion model. This stage applies your LoRA adapters to condition the generation on reference videos.

In ltx_pipelines/ic_lora.py, the pipeline initializes Stage 1 by passing LoRA weights to the DiffusionStage class:

self.stage_1 = DiffusionStage(
    # ... other parameters ...

    loras=tuple(loras)
)

The LoRA files may embed reference_downscale_factor and reference_temporal_scale_factor metadata in their safetensors headers. The pipeline automatically reads these values via read_lora_reference_downscale_factor and read_lora_reference_temporal_scale_factor in ltx_pipelines/iclora_utils.py to resize and subsample reference videos accordingly. If you load multiple LoRAs with conflicting scale factors, the pipeline raises a clear error.

Stage 2: Upscaling and Refinement

Stage 2 upscales the Stage 1 output to the target resolution and refines it using another distilled model. Critically, this stage runs without LoRA adapters (pure distilled model) to ensure high-quality upsampling:

self.stage_2 = DiffusionStage(
    # ... other parameters ...

    loras=()  # Empty tuple - no LoRA applied

)

You can skip Stage 2 using --skip-stage-2 for rapid prototyping or memory-constrained environments, though this yields half-resolution output.

Preparing Video Conditionings

The append_ic_lora_reference_video_conditionings function in ltx_pipelines/iclora_utils.py handles the conversion of reference videos into latent conditioning signals.

Reference Video Scaling

When you provide a reference video path, the system creates a VideoConditionByReferenceLatent object that undergoes automatic spatial downscaling and temporal subsampling based on the LoRA metadata. The utility functions read these parameters directly from the safetensors file:

  • read_lora_reference_downscale_factor: Determines spatial resize ratios
  • read_lora_reference_temporal_scale_factor: Controls frame subsampling rates

Spatial Attention Masks

For spatially varying influence, you can supply a mask video via --conditioning-attention-mask. The downsample_mask_video_to_latent function processes this mask by:

  1. Loading the video via _load_mask_video in ic_lora.py
  2. Normalizing values to [0, 1]
  3. Downsampling to latent token dimensions
  4. Multiplying by conditioning_attention_strength

The resulting mask modulates the reference video's influence per pixel, where 0 ignores the reference and 1 applies full conditioning.

Command-Line Usage

The CLI implementation in ic_lora.py provides direct access to the pipeline through the main() function, which parses arguments including --video-conditioning, --conditioning-attention-mask, and --skip-stage-2.

Basic Video Transformation

python -m ltx_pipelines.ic_lora \
    --distilled-checkpoint-path models/distilled.pt \
    --spatial-upsampler-path models/upsampler.pt \
    --gemma-root /path/to/gemma \
    --lora path/to/ic_lora.safetensors 1.0 \
    --video-conditioning path/to/reference.mp4 0.8 \
    --prompt "A dancer performing in rain" \
    --seed 42 \
    --height 512 \
    --width 512 \
    --num-frames 24 \
    --frame-rate 24 \
    --output-path results/dance.mp4

With Spatial Masking

python -m ltx_pipelines.ic_lora \
    --distilled-checkpoint-path models/distilled.pt \
    --spatial-upsampler-path models/upsampler.pt \
    --gemma-root /path/to/gemma \
    --lora path/to/ic_lora.safetensors 1.0 \
    --video-conditioning path/to/reference.mp4 0.8 \
    --conditioning-attention-mask path/to/mask.mp4 0.5 \
    --prompt "A dancer performing in rain" \
    --height 512 \
    --width 512 \
    --num-frames 24 \
    --output-path results/dance_masked.mp4

Quick Preview (Skip Stage 2)

python -m ltx_pipelines.ic_lora \
    --distilled-checkpoint-path models/distilled.pt \
    --spatial-upsampler-path models/upsampler.pt \
    --gemma-root /path/to/gemma \
    --lora path/to/ic_lora.safetensors 1.0 \
    --video-conditioning path/to/reference.mp4 1.0 \
    --skip-stage-2 \
    --output-path quick_preview.mp4

Python API Implementation

For programmatic control, import ICLoraPipeline from ltx_pipelines.ic_lora and configure it with LoraPathStrengthAndSDOps objects.

Basic Pipeline Setup

from ltx_pipelines.ic_lora import ICLoraPipeline
from ltx_core.loader import LoraPathStrengthAndSDOps
from pathlib import Path

# Configure LoRA adapter

lora = LoraPathStrengthAndSDOps(
    path=Path("models/ic_lora.safetensors"),
    strength=1.0,
    # The ops field is populated automatically by the loader

)

# Initialize pipeline

pipeline = ICLoraPipeline(
    distilled_checkpoint_path="models/distilled.pt",
    spatial_upsampler_path="models/upsampler.pt",
    gemma_root="/path/to/gemma",
    loras=[lora],
)

# Generate video

video, audio = pipeline(
    prompt="A futuristic city skyline at sunset",
    seed=123,
    height=512,
    width=512,
    num_frames=30,
    frame_rate=30,
    images=[],  # No image conditioning

    video_conditioning=[("data/ref_video.mp4", 0.9)],
    conditioning_attention_strength=0.8,
    skip_stage_2=False,
)

# Encode and save

from ltx_pipelines.utils.media_io import encode_video
encode_video(video, fps=30, audio=audio, output_path="outputs/result.mp4")

Adding Spatial Masks Manually

To apply spatially varying attention weights, load a mask video using the internal _load_mask_video helper:

from ltx_pipelines.ic_lora import _load_mask_video

# Load mask at Stage 1 resolution (half target)

mask_tensor = _load_mask_video(
    mask_path="data/mask.mp4",
    height=256,  # 512 // 2

    width=256,
    num_frames=24,
)

# Generate with mask

video, audio = pipeline(
    prompt="A cat dancing on a beach",
    seed=777,
    height=512,
    width=512,
    num_frames=24,
    frame_rate=24,
    images=[],
    video_conditioning=[("data/cat_ref.mp4", 1.0)],
    conditioning_attention_strength=0.7,
    conditioning_attention_mask=mask_tensor,
)

The mask tensor is normalized to [0, 1], downsampled to latent space via downsample_mask_video_to_latent, and multiplied by the conditioning_attention_strength scalar.

Multiple LoRA Adapters

You can pass multiple LoRA files to combine different conditioning signals:

lora1 = LoraPathStrengthAndSDOps(path=Path("models/motion_lora.safetensors"), strength=0.8)
lora2 = LoraPathStrengthAndSDOps(path=Path("models/style_lora.safetensors"), strength=0.6)

pipeline = ICLoraPipeline(
    # ... paths ...

    loras=[lora1, lora2],
)

Ensure that all LoRAs share compatible reference_downscale_factor and reference_temporal_scale_factor metadata, as conflicting values will raise an error during initialization.

Summary

  • IC-LoRA enables video generation conditioned on reference videos through a specialized two-stage pipeline in ltx_pipelines/ic_lora.py.
  • Stage 1 generates low-resolution output at half target resolution using LoRA adapters and automatic scaling via metadata in iclora_utils.py.
  • Stage 2 upscales results using a distilled model without LoRA weights; skip this stage with --skip-stage-2 for faster iteration.
  • Video conditioning supports optional spatial masks via downsample_mask_video_to_latent for per-pixel control over reference influence.
  • Metadata automation eliminates manual calibration by reading reference_downscale_factor and reference_temporal_scale_factor directly from safetensors files.

Frequently Asked Questions

What is the difference between Stage 1 and Stage 2 in IC-LoRA?

Stage 1 generates a low-resolution video (half the target dimensions) using the distilled model with your LoRA adapters attached, processing reference videos according to their embedded metadata. Stage 2 upscales this result to full resolution using a separate distilled model without any LoRA weights, providing quality refinement. Stage 2 can be skipped using --skip-stage-2 to produce half-resolution outputs quickly.

How does LTX-2 handle resolution differences between reference and target videos?

The pipeline automatically reads reference_downscale_factor and reference_temporal_scale_factor from each LoRA's safetensors metadata via read_lora_reference_downscale_factor and read_lora_reference_temporal_scale_factor in iclora_utils.py. These values determine how the reference video is spatially resized and temporally subsampled to match the generation process. If multiple LoRAs specify conflicting scale factors, the pipeline raises an error requesting compatible adapters.

Can I use multiple LoRA adapters simultaneously with IC-LoRA?

Yes, you can pass multiple --lora arguments via CLI or a list of LoraPathStrengthAndSDOps objects via the Python API. However, all LoRAs must share identical reference_downscale_factor and reference_temporal_scale_factor metadata, as the pipeline applies a single scaling configuration per generation. Conflicting metadata triggers a validation error to prevent inconsistent conditioning.

How do I create a spatial attention mask for video conditioning?

Supply a grayscale video mask via --conditioning-attention-mask (CLI) or the conditioning_attention_mask parameter (Python API). The mask is loaded via _load_mask_video, normalized to [0, 1], and downsampled to latent space using downsample_mask_video_to_latent. Values of 0 completely block the reference video's influence at that pixel, while 1 applies full strength multiplied by your conditioning_attention_strength parameter.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →