Using IC-LoRA for Video-to-Video Transformation with Reference Images in LTX-2

The IC-LoRA pipeline in LTX-2 enables video-to-video transformation by conditioning generation on reference videos through In-Context LoRA adapters in a two-stage distilled diffusion process.

This guide explains how to use IC-LoRA (In-Context LoRA) for video-to-video transformation with reference images in the Lightricks/LTX-2 repository. The pipeline allows you to steer video generation using reference videos or image sequences, with fine-grained control over spatial and temporal conditioning.

How the IC-LoRA Pipeline Works

The ICLoraPipeline implements a two-stage distilled diffusion architecture:

Stage Purpose Resolution
Stage 1 Generate low-resolution video (half target size) Half resolution
Stage 2 Upsample and refine to final resolution Full resolution

Both stages share the same prompt encoder (PromptEncoder) and image conditioner (ImageConditioner). The pipeline is defined in [packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py).

IC-LoRA Conditioning Architecture

The conditioning system processes reference videos through four key steps implemented in [iclora_utils.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py):

1. Reading LoRA Metadata

Each LoRA file contains optional scaling metadata:

  • reference_downscale_factor — spatial scaling applied during training
  • reference_temporal_scale_factor — temporal subsampling factor

These are read using read_lora_reference_downscale_factor() and read_lora_reference_temporal_scale_factor() to ensure inference matches training conditions.

2. Preparing the Reference Video

The append_ic_lora_reference_video_conditionings() function:

  • Loads reference video (MP4 or EXR directory)
  • Applies spatial resize using reference_downscale_factor
  • Applies temporal subsampling if reference_temporal_scale_factor > 1

3. Encoding to Latent Space

Reference frames pass through the VAE encoder (video_encoder). For large videos, tiled encoding is used based on tiling_config.

4. Attention Masking

Optional pixel-space masks control conditioning strength regionally:

  • Masks are downsampled to latent resolution via downsample_mask_video_to_latent()
  • Multiplied by conditioning_attention_strength
  • Scalar strengths < 1 apply global attention weighting when no mask is provided

The encoded reference is wrapped in VideoConditionByReferenceLatent and optionally ConditioningItemAttentionStrengthWrapper before being consumed by the diffusion model.

End-to-End Inference Flow

The ICLoraPipeline.__call__ method orchestrates the complete video-to-video transformation:


# Stage 1: Low-resolution generation at half resolution

stage_1_conditionings = self.image_conditioner(
    lambda enc: self._create_conditionings(
        images=images,
        video_conditioning=video_conditioning,
        height=stage_1_output_shape.height,
        width=stage_1_output_shape.width,
        video_encoder=enc,
        conditioning_attention_strength=conditioning_attention_strength,
        conditioning_attention_mask=conditioning_attention_mask,
        color_space=color_space,
    )
)

video_state, audio_state = self.stage_1(
    denoiser=SimpleDenoiser(video_context, audio_context),
    sigmas=stage_1_sigmas,
    noiser=noiser,
    width=stage_1_output_shape.width,
    height=stage_1_output_shape.height,
    frames=num_frames,
    fps=frame_rate,
    video=ModalitySpec(context=video_context, conditionings=stage_1_conditionings),
    audio=ModalitySpec(context=audio_context),
)

For Stage 2, the latent is upsampled and refined:


# Upsample Stage 1 output

upscaled_video_latent = self.upsampler(video_state.latent[:1])

# Stage 2 conditionings (typically without video conditioning)

stage_2_conditionings = self.image_conditioner(
    lambda enc: combined_image_conditionings(
        images=images,
        height=stage_2_output_shape.height,
        width=stage_2_output_shape.width,
        video_encoder=enc,
        dtype=self.dtype,
        device=self.device,
        color_space=color_space,
    )
)

# Refine at full resolution

video_state, audio_state = self.stage_2(
    denoiser=SimpleDenoiser(video_context, audio_context),
    sigmas=stage_2_sigmas,
    noiser=noiser,
    width=width,
    height=height,
    frames=num_frames,
    fps=frame_rate,
    video=ModalitySpec(
        context=video_context,
        conditionings=stage_2_conditionings,
        noise_scale=stage_2_sigmas[0].item(),
        initial_latent=upscaled_video_latent,
    ),
    audio=ModalitySpec(
        context=audio_context,
        noise_scale=stage_2_sigmas[0].item(),
        initial_latent=audio_state.latent,
    ),
)

Finally, decode to pixel space:

decoded_video = self.video_decoder(
    video_state.latent, 
    tiling_config, 
    generator, 
    dtype=vae_dtype
)

CLI Usage for Video-to-Video Transformation

Run IC-LoRA video-to-video transformation from the command line:

python -m ltx_pipelines.ic_lora \
    --prompt "A bustling cyberpunk street at night" \
    --seed 42 \
    --height 512 --width 512 \
    --num-frames 24 --frame-rate 24 \
    --images img1.png:0:1.0 img2.png:10:0.8 \
    --video-conditioning ref_video.mp4 1.0 \
    --conditioning-attention-mask mask_video.mp4 0.5 \
    --output-path ./output.mp4

Key CLI arguments:

Argument Description
--video-conditioning Reference video path and strength (e.g., ref.mp4 1.0)
--conditioning-attention-mask Optional mask video and scalar strength
--images Image conditioning tuples: path:frame_index:strength
--height/--width Target resolution (must be divisible by 64)

Python API for IC-LoRA Video-to-Video

For programmatic control:

from ltx_pipelines.ic_lora import ICLoraPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Initialize pipeline

pipeline = ICLoraPipeline(
    model_paths=ModelPaths(),
    spatial_upsampler_path="spatial_upsampler.pt",
    loras=[("ic_lora.safetensors", 1.0, None)],  # (path, strength, state_dict)

)

# Run video-to-video transformation

video, audio, metadata = pipeline(
    prompt="A serene sunrise over mountains",
    seed=123,
    height=640,
    width=640,
    num_frames=30,
    frame_rate=30,
    images=[("sky.png", 0, 1.0)],
    video_conditioning=[("ref_video.mp4", 1.0)],
    conditioning_attention_strength=0.8,
    conditioning_attention_mask=None,
)

Key Implementation Files

File Purpose
[ic_lora.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py) Main pipeline: two-stage diffusion, LoRA loading, reference conditioning
[iclora_utils.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py) Metadata reading, mask downsampling, temporal subsampling, conditioning items
[args.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) CLI argument parsing for video conditioning and mask options
[media_io.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) Video/EXR loading and frame preprocessing
[constants.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py) Sigma schedules: DISTILLED_SIGMAS, STAGE_2_DISTILLED_SIGMAS

Critical Design Considerations

Reference scaling fidelity — The pipeline automatically applies reference_downscale_factor and reference_temporal_scale_factor from LoRA metadata. This ensures the reference video matches the resolution and frame rate used during adapter training.

Memory-efficient two-stage design — Stage 1 operates at half resolution, reducing VRAM requirements. The upsampler (spatial upsampler) expands the latent before Stage 2 refinement.

Flexible attention control — Spatial masks enable regional conditioning: multiply reference influence in foreground regions while allowing free generation elsewhere.

Summary

  • IC-LoRA enables video-to-video transformation by encoding reference videos into latent conditioning tokens
  • Two-stage pipeline: half-resolution generation → upsampling → full-resolution refinement
  • Metadata-driven scaling preserves training-time resolution matching via reference_downscale_factor and reference_temporal_scale_factor
  • Attention masking provides spatial control over reference influence
  • Entry points: CLI (python -m ltx_pipelines.ic_lora) or Python API (ICLoraPipeline)

Frequently Asked Questions

What video formats are supported for reference conditioning?

The pipeline accepts MP4 files or directories of EXR frames. The append_ic_lora_reference_video_conditionings() function in iclora_utils.py handles both cases, loading frames and preprocessing them to the target color space before VAE encoding.

How does temporal subsampling work for reference videos?

When reference_temporal_scale_factor > 1 is stored in LoRA metadata, the pipeline keeps only every N-th frame from the reference video. This matches the temporal resolution used during adapter training and prevents frame rate mismatches that would degrade generation quality.

Can I skip Stage 2 for faster inference?

Yes. Set skip_stage_2=True when calling the pipeline. The output decodes directly from Stage 1 latents using self.video_decoder. This trades quality for speed, producing half-resolution results without the final upsampling and refinement pass.

What is the purpose of the conditioning attention mask?

The attention mask modulates how strongly different spatial regions of the reference video influence generation. A mask value of 0 ignores conditioning in that region; 1 applies full strength. The mask is downsampled to latent resolution via downsample_mask_video_to_latent() and combined with the scalar conditioning_attention_strength.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →