How to Apply IC-LoRA for Video-to-Video Transformations Using LTX-2

IC-LoRA for video-to-video transformations in LTX-2 requires training a LoRA with the VideoToVideoStrategy and running inference through the ICLoraPipeline with reference video conditioning.

LTX-2 implements In-Context LoRA (IC-LoRA) through a specialized two-stage architecture that enables powerful video-to-video (V2V) transformations. This guide walks through the complete workflow—from training configuration to inference—based on the official Lightricks/LTX-2 source code.

Architecture Overview

IC-LoRA in LTX-2 splits responsibilities between training and inference components:

Component Purpose Source Location
VideoToVideoStrategy Training-side batch processing, reference/target latent concatenation, and conditioning mask construction packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py
ICLoraPipeline Inference-side two-stage generation with distilled transformer packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py
append_ic_lora_reference_video_conditionings() Converts video paths into structured conditioning items with automatic down-sampling packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py
read_lora_reference_downscale_factor() / read_lora_reference_temporal_scale_factor() Extract scaling metadata from LoRA safetensors for consistent resizing packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py

Training IC-LoRA for Video-to-Video

Dataset Requirements

Each training sample must contain a reference video latent directory (reference_latents/). The dataset layout follows the standard LTX-2 structure with this additional subdirectory.

Training Flow

The VideoToVideoStrategy.prepare_training_inputs method handles five critical operations:

  1. Loads reference latents (ref_latents) alongside target latents
  2. Infers down-scale factor—the ratio of reference resolution to target resolution
  3. Builds concatenated latent tensor [ref, noisy_target]
  4. Constructs conditioning masks—reference tokens always use timestep 0; target tokens may condition on first frame with probability first_frame_conditioning_p
  5. Computes loss only on target portion—reference latents remain untouched

This design forces the model to learn style/structure transfer rather than reconstruction.

Configuration

Use the example training config at packages/ltx-trainer/configs/v2v_ic_lora.yaml. Key parameters include:

  • first_frame_conditioning_p: Probability of conditioning target on reference first frame
  • Reference latent paths resolved relative to each sample

Inference with IC-LoRA for Video-to-Video

IC-LoRA inference requires distilled model checkpoints. Standard full-model pipelines do not support this feature.

Setup and Initialization

from ltx_pipelines import ICLoraPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Initialize with distilled transformer and upsampler

model_paths = ModelPaths(
    transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
    video_vae="models/ltx-2.5/vae/video_vae.safetensors",
    audio_vae="models/ltx-2.5/vae/audio_vae.safetensors",
    spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)

# LoRA must contain reference_downscale_factor in metadata

lora_paths = [
    "models/ltx-2.5/loras/ltx-2.5-22b-ic-lora-video.safetensors"
]

pipeline = ICLoraPipeline(
    model_paths=model_paths,
    spatial_upsampler_path=model_paths.spatial_upsampler,
    loras=lora_paths,
)

The ICLoraPipeline.__init__ automatically calls read_lora_reference_downscale_factor() and read_lora_reference_temporal_scale_factor() to extract scaling metadata. Conflicting factors across multiple LoRAs raise an error.

Generation with Reference Video


# Reference video provides style/structure to transfer

video_conditioning = [
    ("assets/reference_style.mp4", 1.0)  # (path, strength)

]

frames_iter, audio, tiling_cfg = pipeline(
    prompt="A sunrise over a futuristic city, in the style of the reference video",
    seed=12345,
    height=512,
    width=768,
    num_frames=121,        # Must satisfy 8k+1 convention

    frame_rate=30.0,
    images=[],             # No image conditioning

    video_conditioning=video_conditioning,
    enhance_prompt=False,
)

generated_frames = list(frames_iter)

The pipeline executes:

  1. Stage 1: Distilled transformer denoises concatenated [ref, noisy_target] latents using VideoConditionByReferenceLatent items
  2. Stage 2: Spatial upsampler (2×) refines to final resolution

The append_ic_lora_reference_video_conditionings helper in iclora_utils.py automatically down-samples the reference video to match target dimensions using the stored reference_downscale_factor.

Spatial Masking for Selective Application

Control where IC-LoRA influence applies using attention masks:

import torch

# Binary mask: True = apply reference conditioning

mask = torch.zeros((height, width), dtype=torch.bool)
mask[100:400, 200:600] = True  # Center region only

frames_iter, audio, tiling_cfg = pipeline(
    prompt="Futuristic city with selective style transfer",
    seed=12345,
    height=512,
    width=768,
    num_frames=121,
    video_conditioning=[("assets/reference_style.mp4", 1.0)],
    conditioning_attention_mask=mask,
    conditioning_attention_strength=0.8,  # Scale mask influence

)

The downsample_mask_video_to_latent utility (also in iclora_utils.py) handles resolution matching automatically.

Key Implementation Files

Path Description
packages/ltx-trainer/src/ltx_trainer/training_strategies/video_to_video.py VideoToVideoStrategy class with prepare_training_inputs() method
packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py ICLoraPipeline two-stage inference implementation
packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py Metadata readers, mask utilities, conditioning helpers
packages/ltx-trainer/configs/v2v_ic_lora.yaml Example training configuration
packages/ltx-trainer/docs/training-modes.md IC-LoRA training mode documentation

Summary

  • IC-LoRA enables video-to-video style transfer by conditioning generation on clean reference latents
  • Training requires VideoToVideoStrategy with reference latent directories and proper metadata storage
  • Inference requires ICLoraPipeline with distilled models and LoRAs containing reference_downscale_factor
  • Reference scaling is automatic—the pipeline reads metadata and applies consistent spatial/temporal subsampling
  • Spatial control available via conditioning_attention_mask for selective region application

Frequently Asked Questions

What model type does IC-LoRA require?

IC-LoRA requires distilled transformer checkpoints only. The two-stage ICLoraPipeline is incompatible with standard full-model pipelines. Use paths like ltx-2.5-22b-distilled-transformer-bf16.safetensors.

How is the reference video resolution matched to the target?

The read_lora_reference_downscale_factor() function extracts the ratio from LoRA safetensors metadata. The reference is down-sampled to height // downscale_factor by width // downscale_factor before latent encoding.

Can I use multiple reference videos?

Yes. Pass multiple tuples to video_conditioning: [("ref1.mp4", 0.8), ("ref2.mp4", 0.5)]. Each becomes a separate VideoConditionByReferenceLatent. Ensure all referenced LoRAs share identical down-scale and temporal scale factors.

What frame count should I use for generation?

LTX-2 requires num_frames = 8k + 1 (e.g., 121, 129, 137). The temporal pipeline architecture enforces this constraint. Frame rates are flexible but 24-30 fps works best with most trained LoRAs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →