# How to Use Image-Conditioned LoRA (IC-LoRA) for Video Transformations in LTX-2

> Learn to use image-conditioned LoRA IC LoRA in LTX-2 for video transformations. This guide details conditioning generation on reference videos with spatial scaling and temporal subsampling.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-06-20

---

**LTX-2 provides a two-stage diffusion pipeline that uses Image-Conditioned LoRA (IC-LoRA) to transform videos by conditioning generation on reference videos, with automatic handling of spatial scaling, temporal subsampling, and optional attention masks.**

LTX-2 by Lightricks introduces **Image-Conditioned LoRA (IC-LoRA)**, a powerful mechanism for guiding video generation through reference videos and visual signals such as depth maps or pose sequences. This article explores how to leverage the `ICLoraPipeline` class and associated utilities in [`ltx_pipelines/ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/ic_lora.py) to perform video transformations using LoRA adapters with video conditioning.

## The Two-Stage IC-LoRA Architecture

The `ICLoraPipeline` orchestrates a two-stage diffusion process that balances quality and computational efficiency. Understanding this architecture is essential for configuring your video transformations correctly.

### Stage 1: Low-Resolution Generation with LoRA

Stage 1 generates a low-resolution video at half the target resolution (e.g., 256×256 for a 512×512 output) using a distilled diffusion model. This stage applies your **LoRA adapters** to condition the generation on reference videos.

In [`ltx_pipelines/ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/ic_lora.py), the pipeline initializes Stage 1 by passing LoRA weights to the `DiffusionStage` class:

```python
self.stage_1 = DiffusionStage(
    # ... other parameters ...

    loras=tuple(loras)
)

```

The LoRA files may embed `reference_downscale_factor` and `reference_temporal_scale_factor` metadata in their safetensors headers. The pipeline automatically reads these values via `read_lora_reference_downscale_factor` and `read_lora_reference_temporal_scale_factor` in [`ltx_pipelines/iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/iclora_utils.py) to resize and subsample reference videos accordingly. If you load multiple LoRAs with conflicting scale factors, the pipeline raises a clear error.

### Stage 2: Upscaling and Refinement

Stage 2 upscales the Stage 1 output to the target resolution and refines it using another distilled model. Critically, this stage runs **without LoRA adapters** (pure distilled model) to ensure high-quality upsampling:

```python
self.stage_2 = DiffusionStage(
    # ... other parameters ...

    loras=()  # Empty tuple - no LoRA applied

)

```

You can skip Stage 2 using `--skip-stage-2` for rapid prototyping or memory-constrained environments, though this yields half-resolution output.

## Preparing Video Conditionings

The `append_ic_lora_reference_video_conditionings` function in [`ltx_pipelines/iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/iclora_utils.py) handles the conversion of reference videos into latent conditioning signals.

### Reference Video Scaling

When you provide a reference video path, the system creates a `VideoConditionByReferenceLatent` object that undergoes automatic spatial downscaling and temporal subsampling based on the LoRA metadata. The utility functions read these parameters directly from the safetensors file:

- `read_lora_reference_downscale_factor`: Determines spatial resize ratios
- `read_lora_reference_temporal_scale_factor`: Controls frame subsampling rates

### Spatial Attention Masks

For spatially varying influence, you can supply a mask video via `--conditioning-attention-mask`. The `downsample_mask_video_to_latent` function processes this mask by:

1. Loading the video via `_load_mask_video` in [`ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/ic_lora.py)
2. Normalizing values to `[0, 1]`
3. Downsampling to latent token dimensions
4. Multiplying by `conditioning_attention_strength`

The resulting mask modulates the reference video's influence per pixel, where 0 ignores the reference and 1 applies full conditioning.

## Command-Line Usage

The CLI implementation in [`ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/ic_lora.py) provides direct access to the pipeline through the `main()` function, which parses arguments including `--video-conditioning`, `--conditioning-attention-mask`, and `--skip-stage-2`.

### Basic Video Transformation

```bash
python -m ltx_pipelines.ic_lora \
    --distilled-checkpoint-path models/distilled.pt \
    --spatial-upsampler-path models/upsampler.pt \
    --gemma-root /path/to/gemma \
    --lora path/to/ic_lora.safetensors 1.0 \
    --video-conditioning path/to/reference.mp4 0.8 \
    --prompt "A dancer performing in rain" \
    --seed 42 \
    --height 512 \
    --width 512 \
    --num-frames 24 \
    --frame-rate 24 \
    --output-path results/dance.mp4

```

### With Spatial Masking

```bash
python -m ltx_pipelines.ic_lora \
    --distilled-checkpoint-path models/distilled.pt \
    --spatial-upsampler-path models/upsampler.pt \
    --gemma-root /path/to/gemma \
    --lora path/to/ic_lora.safetensors 1.0 \
    --video-conditioning path/to/reference.mp4 0.8 \
    --conditioning-attention-mask path/to/mask.mp4 0.5 \
    --prompt "A dancer performing in rain" \
    --height 512 \
    --width 512 \
    --num-frames 24 \
    --output-path results/dance_masked.mp4

```

### Quick Preview (Skip Stage 2)

```bash
python -m ltx_pipelines.ic_lora \
    --distilled-checkpoint-path models/distilled.pt \
    --spatial-upsampler-path models/upsampler.pt \
    --gemma-root /path/to/gemma \
    --lora path/to/ic_lora.safetensors 1.0 \
    --video-conditioning path/to/reference.mp4 1.0 \
    --skip-stage-2 \
    --output-path quick_preview.mp4

```

## Python API Implementation

For programmatic control, import `ICLoraPipeline` from `ltx_pipelines.ic_lora` and configure it with `LoraPathStrengthAndSDOps` objects.

### Basic Pipeline Setup

```python
from ltx_pipelines.ic_lora import ICLoraPipeline
from ltx_core.loader import LoraPathStrengthAndSDOps
from pathlib import Path

# Configure LoRA adapter

lora = LoraPathStrengthAndSDOps(
    path=Path("models/ic_lora.safetensors"),
    strength=1.0,
    # The ops field is populated automatically by the loader

)

# Initialize pipeline

pipeline = ICLoraPipeline(
    distilled_checkpoint_path="models/distilled.pt",
    spatial_upsampler_path="models/upsampler.pt",
    gemma_root="/path/to/gemma",
    loras=[lora],
)

# Generate video

video, audio = pipeline(
    prompt="A futuristic city skyline at sunset",
    seed=123,
    height=512,
    width=512,
    num_frames=30,
    frame_rate=30,
    images=[],  # No image conditioning

    video_conditioning=[("data/ref_video.mp4", 0.9)],
    conditioning_attention_strength=0.8,
    skip_stage_2=False,
)

# Encode and save

from ltx_pipelines.utils.media_io import encode_video
encode_video(video, fps=30, audio=audio, output_path="outputs/result.mp4")

```

### Adding Spatial Masks Manually

To apply spatially varying attention weights, load a mask video using the internal `_load_mask_video` helper:

```python
from ltx_pipelines.ic_lora import _load_mask_video

# Load mask at Stage 1 resolution (half target)

mask_tensor = _load_mask_video(
    mask_path="data/mask.mp4",
    height=256,  # 512 // 2

    width=256,
    num_frames=24,
)

# Generate with mask

video, audio = pipeline(
    prompt="A cat dancing on a beach",
    seed=777,
    height=512,
    width=512,
    num_frames=24,
    frame_rate=24,
    images=[],
    video_conditioning=[("data/cat_ref.mp4", 1.0)],
    conditioning_attention_strength=0.7,
    conditioning_attention_mask=mask_tensor,
)

```

The mask tensor is normalized to `[0, 1]`, downsampled to latent space via `downsample_mask_video_to_latent`, and multiplied by the `conditioning_attention_strength` scalar.

### Multiple LoRA Adapters

You can pass multiple LoRA files to combine different conditioning signals:

```python
lora1 = LoraPathStrengthAndSDOps(path=Path("models/motion_lora.safetensors"), strength=0.8)
lora2 = LoraPathStrengthAndSDOps(path=Path("models/style_lora.safetensors"), strength=0.6)

pipeline = ICLoraPipeline(
    # ... paths ...

    loras=[lora1, lora2],
)

```

Ensure that all LoRAs share compatible `reference_downscale_factor` and `reference_temporal_scale_factor` metadata, as conflicting values will raise an error during initialization.

## Summary

- **IC-LoRA** enables video generation conditioned on reference videos through a specialized two-stage pipeline in [`ltx_pipelines/ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/ic_lora.py).
- **Stage 1** generates low-resolution output at half target resolution using LoRA adapters and automatic scaling via metadata in [`iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/iclora_utils.py).
- **Stage 2** upscales results using a distilled model without LoRA weights; skip this stage with `--skip-stage-2` for faster iteration.
- **Video conditioning** supports optional spatial masks via `downsample_mask_video_to_latent` for per-pixel control over reference influence.
- **Metadata automation** eliminates manual calibration by reading `reference_downscale_factor` and `reference_temporal_scale_factor` directly from safetensors files.

## Frequently Asked Questions

### What is the difference between Stage 1 and Stage 2 in IC-LoRA?

Stage 1 generates a low-resolution video (half the target dimensions) using the distilled model with your LoRA adapters attached, processing reference videos according to their embedded metadata. Stage 2 upscales this result to full resolution using a separate distilled model without any LoRA weights, providing quality refinement. Stage 2 can be skipped using `--skip-stage-2` to produce half-resolution outputs quickly.

### How does LTX-2 handle resolution differences between reference and target videos?

The pipeline automatically reads `reference_downscale_factor` and `reference_temporal_scale_factor` from each LoRA's safetensors metadata via `read_lora_reference_downscale_factor` and `read_lora_reference_temporal_scale_factor` in [`iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/iclora_utils.py). These values determine how the reference video is spatially resized and temporally subsampled to match the generation process. If multiple LoRAs specify conflicting scale factors, the pipeline raises an error requesting compatible adapters.

### Can I use multiple LoRA adapters simultaneously with IC-LoRA?

Yes, you can pass multiple `--lora` arguments via CLI or a list of `LoraPathStrengthAndSDOps` objects via the Python API. However, all LoRAs must share identical `reference_downscale_factor` and `reference_temporal_scale_factor` metadata, as the pipeline applies a single scaling configuration per generation. Conflicting metadata triggers a validation error to prevent inconsistent conditioning.

### How do I create a spatial attention mask for video conditioning?

Supply a grayscale video mask via `--conditioning-attention-mask` (CLI) or the `conditioning_attention_mask` parameter (Python API). The mask is loaded via `_load_mask_video`, normalized to `[0, 1]`, and downsampled to latent space using `downsample_mask_video_to_latent`. Values of 0 completely block the reference video's influence at that pixel, while 1 applies full strength multiplied by your `conditioning_attention_strength` parameter.