# Using IC-LoRA for Video-to-Video Transformation with Reference Images in LTX-2

> Discover how IC-LoRA in LTX-2 transforms videos using reference images. Explore two-stage diffusion for stunning video-to-video generation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: tutorial
- Published: 2026-08-15

---

**The IC-LoRA pipeline in LTX-2 enables video-to-video transformation by conditioning generation on reference videos through In-Context LoRA adapters in a two-stage distilled diffusion process.**

This guide explains how to use **IC-LoRA (In-Context LoRA)** for video-to-video transformation with reference images in the [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2) repository. The pipeline allows you to steer video generation using reference videos or image sequences, with fine-grained control over spatial and temporal conditioning.

## How the IC-LoRA Pipeline Works

The `ICLoraPipeline` implements a **two-stage distilled diffusion architecture**:

| Stage | Purpose | Resolution |
|-------|---------|------------|
| **Stage 1** | Generate low-resolution video (half target size) | Half resolution |
| **Stage 2** | Upsample and refine to final resolution | Full resolution |

Both stages share the same **prompt encoder** (`PromptEncoder`) and **image conditioner** (`ImageConditioner`). The pipeline is defined in [[`packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py).

## IC-LoRA Conditioning Architecture

The conditioning system processes reference videos through four key steps implemented in [[`iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/iclora_utils.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py):

### 1. Reading LoRA Metadata

Each LoRA file contains optional scaling metadata:

- `reference_downscale_factor` — spatial scaling applied during training
- `reference_temporal_scale_factor` — temporal subsampling factor

These are read using `read_lora_reference_downscale_factor()` and `read_lora_reference_temporal_scale_factor()` to ensure inference matches training conditions.

### 2. Preparing the Reference Video

The `append_ic_lora_reference_video_conditionings()` function:
- Loads reference video (MP4 or EXR directory)
- Applies spatial resize using `reference_downscale_factor`
- Applies temporal subsampling if `reference_temporal_scale_factor > 1`

### 3. Encoding to Latent Space

Reference frames pass through the VAE encoder (`video_encoder`). For large videos, tiled encoding is used based on `tiling_config`.

### 4. Attention Masking

Optional pixel-space masks control conditioning strength regionally:

- Masks are downsampled to latent resolution via `downsample_mask_video_to_latent()`
- Multiplied by `conditioning_attention_strength`
- Scalar strengths < 1 apply global attention weighting when no mask is provided

The encoded reference is wrapped in `VideoConditionByReferenceLatent` and optionally `ConditioningItemAttentionStrengthWrapper` before being consumed by the diffusion model.

## End-to-End Inference Flow

The `ICLoraPipeline.__call__` method orchestrates the complete video-to-video transformation:

```python

# Stage 1: Low-resolution generation at half resolution

stage_1_conditionings = self.image_conditioner(
    lambda enc: self._create_conditionings(
        images=images,
        video_conditioning=video_conditioning,
        height=stage_1_output_shape.height,
        width=stage_1_output_shape.width,
        video_encoder=enc,
        conditioning_attention_strength=conditioning_attention_strength,
        conditioning_attention_mask=conditioning_attention_mask,
        color_space=color_space,
    )
)

video_state, audio_state = self.stage_1(
    denoiser=SimpleDenoiser(video_context, audio_context),
    sigmas=stage_1_sigmas,
    noiser=noiser,
    width=stage_1_output_shape.width,
    height=stage_1_output_shape.height,
    frames=num_frames,
    fps=frame_rate,
    video=ModalitySpec(context=video_context, conditionings=stage_1_conditionings),
    audio=ModalitySpec(context=audio_context),
)

```

For **Stage 2**, the latent is upsampled and refined:

```python

# Upsample Stage 1 output

upscaled_video_latent = self.upsampler(video_state.latent[:1])

# Stage 2 conditionings (typically without video conditioning)

stage_2_conditionings = self.image_conditioner(
    lambda enc: combined_image_conditionings(
        images=images,
        height=stage_2_output_shape.height,
        width=stage_2_output_shape.width,
        video_encoder=enc,
        dtype=self.dtype,
        device=self.device,
        color_space=color_space,
    )
)

# Refine at full resolution

video_state, audio_state = self.stage_2(
    denoiser=SimpleDenoiser(video_context, audio_context),
    sigmas=stage_2_sigmas,
    noiser=noiser,
    width=width,
    height=height,
    frames=num_frames,
    fps=frame_rate,
    video=ModalitySpec(
        context=video_context,
        conditionings=stage_2_conditionings,
        noise_scale=stage_2_sigmas[0].item(),
        initial_latent=upscaled_video_latent,
    ),
    audio=ModalitySpec(
        context=audio_context,
        noise_scale=stage_2_sigmas[0].item(),
        initial_latent=audio_state.latent,
    ),
)

```

Finally, decode to pixel space:

```python
decoded_video = self.video_decoder(
    video_state.latent, 
    tiling_config, 
    generator, 
    dtype=vae_dtype
)

```

## CLI Usage for Video-to-Video Transformation

Run IC-LoRA video-to-video transformation from the command line:

```bash
python -m ltx_pipelines.ic_lora \
    --prompt "A bustling cyberpunk street at night" \
    --seed 42 \
    --height 512 --width 512 \
    --num-frames 24 --frame-rate 24 \
    --images img1.png:0:1.0 img2.png:10:0.8 \
    --video-conditioning ref_video.mp4 1.0 \
    --conditioning-attention-mask mask_video.mp4 0.5 \
    --output-path ./output.mp4

```

**Key CLI arguments:**

| Argument | Description |
|----------|-------------|
| `--video-conditioning` | Reference video path and strength (e.g., `ref.mp4 1.0`) |
| `--conditioning-attention-mask` | Optional mask video and scalar strength |
| `--images` | Image conditioning tuples: `path:frame_index:strength` |
| `--height/--width` | Target resolution (must be divisible by 64) |

## Python API for IC-LoRA Video-to-Video

For programmatic control:

```python
from ltx_pipelines.ic_lora import ICLoraPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Initialize pipeline

pipeline = ICLoraPipeline(
    model_paths=ModelPaths(),
    spatial_upsampler_path="spatial_upsampler.pt",
    loras=[("ic_lora.safetensors", 1.0, None)],  # (path, strength, state_dict)

)

# Run video-to-video transformation

video, audio, metadata = pipeline(
    prompt="A serene sunrise over mountains",
    seed=123,
    height=640,
    width=640,
    num_frames=30,
    frame_rate=30,
    images=[("sky.png", 0, 1.0)],
    video_conditioning=[("ref_video.mp4", 1.0)],
    conditioning_attention_strength=0.8,
    conditioning_attention_mask=None,
)

```

## Key Implementation Files

| File | Purpose |
|------|---------|
| [[`ic_lora.py`](https://github.com/Lightricks/LTX-2/blob/main/ic_lora.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ic_lora.py) | Main pipeline: two-stage diffusion, LoRA loading, reference conditioning |
| [[`iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/iclora_utils.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/iclora_utils.py) | Metadata reading, mask downsampling, temporal subsampling, conditioning items |
| [[`args.py`](https://github.com/Lightricks/LTX-2/blob/main/args.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) | CLI argument parsing for video conditioning and mask options |
| [[`media_io.py`](https://github.com/Lightricks/LTX-2/blob/main/media_io.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) | Video/EXR loading and frame preprocessing |
| [[`constants.py`](https://github.com/Lightricks/LTX-2/blob/main/constants.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py) | Sigma schedules: `DISTILLED_SIGMAS`, `STAGE_2_DISTILLED_SIGMAS` |

## Critical Design Considerations

**Reference scaling fidelity** — The pipeline automatically applies `reference_downscale_factor` and `reference_temporal_scale_factor` from LoRA metadata. This ensures the reference video matches the resolution and frame rate used during adapter training.

**Memory-efficient two-stage design** — Stage 1 operates at half resolution, reducing VRAM requirements. The `upsampler` (spatial upsampler) expands the latent before Stage 2 refinement.

**Flexible attention control** — Spatial masks enable regional conditioning: multiply reference influence in foreground regions while allowing free generation elsewhere.

## Summary

- **IC-LoRA enables video-to-video transformation** by encoding reference videos into latent conditioning tokens
- **Two-stage pipeline**: half-resolution generation → upsampling → full-resolution refinement
- **Metadata-driven scaling** preserves training-time resolution matching via `reference_downscale_factor` and `reference_temporal_scale_factor`
- **Attention masking** provides spatial control over reference influence
- **Entry points**: CLI (`python -m ltx_pipelines.ic_lora`) or Python API (`ICLoraPipeline`)

## Frequently Asked Questions

### What video formats are supported for reference conditioning?

The pipeline accepts **MP4 files** or **directories of EXR frames**. The `append_ic_lora_reference_video_conditionings()` function in [`iclora_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/iclora_utils.py) handles both cases, loading frames and preprocessing them to the target color space before VAE encoding.

### How does temporal subsampling work for reference videos?

When `reference_temporal_scale_factor > 1` is stored in LoRA metadata, the pipeline keeps only every N-th frame from the reference video. This matches the temporal resolution used during adapter training and prevents frame rate mismatches that would degrade generation quality.

### Can I skip Stage 2 for faster inference?

Yes. Set `skip_stage_2=True` when calling the pipeline. The output decodes directly from Stage 1 latents using `self.video_decoder`. This trades quality for speed, producing half-resolution results without the final upsampling and refinement pass.

### What is the purpose of the conditioning attention mask?

The attention mask modulates how strongly different spatial regions of the reference video influence generation. A mask value of 0 ignores conditioning in that region; 1 applies full strength. The mask is downsampled to latent resolution via `downsample_mask_video_to_latent()` and combined with the scalar `conditioning_attention_strength`.