# How the Keyframe Interpolation Pipeline Works in LTX-2: A Two-Stage Diffusion Approach

> Discover how the LTX-2 keyframe interpolation pipeline uses a two-stage diffusion approach to create smooth, high-quality videos by generating a coarse draft and then refining it.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-06-20

---

**The KeyframeInterpolationPipeline in Lightricks/LTX-2 implements a two-stage diffusion workflow that generates high-quality video by first creating a low-resolution coarse draft at half resolution and then refining it with a distilled LoRA upsampler to produce smooth, temporally consistent output.**

The keyframe interpolation pipeline in Lightricks/LTX-2 enables users to generate high-fidelity video sequences by interpolating between provided keyframes using a sophisticated diffusion-based architecture. Located in [`packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py), this pipeline orchestrates multiple modular components to transform text prompts and visual references into coherent video and audio outputs.

## Pipeline Architecture and Component Initialization

The `KeyframeInterpolationPipeline` class is constructed by wiring together specialized components that enable two-stage generation. In the `__init__` method (lines 46-104), the pipeline initializes a **prompt encoder**, an **image conditioner**, two **DiffusionStage** objects for sequential processing, a **video upsampler**, and separate decoders for video and audio modalities.

Notably, Stage 2 receives both regular LoRA weights and **distilled LoRA weights**, allowing the refinement phase to efficiently utilize compressed knowledge for high-resolution output without recomputing the full diffusion process from scratch.

## The Two-Stage Generation Process

The `__call__` method (starting at line 117) executes the generation workflow through a deterministic sequence of validation, encoding, denoising, and decoding phases.

### Input Validation and Seed Configuration

The pipeline first validates compatibility between the requested resolution and the two-stage architecture using `assert_resolution` (line 123). It then initializes a `torch.Generator` seeded with the user-provided seed and wraps it in a `GaussianNoiser` (lines 125-127) to ensure reproducible noise sampling throughout the diffusion steps.

### Prompt Encoding and Multimodal Conditioning

The `PromptEncoder` produces text-conditioned embeddings for both positive and negative prompts, generating separate video and audio context vectors (lines 129-136). These embeddings guide the diffusion process toward the semantic content described in the text prompts while avoiding unwanted artifacts specified in the negative prompts.

### Stage 1: Low-Resolution Coarse Generation

Stage 1 operates at **half the target resolution** (e.g., 360p for a 720p output) to establish the initial temporal structure:

1. **Sigma Schedule Construction**: The `LTX2Scheduler` (defined in [`packages/ltx-core/src/ltx_core/components/schedulers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/schedulers.py)) builds a noise schedule governing the diffusion timesteps.
2. **Image Conditioning**: Keyframe images are transformed into guidance latents via `image_conditionings_by_adding_guiding_latent`, enriching the conditioning signals to steer generation toward the supplied visual references.
3. **Guider Factory Instantiation**: The `create_multimodal_guider_factory` (from [`packages/ltx-core/src/ltx_core/components/guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/guiders.py)) creates video and audio guiders with configurable CFG scales.
4. **Denoising Execution**: A `FactoryGuidedDenoiser` processes the noisy latent through the first `DiffusionStage` using the computed `stage_1_output_shape` (lines 143-149) and `stage_1_conditionings` (lines 150-158), producing a coarse latent representation (lines 170-191).

### Upsampling and Stage 2: High-Resolution Refinement

Following Stage 1, the pipeline upscales the coarse latent by a factor of 2 using the `VideoUpsampler` (line 194):

```python
upscaled_video_latent = self.upsampler(video_state.latent[:1])

```

Stage 2 leverages **distilled LoRA weights** alongside the original LoRAs to refine the upsampled latent efficiently:

- A new `stage_2_sigmas` schedule drives the refinement (line 196)
- `stage_2_conditionings` are computed for the higher resolution (lines 198-207)
- A simpler denoiser refines the visual details without the full computational cost of Stage 1 (lines 209-228)
- Audio latents are simultaneously refined using the first sigma value as the noise scale

### Decoding and Output Generation

The final decoding phase converts latents to pixel-space outputs:

- The video latent is decoded by the `VideoDecoder` (optionally using tiled decoding for memory efficiency), located at line 230
- The audio latent is processed by the `AudioDecoder` (line 231)
- The method returns an iterator yielding video frames and the decoded audio waveform (line 232)

## Programmatic Usage and CLI

You can instantiate the pipeline programmatically to integrate keyframe interpolation into custom workflows:

```python
from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_core.loader import LoraPathStrengthAndSDOps
from ltx_core.types import MultiModalGuiderParams
import torch

# Initialize pipeline with checkpoint and LoRA configurations

pipeline = KeyframeInterpolationPipeline(
    checkpoint_path="/path/to/checkpoint",
    distilled_lora=[
        LoraPathStrengthAndSDOps(
            path="/path/to/distilled_lora.pt", 
            strength=0.7, 
            sd_ops=[]
        )
    ],
    spatial_upsampler_path="/path/to/spatial_upsampler.pt",
    gemma_root="/path/to/gemma",
    loras=[],  # Optional additional LoRAs

)

# Prepare keyframe images as list of (tensor, strength) tuples

images = []  # Populate with actual torch tensors

# Execute generation

video_iter, audio = pipeline(
    prompt="A sunrise over a mountain lake",
    negative_prompt="low quality",
    seed=42,
    height=512,
    width=512,
    num_frames=24,
    frame_rate=24.0,
    num_inference_steps=30,
    video_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
    audio_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
    images=images,
)

```

Alternatively, use the built-in CLI entry point (lines 335-393) for end-to-end generation:

```bash
python -m ltx_pipelines.keyframe_interpolation \
    --checkpoint_path /path/to/checkpoint \
    --distilled_lora /path/to/distilled_lora.pt \
    --spatial_upsampler_path /path/to/spatial_upsampler.pt \
    --gemma_root /path/to/gemma \
    --prompt "A futuristic cityscape at night" \
    --height 720 --width 1280 --num_frames 60 --frame_rate 30 \
    --output_path result.mp4

```

## Summary

- The **KeyframeInterpolationPipeline** employs a **two-stage diffusion** architecture: Stage 1 generates a coarse half-resolution draft, while Stage 2 refines the 2× upsampled output using distilled LoRA weights.
- **Modality-specific conditioning** via `image_conditionings_by_adding_guiding_latent` and multimodal guider factories ensures precise alignment between keyframes, text prompts, and generated output.
- **Memory-efficient decoding** supports tiled video decoding to handle high-resolution outputs without excessive GPU memory requirements.
- The pipeline is fully modular, with components defined in [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) and core utilities in `packages/ltx-core`, enabling reuse across different video generation tasks.

## Frequently Asked Questions

### What is the purpose of the two-stage design in the LTX-2 keyframe interpolation pipeline?

The two-stage design optimizes the quality-efficiency trade-off by first generating a temporally coherent coarse video at half resolution using the full diffusion model, then applying a computationally efficient refinement stage with distilled LoRAs to upscale and enhance details. This approach avoids the prohibitive cost of running full diffusion at the target resolution while maintaining high-fidelity output.

### How does the pipeline handle memory efficiency during video decoding?

The `VideoDecoder` supports tiled decoding strategies that process the video latent in spatial chunks rather than loading the entire high-resolution frame into memory simultaneously. This tiling mechanism, referenced in the decoding call at line 230, enables the generation of high-resolution videos (e.g., 720p or 1080p) on consumer-grade hardware without out-of-memory errors.

### What role do distilled LoRA weights play in Stage 2 of the interpolation process?

Distilled LoRA weights provide a compressed, efficient representation of the full model's knowledge for the refinement task. In Stage 2 (lines 196-228), these weights enable rapid denoising of the upsampled latent without reloading the full pretrained weights, significantly reducing computational overhead while preserving the visual quality established in Stage 1.

### Can the keyframe interpolation pipeline generate audio alongside video?

Yes, the pipeline is architecturally multimodal. The `PromptEncoder` generates separate audio embeddings (lines 129-136), and the `AudioDecoder` (line 231) processes the refined audio latent into a waveform. The pipeline returns both a video frame iterator and the decoded audio array, synchronized through the shared conditioning and generation process.