# LTX-2 KeyframeInterpolationPipeline: Generate Smooth Video from Image Keyframes

> Generate smooth video from image keyframes with LTX-2 KeyframeInterpolationPipeline. Discover its two-stage diffusion workflow for temporally coherent full-resolution video.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: tutorial
- Published: 2026-08-15

---

**The LTX-2 KeyframeInterpolationPipeline uses a two-stage diffusion workflow with additive guiding latents to interpolate between user-supplied image keyframes, producing temporally coherent full-resolution video.** This pipeline lives in `packages/ltx-pipelines` and implements a unique conditioning strategy that preserves underlying diffusion dynamics while steering output toward specified keyframes.

## How KeyframeInterpolationPipeline Works in LTX-2

The **KeyframeInterpolationPipeline** follows the same architectural pattern as LTX-2's other two-stage pipelines (like `TI2VidTwoStagesPipeline`), but replaces latent replacement with **additive guiding latents**. This approach enables smooth transitions between keyframes without disrupting the diffusion process.

### Stage 1: Low-Resolution Generation with Guiding Latents

The pipeline begins by encoding prompts and preparing image conditioning:

1. **Prompt encoding** — `PromptEncoder` (from [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py)) converts positive/negative text prompts into video and audio conditioning contexts: `v_context_p`, `a_context_p`, `v_context_n`, `a_context_n`.

2. **Image conditioning** — The `ImageConditioner` loads the VAE checkpoint and creates conditioning latents by **adding** a guiding latent derived from each keyframe image to the latent being denoised. This happens in `image_conditionings_by_adding_guiding_latent` within [`packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py).

3. **Diffusion pass** — A `DiffusionStage` built from the full transformer checkpoint runs at **half target resolution**. The `FactoryGuidedDenoiser` receives video/audio contexts plus multimodal guider factories from `create_multimodal_guider_factory`.

The result is a low-resolution latent video (`video_state.latent`) that already follows the keyframe trajectory.

### Stage 2: Upsampling and Refinement

4. **Spatial upsampling** — `VideoUpsampler` (also in [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py)) lifts the latent from [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) to full resolution (×2) using the pretrained spatial upsampler checkpoint. Temporal continuity is preserved through direct latent feed-forward.

5. **Refinement with distilled LoRA** — The second `DiffusionStage` loads the same transformer checkpoint but applies **distilled LoRA weights**. It uses `SimpleDenoiser` because the upscaled video requires only minimal refinement steps (`stage_2_sigmas`). The initial latent combines upscaled video with the same guiding latents from keyframes.

6. **Decoding and output** — `VideoDecoder` and `AudioDecoder` convert final latents to pixels and waveforms, respecting `AUTO_TILING` configuration for memory efficiency. `encode_video` writes the result to disk.

All heavy components (`PromptEncoder`, `ImageConditioner`, `DiffusionStage`, `VideoUpsampler`, `VideoDecoder`, `AudioDecoder`) are **self-contained blocks** that allocate, run, and immediately free GPU memory—critical for preventing OOM errors in multi-stage workflows.

## Running KeyframeInterpolationPipeline from CLI

The pipeline exposes a full command-line interface through [`keyframe_interpolation.py`](https://github.com/Lightricks/LTX-2/blob/main/keyframe_interpolation.py):

```bash
uv run python -m ltx_pipelines.keyframe_interpolation \
    --transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors \
    --text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
    --video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
    --audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
    --spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
    --distilled-lora models/ltx-2.5/loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors \
    --prompt "A sunrise over a quiet lake" \
    --negative-prompt "" \
    --seed 1234 \
    --height 720 \
    --width 1280 \
    --num-frames 121 \
    --frame-rate 30 \
    --num-inference-steps 30 \
    --images path/to/keyframe1.png path/to/keyframe2.png path/to/keyframe3.png \
    --output-path output.mp4

```

Key parameters specific to keyframe interpolation:
- `--images` — List of keyframe image paths (minimum 2 for meaningful interpolation)
- `--distilled-lora` — Required for Stage-2 refinement speedup
- `--num-frames` — Total output length; intermediate frames are generated between keyframes

## Using KeyframeInterpolationPipeline in Python

For programmatic control, instantiate `KeyframeInterpolationPipeline` directly:

```python
from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.args import ImageConditioningInput
from PIL import Image

# Configure model paths (download from HuggingFace first)

paths = ModelPaths(
    transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors",
    text_encoder="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
    video_vae="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
    audio_vae="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
    spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)

pipeline = KeyframeInterpolationPipeline(
    model_paths=paths,
    distilled_lora=[("models/ltx-2.5/loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors", 1.0, None)],
    spatial_upsampler_path=paths.spatial_upsampler,
    loras=[],  # Additional LoRAs can be loaded here

)

# Prepare keyframe inputs

keyframes = [
    ImageConditioningInput(image=Image.open("keyframe1.png")),
    ImageConditioningInput(image=Image.open("keyframe2.png")),
    ImageConditioningInput(image=Image.open("keyframe3.png")),
]

# Generate video

video, audio, metadata = pipeline(
    prompt="A sunrise over a quiet lake",
    negative_prompt="",
    seed=1234,
    height=720,
    width=1280,
    num_frames=121,
    frame_rate=30.0,
    num_inference_steps=30,
    video_guider_params=...,   # MultiModalGuiderParams for CFG/STG guidance

    audio_guider_params=...,   # Audio guidance configuration

    images=keyframes,
)

# video: iterator of torch.Tensor frames

# audio: raw waveform array

# Export with encode_video from ltx_pipelines.utils.media_io

```

## Core Architecture Concepts

| Concept | Implementation | Location |
|---------|---------------|----------|
| **Guiding latent** | Added (not replaced) to denoising latent; derived from keyframe VAE encoding | [`helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/helpers.py): `image_conditionings_by_adding_guiding_latent` |
| **Multimodal guidance** | Video and audio guidance factories providing CFG/STG/rescale guidance | Passed to `FactoryGuidedDenoiser` |
| **Distilled LoRA** | Lightweight adapter (450 steps) applied in Stage-2 for fast refinement | `KeyframeInterpolationPipeline` constructor |
| **Memory-safe blocks** | Self-contained components that free GPU memory after each stage | [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py): all block classes |
| **Adaptive tiling** | `AUTO_TILING` selects tile sizes based on VAE checkpoint specifications | [`helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/helpers.py) tiling utilities |

## Key Source Files in LTX-2

- [`packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py) — Main `KeyframeInterpolationPipeline` class and CLI entry point
- [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) — Reusable pipeline blocks (`PromptEncoder`, `ImageConditioner`, `DiffusionStage`, `VideoUpsampler`, `VideoDecoder`, `AudioDecoder`)
- [`packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) — Guiding latent creation, tiling logic, memory cleanup
- [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) — `ImageConditioningInput` and argument parsing
- [`packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) — Video encoding, HDR handling, VAE dtype selection

## Summary

- **KeyframeInterpolationPipeline** generates smooth video from 2+ image keyframes using a **two-stage diffusion process** with **additive guiding latents**
- **Stage 1** creates low-resolution video at half resolution using `FactoryGuidedDenoiser` with multimodal guidance
- **Stage 2** upsamples via `VideoUpsampler` and refines with distilled LoRA weights via `SimpleDenoiser`
- **Additive conditioning** (in [`helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/helpers.py)) preserves diffusion dynamics while steering toward keyframes—distinct from replacement strategies in other pipelines
- All components are **memory-safe blocks** that prevent OOM errors during multi-stage generation
- Available via **CLI** (`python -m ltx_pipelines.keyframe_interpolation`) or **Python API** (`KeyframeInterpolationPipeline` class)

## Frequently Asked Questions

### What makes KeyframeInterpolationPipeline different from image-to-video pipelines?

**KeyframeInterpolationPipeline specifically handles multiple image inputs with temporal spacing**, using additive guiding latents to interpolate smooth transitions between them. Standard image-to-video pipelines (like `TI2VidTwoStagesPipeline`) typically use latent replacement rather than addition, and don't optimize for multi-keyframe trajectories. The additive approach in `image_conditionings_by_adding_guiding_latent` preserves more of the base diffusion model's generation quality while still respecting keyframe constraints.

### Why does the pipeline require distilled LoRA weights?

**The distilled LoRA enables fast Stage-2 refinement** without quality degradation. According to the LTX-2 source, Stage 2 uses `SimpleDenoiser` with few steps (`stage_2_sigmas`) because the upsampled video from Stage 1 already looks good—the LoRA provides the necessary adaptation for this lightweight refinement mode. The checkpoint `ltx-2.5-22b-distilled-lora-450-bf16.safetensors` was specifically trained for this purpose.

### How does the pipeline handle GPU memory with long videos?

**Memory management relies on self-contained block allocation and `AUTO_TILING`**. Each heavy component (`PromptEncoder`, `DiffusionStage`, `VideoDecoder`, etc.) allocates GPU memory, runs its computation, and immediately frees it before returning. Additionally, the VAE decoder uses adaptive tiling based on the checkpoint's specifications, splitting large frames into manageable tiles without manual configuration.

### Can I use more than three keyframes, and how are they temporally distributed?

**Yes—provide any number of keyframes via the `--images` argument or `images` parameter**. The pipeline distributes them evenly across `num_frames` and interpolates between consecutive keyframes. The guiding latents are computed per-frame in `image_conditionings_by_adding_guiding_latent`, ensuring each output frame receives appropriate conditioning based on its temporal proximity to the nearest keyframes.