LTX-2 KeyframeInterpolationPipeline: Generate Smooth Video from Image Keyframes

The LTX-2 KeyframeInterpolationPipeline uses a two-stage diffusion workflow with additive guiding latents to interpolate between user-supplied image keyframes, producing temporally coherent full-resolution video. This pipeline lives in packages/ltx-pipelines and implements a unique conditioning strategy that preserves underlying diffusion dynamics while steering output toward specified keyframes.

How KeyframeInterpolationPipeline Works in LTX-2

The KeyframeInterpolationPipeline follows the same architectural pattern as LTX-2's other two-stage pipelines (like TI2VidTwoStagesPipeline), but replaces latent replacement with additive guiding latents. This approach enables smooth transitions between keyframes without disrupting the diffusion process.

Stage 1: Low-Resolution Generation with Guiding Latents

The pipeline begins by encoding prompts and preparing image conditioning:

  1. Prompt encoding — PromptEncoder (from packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) converts positive/negative text prompts into video and audio conditioning contexts: v_context_p, a_context_p, v_context_n, a_context_n.

  2. Image conditioning — The ImageConditioner loads the VAE checkpoint and creates conditioning latents by adding a guiding latent derived from each keyframe image to the latent being denoised. This happens in image_conditionings_by_adding_guiding_latent within packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py.

  3. Diffusion pass — A DiffusionStage built from the full transformer checkpoint runs at half target resolution. The FactoryGuidedDenoiser receives video/audio contexts plus multimodal guider factories from create_multimodal_guider_factory.

The result is a low-resolution latent video (video_state.latent) that already follows the keyframe trajectory.

Stage 2: Upsampling and Refinement

  1. Spatial upsampling — VideoUpsampler (also in blocks.py) lifts the latent from packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py to full resolution (×2) using the pretrained spatial upsampler checkpoint. Temporal continuity is preserved through direct latent feed-forward.

  2. Refinement with distilled LoRA — The second DiffusionStage loads the same transformer checkpoint but applies distilled LoRA weights. It uses SimpleDenoiser because the upscaled video requires only minimal refinement steps (stage_2_sigmas). The initial latent combines upscaled video with the same guiding latents from keyframes.

  3. Decoding and output — VideoDecoder and AudioDecoder convert final latents to pixels and waveforms, respecting AUTO_TILING configuration for memory efficiency. encode_video writes the result to disk.

All heavy components (PromptEncoder, ImageConditioner, DiffusionStage, VideoUpsampler, VideoDecoder, AudioDecoder) are self-contained blocks that allocate, run, and immediately free GPU memory—critical for preventing OOM errors in multi-stage workflows.

Running KeyframeInterpolationPipeline from CLI

The pipeline exposes a full command-line interface through keyframe_interpolation.py:

uv run python -m ltx_pipelines.keyframe_interpolation \
    --transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors \
    --text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
    --video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
    --audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
    --spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
    --distilled-lora models/ltx-2.5/loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors \
    --prompt "A sunrise over a quiet lake" \
    --negative-prompt "" \
    --seed 1234 \
    --height 720 \
    --width 1280 \
    --num-frames 121 \
    --frame-rate 30 \
    --num-inference-steps 30 \
    --images path/to/keyframe1.png path/to/keyframe2.png path/to/keyframe3.png \
    --output-path output.mp4

Key parameters specific to keyframe interpolation:

  • --images — List of keyframe image paths (minimum 2 for meaningful interpolation)
  • --distilled-lora — Required for Stage-2 refinement speedup
  • --num-frames — Total output length; intermediate frames are generated between keyframes

Using KeyframeInterpolationPipeline in Python

For programmatic control, instantiate KeyframeInterpolationPipeline directly:

from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.args import ImageConditioningInput
from PIL import Image

# Configure model paths (download from HuggingFace first)

paths = ModelPaths(
    transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors",
    text_encoder="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
    video_vae="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
    audio_vae="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
    spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)

pipeline = KeyframeInterpolationPipeline(
    model_paths=paths,
    distilled_lora=[("models/ltx-2.5/loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors", 1.0, None)],
    spatial_upsampler_path=paths.spatial_upsampler,
    loras=[],  # Additional LoRAs can be loaded here

)

# Prepare keyframe inputs

keyframes = [
    ImageConditioningInput(image=Image.open("keyframe1.png")),
    ImageConditioningInput(image=Image.open("keyframe2.png")),
    ImageConditioningInput(image=Image.open("keyframe3.png")),
]

# Generate video

video, audio, metadata = pipeline(
    prompt="A sunrise over a quiet lake",
    negative_prompt="",
    seed=1234,
    height=720,
    width=1280,
    num_frames=121,
    frame_rate=30.0,
    num_inference_steps=30,
    video_guider_params=...,   # MultiModalGuiderParams for CFG/STG guidance

    audio_guider_params=...,   # Audio guidance configuration

    images=keyframes,
)

# video: iterator of torch.Tensor frames

# audio: raw waveform array

# Export with encode_video from ltx_pipelines.utils.media_io

Core Architecture Concepts

Concept Implementation Location
Guiding latent Added (not replaced) to denoising latent; derived from keyframe VAE encoding helpers.py: image_conditionings_by_adding_guiding_latent
Multimodal guidance Video and audio guidance factories providing CFG/STG/rescale guidance Passed to FactoryGuidedDenoiser
Distilled LoRA Lightweight adapter (450 steps) applied in Stage-2 for fast refinement KeyframeInterpolationPipeline constructor
Memory-safe blocks Self-contained components that free GPU memory after each stage blocks.py: all block classes
Adaptive tiling AUTO_TILING selects tile sizes based on VAE checkpoint specifications helpers.py tiling utilities

Key Source Files in LTX-2

Summary

  • KeyframeInterpolationPipeline generates smooth video from 2+ image keyframes using a two-stage diffusion process with additive guiding latents
  • Stage 1 creates low-resolution video at half resolution using FactoryGuidedDenoiser with multimodal guidance
  • Stage 2 upsamples via VideoUpsampler and refines with distilled LoRA weights via SimpleDenoiser
  • Additive conditioning (in helpers.py) preserves diffusion dynamics while steering toward keyframes—distinct from replacement strategies in other pipelines
  • All components are memory-safe blocks that prevent OOM errors during multi-stage generation
  • Available via CLI (python -m ltx_pipelines.keyframe_interpolation) or Python API (KeyframeInterpolationPipeline class)

Frequently Asked Questions

What makes KeyframeInterpolationPipeline different from image-to-video pipelines?

KeyframeInterpolationPipeline specifically handles multiple image inputs with temporal spacing, using additive guiding latents to interpolate smooth transitions between them. Standard image-to-video pipelines (like TI2VidTwoStagesPipeline) typically use latent replacement rather than addition, and don't optimize for multi-keyframe trajectories. The additive approach in image_conditionings_by_adding_guiding_latent preserves more of the base diffusion model's generation quality while still respecting keyframe constraints.

Why does the pipeline require distilled LoRA weights?

The distilled LoRA enables fast Stage-2 refinement without quality degradation. According to the LTX-2 source, Stage 2 uses SimpleDenoiser with few steps (stage_2_sigmas) because the upsampled video from Stage 1 already looks good—the LoRA provides the necessary adaptation for this lightweight refinement mode. The checkpoint ltx-2.5-22b-distilled-lora-450-bf16.safetensors was specifically trained for this purpose.

How does the pipeline handle GPU memory with long videos?

Memory management relies on self-contained block allocation and AUTO_TILING. Each heavy component (PromptEncoder, DiffusionStage, VideoDecoder, etc.) allocates GPU memory, runs its computation, and immediately frees it before returning. Additionally, the VAE decoder uses adaptive tiling based on the checkpoint's specifications, splitting large frames into manageable tiles without manual configuration.

Can I use more than three keyframes, and how are they temporally distributed?

Yes—provide any number of keyframes via the --images argument or images parameter. The pipeline distributes them evenly across num_frames and interpolates between consecutive keyframes. The guiding latents are computed per-frame in image_conditionings_by_adding_guiding_latent, ensuring each output frame receives appropriate conditioning based on its temporal proximity to the nearest keyframes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →