How to Use LTX-2 for Keyframe Interpolation: Two-Stage Pipeline Guide

LTX-2 interpolates sparse keyframes into smooth high-resolution video using a two-stage diffusion pipeline that first generates low-resolution latents, then upscales and refines them with a distilled LoRA.

The KeyframeInterpolationPipeline in the Lightricks/LTX-2 repository transforms a sparse set of image conditions into fluid, high-resolution video sequences. Implemented in packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py, this pipeline orchestrates a full-weight diffusion model, a specialized upsampler, and multimodal guidance to bridge the gaps between your keyframes.

Understanding the Two-Stage Architecture

LTX-2 keyframe interpolation operates through a cascaded generation process that balances computational efficiency with output quality.

Stage 1: Low-Resolution Latent Generation

The first stage generates video at ½ × target resolution using the full diffusion model. In packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py (lines 70-78), the DiffusionStage class receives a FactoryGuidedDenoiser that combines text and audio embeddings with multimodal guidance from MultiModalGuiderFactory (defined in ltx_core/components/guiders.py).

Image conditions are injected into the latent space via image_conditionings_by_adding_guiding_latent (lines 50-58), allowing the model to anchor the interpolation to your specific keyframes.

Stage 2: Upsampling and Refinement

The second stage operates on the low-resolution latent output:

  1. Spatial Upsampling: The VideoUpsampler (from ltx_pipelines.utils.blocks, referenced at lines 98-100) scales the latent video 2× to reach the target resolution.
  2. Distilled Refinement: A distilled LoRA checkpoint—lighter than the full model—is added to the LoRA list (lines 86-97) for fast refinement while preserving quality.
  3. Simplified Denoising: This stage uses SimpleDenoiser (lines 108-112) instead of the full guided denoiser, as the upsampled video already contains most structural details.

Final Decoding

The pipeline concludes with VideoDecoder and AudioDecoder (from blocks.py) converting the final latents into tensors. The encode_video function (lines 30-33 in packages/ltx-pipelines/src/ltx_pipelines/media_io.py) writes the combined output to an MP4 file.

Running LTX-2 Keyframe Interpolation

Command-Line Interface

The CLI entry point uses default_2_stage_arg_parser from packages/ltx-pipelines/src/ltx_pipelines/utils/args.py to handle arguments, then executes via KeyframeInterpolationPipeline.__call__ (lines 36-44).

python -m ltx_pipelines.keyframe_interpolation \
    --checkpoint_path /path/to/checkpoint \
    --distilled_lora /path/to/distilled_lora.safetensors \
    --spatial_upsampler_path /path/to/upsampler.safetensors \
    --gemma_root /path/to/gemma \
    --prompt "A sunrise over a mountain lake" \
    --negative_prompt "" \
    --seed 42 \
    --height 720 \
    --width 1280 \
    --num_frames 48 \
    --frame_rate 24 \
    --num_inference_steps 50 \
    --images img1.png img2.png img3.png \
    --output_path result.mp4

Python API

For programmatic control, import the pipeline directly and configure the two-stage parameters:

from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_pipelines.utils.args import detect_checkpoint_path, detect_params
import torch

# Auto-detect checkpoint and parameters

ckpt = detect_checkpoint_path()
params = detect_params(ckpt)

# Initialize the two-stage pipeline

pipeline = KeyframeInterpolationPipeline(
    checkpoint_path=ckpt,
    distilled_lora=[("distilled_lora.safetensors", 0.8, None)],
    spatial_upsampler_path="spatial_upsampler.safetensors",
    gemma_root="gemma",
    loras=[],                     # Optional additional LoRAs

    quantization=None,
    compilation_config=None,
    offload_mode="none",
)

# Prepare image conditionings (list of (image_tensor, weight) tuples)

image_conditionings = []

# Execute interpolation

video_tensor, audio_tensor = pipeline(
    prompt="A futuristic city at night",
    negative_prompt="",
    seed=1234,
    height=720,
    width=1280,
    num_frames=48,
    frame_rate=24.0,
    num_inference_steps=50,
    video_guider_params=None,   # Use defaults or supply MultiModalGuiderParams

    audio_guider_params=None,
    images=image_conditionings,
)

The pipeline returns a video tensor (shape B, T, C, H, W) and an audio tensor, as defined in the return signature at lines 104-112 of keyframe_interpolation.py.

Key Components and Source Files

The following classes and modules work together to enable LTX-2 keyframe interpolation:

Summary

  • LTX-2 keyframe interpolation uses a KeyframeInterpolationPipeline that runs two distinct diffusion stages to convert sparse keyframes into high-resolution video.
  • Stage 1 generates at half resolution using FactoryGuidedDenoiser with multimodal guidance, while Stage 2 upscales via VideoUpsampler and refines with a distilled LoRA using SimpleDenoiser.
  • Image conditions are injected through image_conditionings_by_adding_guiding_latent in the first stage to anchor the motion between keyframes.
  • The pipeline supports both CLI execution via default_2_stage_arg_parser and programmatic Python API access, with final output encoded through encode_video in media_io.py.

Frequently Asked Questions

What is the difference between Stage 1 and Stage 2 in LTX-2 keyframe interpolation?

Stage 1 generates the video at half the target resolution using the full-weight diffusion model with FactoryGuidedDenoiser to incorporate text, audio, and image guidance. Stage 2 upscales the result to full resolution using VideoUpsampler and applies a distilled LoRA with SimpleDenoiser for fast refinement, eliminating the complex guidance since the structural details are already established.

How are keyframe images processed by the pipeline?

Keyframe images are converted into latent conditionings by the ImageConditioner and injected into the diffusion process via image_conditionings_by_adding_guiding_latent (lines 50-58). This mechanism guides the GaussianNoiser (from ltx_core/components/noisers.py) to preserve the visual characteristics of your keyframes while generating intermediate frames.

Why does the pipeline use different denoisers for each stage?

The first stage requires FactoryGuidedDenoiser to combine multimodal inputs—text prompts, audio embeddings, and image conditions—into a coherent low-resolution video. The second stage uses SimpleDenoiser because the upsampled latent already contains the necessary structural information, allowing for faster, lighter processing without the overhead of full multimodal guidance.

Can I customize the LoRA weights used during refinement?

Yes, the KeyframeInterpolationPipeline accepts a distilled_lora parameter (lines 86-97) that accepts a list of tuples specifying the LoRA path, scale factor, and optional trigger words. You can also add supplementary LoRAs via the loras parameter to influence the style or motion characteristics beyond the distilled checkpoint.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →