How to Use LTX-2 for Keyframe Interpolation: Two-Stage Pipeline Guide
LTX-2 interpolates sparse keyframes into smooth high-resolution video using a two-stage diffusion pipeline that first generates low-resolution latents, then upscales and refines them with a distilled LoRA.
The KeyframeInterpolationPipeline in the Lightricks/LTX-2 repository transforms a sparse set of image conditions into fluid, high-resolution video sequences. Implemented in packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py, this pipeline orchestrates a full-weight diffusion model, a specialized upsampler, and multimodal guidance to bridge the gaps between your keyframes.
Understanding the Two-Stage Architecture
LTX-2 keyframe interpolation operates through a cascaded generation process that balances computational efficiency with output quality.
Stage 1: Low-Resolution Latent Generation
The first stage generates video at ½ × target resolution using the full diffusion model. In packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py (lines 70-78), the DiffusionStage class receives a FactoryGuidedDenoiser that combines text and audio embeddings with multimodal guidance from MultiModalGuiderFactory (defined in ltx_core/components/guiders.py).
Image conditions are injected into the latent space via image_conditionings_by_adding_guiding_latent (lines 50-58), allowing the model to anchor the interpolation to your specific keyframes.
Stage 2: Upsampling and Refinement
The second stage operates on the low-resolution latent output:
- Spatial Upsampling: The
VideoUpsampler(fromltx_pipelines.utils.blocks, referenced at lines 98-100) scales the latent video 2× to reach the target resolution. - Distilled Refinement: A distilled LoRA checkpoint—lighter than the full model—is added to the LoRA list (lines 86-97) for fast refinement while preserving quality.
- Simplified Denoising: This stage uses
SimpleDenoiser(lines 108-112) instead of the full guided denoiser, as the upsampled video already contains most structural details.
Final Decoding
The pipeline concludes with VideoDecoder and AudioDecoder (from blocks.py) converting the final latents into tensors. The encode_video function (lines 30-33 in packages/ltx-pipelines/src/ltx_pipelines/media_io.py) writes the combined output to an MP4 file.
Running LTX-2 Keyframe Interpolation
Command-Line Interface
The CLI entry point uses default_2_stage_arg_parser from packages/ltx-pipelines/src/ltx_pipelines/utils/args.py to handle arguments, then executes via KeyframeInterpolationPipeline.__call__ (lines 36-44).
python -m ltx_pipelines.keyframe_interpolation \
--checkpoint_path /path/to/checkpoint \
--distilled_lora /path/to/distilled_lora.safetensors \
--spatial_upsampler_path /path/to/upsampler.safetensors \
--gemma_root /path/to/gemma \
--prompt "A sunrise over a mountain lake" \
--negative_prompt "" \
--seed 42 \
--height 720 \
--width 1280 \
--num_frames 48 \
--frame_rate 24 \
--num_inference_steps 50 \
--images img1.png img2.png img3.png \
--output_path result.mp4
Python API
For programmatic control, import the pipeline directly and configure the two-stage parameters:
from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_pipelines.utils.args import detect_checkpoint_path, detect_params
import torch
# Auto-detect checkpoint and parameters
ckpt = detect_checkpoint_path()
params = detect_params(ckpt)
# Initialize the two-stage pipeline
pipeline = KeyframeInterpolationPipeline(
checkpoint_path=ckpt,
distilled_lora=[("distilled_lora.safetensors", 0.8, None)],
spatial_upsampler_path="spatial_upsampler.safetensors",
gemma_root="gemma",
loras=[], # Optional additional LoRAs
quantization=None,
compilation_config=None,
offload_mode="none",
)
# Prepare image conditionings (list of (image_tensor, weight) tuples)
image_conditionings = []
# Execute interpolation
video_tensor, audio_tensor = pipeline(
prompt="A futuristic city at night",
negative_prompt="",
seed=1234,
height=720,
width=1280,
num_frames=48,
frame_rate=24.0,
num_inference_steps=50,
video_guider_params=None, # Use defaults or supply MultiModalGuiderParams
audio_guider_params=None,
images=image_conditionings,
)
The pipeline returns a video tensor (shape B, T, C, H, W) and an audio tensor, as defined in the return signature at lines 104-112 of keyframe_interpolation.py.
Key Components and Source Files
The following classes and modules work together to enable LTX-2 keyframe interpolation:
-
KeyframeInterpolationPipeline(packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py): Orchestrates the two-stage generation, upsampling, and decoding workflow. -
DiffusionStage(packages/ltx-pipelines/src/ltx_pipelines/blocks.py): Executes diffusion passes with optional LoRA weight integration. -
FactoryGuidedDenoiserandSimpleDenoiser(packages/ltx-pipelines/src/ltx_pipelines/denoisers.py): Apply guidance during Stage 1 and lightweight refinement during Stage 2, respectively. -
VideoUpsampler(packages/ltx-pipelines/src/ltx_pipelines/blocks.py): Handles 2× spatial upsampling of latent video between stages. -
PromptEncoderandImageConditioner(packages/ltx-pipelines/src/ltx_pipelines/blocks.py): Encode text prompts and transform keyframe images into latent conditionings compatible with the diffusion model. -
MultiModalGuiderFactory(packages/ltx-core/src/ltx_core/components/guiders.py): Provides the guidance mechanism that combines text, audio, and image modalities during the first stage.
Summary
- LTX-2 keyframe interpolation uses a
KeyframeInterpolationPipelinethat runs two distinct diffusion stages to convert sparse keyframes into high-resolution video. - Stage 1 generates at half resolution using
FactoryGuidedDenoiserwith multimodal guidance, while Stage 2 upscales viaVideoUpsamplerand refines with a distilled LoRA usingSimpleDenoiser. - Image conditions are injected through
image_conditionings_by_adding_guiding_latentin the first stage to anchor the motion between keyframes. - The pipeline supports both CLI execution via
default_2_stage_arg_parserand programmatic Python API access, with final output encoded throughencode_videoinmedia_io.py.
Frequently Asked Questions
What is the difference between Stage 1 and Stage 2 in LTX-2 keyframe interpolation?
Stage 1 generates the video at half the target resolution using the full-weight diffusion model with FactoryGuidedDenoiser to incorporate text, audio, and image guidance. Stage 2 upscales the result to full resolution using VideoUpsampler and applies a distilled LoRA with SimpleDenoiser for fast refinement, eliminating the complex guidance since the structural details are already established.
How are keyframe images processed by the pipeline?
Keyframe images are converted into latent conditionings by the ImageConditioner and injected into the diffusion process via image_conditionings_by_adding_guiding_latent (lines 50-58). This mechanism guides the GaussianNoiser (from ltx_core/components/noisers.py) to preserve the visual characteristics of your keyframes while generating intermediate frames.
Why does the pipeline use different denoisers for each stage?
The first stage requires FactoryGuidedDenoiser to combine multimodal inputs—text prompts, audio embeddings, and image conditions—into a coherent low-resolution video. The second stage uses SimpleDenoiser because the upsampled latent already contains the necessary structural information, allowing for faster, lighter processing without the overhead of full multimodal guidance.
Can I customize the LoRA weights used during refinement?
Yes, the KeyframeInterpolationPipeline accepts a distilled_lora parameter (lines 86-97) that accepts a list of tuples specifying the LoRA path, scale factor, and optional trigger words. You can also add supplementary LoRAs via the loras parameter to influence the style or motion characteristics beyond the distilled checkpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →