LTX-2 KeyframeInterpolationPipeline: Generate Smooth Video from Image Keyframes
The LTX-2 KeyframeInterpolationPipeline uses a two-stage diffusion workflow with additive guiding latents to interpolate between user-supplied image keyframes, producing temporally coherent full-resolution video. This pipeline lives in packages/ltx-pipelines and implements a unique conditioning strategy that preserves underlying diffusion dynamics while steering output toward specified keyframes.
How KeyframeInterpolationPipeline Works in LTX-2
The KeyframeInterpolationPipeline follows the same architectural pattern as LTX-2's other two-stage pipelines (like TI2VidTwoStagesPipeline), but replaces latent replacement with additive guiding latents. This approach enables smooth transitions between keyframes without disrupting the diffusion process.
Stage 1: Low-Resolution Generation with Guiding Latents
The pipeline begins by encoding prompts and preparing image conditioning:
-
Prompt encoding —
PromptEncoder(frompackages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) converts positive/negative text prompts into video and audio conditioning contexts:v_context_p,a_context_p,v_context_n,a_context_n. -
Image conditioning — The
ImageConditionerloads the VAE checkpoint and creates conditioning latents by adding a guiding latent derived from each keyframe image to the latent being denoised. This happens inimage_conditionings_by_adding_guiding_latentwithinpackages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py. -
Diffusion pass — A
DiffusionStagebuilt from the full transformer checkpoint runs at half target resolution. TheFactoryGuidedDenoiserreceives video/audio contexts plus multimodal guider factories fromcreate_multimodal_guider_factory.
The result is a low-resolution latent video (video_state.latent) that already follows the keyframe trajectory.
Stage 2: Upsampling and Refinement
-
Spatial upsampling —
VideoUpsampler(also inblocks.py) lifts the latent frompackages/ltx-pipelines/src/ltx_pipelines/utils/blocks.pyto full resolution (×2) using the pretrained spatial upsampler checkpoint. Temporal continuity is preserved through direct latent feed-forward. -
Refinement with distilled LoRA — The second
DiffusionStageloads the same transformer checkpoint but applies distilled LoRA weights. It usesSimpleDenoiserbecause the upscaled video requires only minimal refinement steps (stage_2_sigmas). The initial latent combines upscaled video with the same guiding latents from keyframes. -
Decoding and output —
VideoDecoderandAudioDecoderconvert final latents to pixels and waveforms, respectingAUTO_TILINGconfiguration for memory efficiency.encode_videowrites the result to disk.
All heavy components (PromptEncoder, ImageConditioner, DiffusionStage, VideoUpsampler, VideoDecoder, AudioDecoder) are self-contained blocks that allocate, run, and immediately free GPU memory—critical for preventing OOM errors in multi-stage workflows.
Running KeyframeInterpolationPipeline from CLI
The pipeline exposes a full command-line interface through keyframe_interpolation.py:
uv run python -m ltx_pipelines.keyframe_interpolation \
--transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors \
--text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
--video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
--audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
--spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
--distilled-lora models/ltx-2.5/loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors \
--prompt "A sunrise over a quiet lake" \
--negative-prompt "" \
--seed 1234 \
--height 720 \
--width 1280 \
--num-frames 121 \
--frame-rate 30 \
--num-inference-steps 30 \
--images path/to/keyframe1.png path/to/keyframe2.png path/to/keyframe3.png \
--output-path output.mp4
Key parameters specific to keyframe interpolation:
--images— List of keyframe image paths (minimum 2 for meaningful interpolation)--distilled-lora— Required for Stage-2 refinement speedup--num-frames— Total output length; intermediate frames are generated between keyframes
Using KeyframeInterpolationPipeline in Python
For programmatic control, instantiate KeyframeInterpolationPipeline directly:
from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.args import ImageConditioningInput
from PIL import Image
# Configure model paths (download from HuggingFace first)
paths = ModelPaths(
transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors",
text_encoder="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
video_vae="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
audio_vae="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)
pipeline = KeyframeInterpolationPipeline(
model_paths=paths,
distilled_lora=[("models/ltx-2.5/loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors", 1.0, None)],
spatial_upsampler_path=paths.spatial_upsampler,
loras=[], # Additional LoRAs can be loaded here
)
# Prepare keyframe inputs
keyframes = [
ImageConditioningInput(image=Image.open("keyframe1.png")),
ImageConditioningInput(image=Image.open("keyframe2.png")),
ImageConditioningInput(image=Image.open("keyframe3.png")),
]
# Generate video
video, audio, metadata = pipeline(
prompt="A sunrise over a quiet lake",
negative_prompt="",
seed=1234,
height=720,
width=1280,
num_frames=121,
frame_rate=30.0,
num_inference_steps=30,
video_guider_params=..., # MultiModalGuiderParams for CFG/STG guidance
audio_guider_params=..., # Audio guidance configuration
images=keyframes,
)
# video: iterator of torch.Tensor frames
# audio: raw waveform array
# Export with encode_video from ltx_pipelines.utils.media_io
Core Architecture Concepts
| Concept | Implementation | Location |
|---|---|---|
| Guiding latent | Added (not replaced) to denoising latent; derived from keyframe VAE encoding | helpers.py: image_conditionings_by_adding_guiding_latent |
| Multimodal guidance | Video and audio guidance factories providing CFG/STG/rescale guidance | Passed to FactoryGuidedDenoiser |
| Distilled LoRA | Lightweight adapter (450 steps) applied in Stage-2 for fast refinement | KeyframeInterpolationPipeline constructor |
| Memory-safe blocks | Self-contained components that free GPU memory after each stage | blocks.py: all block classes |
| Adaptive tiling | AUTO_TILING selects tile sizes based on VAE checkpoint specifications |
helpers.py tiling utilities |
Key Source Files in LTX-2
packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py— MainKeyframeInterpolationPipelineclass and CLI entry pointpackages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py— Reusable pipeline blocks (PromptEncoder,ImageConditioner,DiffusionStage,VideoUpsampler,VideoDecoder,AudioDecoder)packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py— Guiding latent creation, tiling logic, memory cleanuppackages/ltx-pipelines/src/ltx_pipelines/utils/args.py—ImageConditioningInputand argument parsingpackages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py— Video encoding, HDR handling, VAE dtype selection
Summary
- KeyframeInterpolationPipeline generates smooth video from 2+ image keyframes using a two-stage diffusion process with additive guiding latents
- Stage 1 creates low-resolution video at half resolution using
FactoryGuidedDenoiserwith multimodal guidance - Stage 2 upsamples via
VideoUpsamplerand refines with distilled LoRA weights viaSimpleDenoiser - Additive conditioning (in
helpers.py) preserves diffusion dynamics while steering toward keyframes—distinct from replacement strategies in other pipelines - All components are memory-safe blocks that prevent OOM errors during multi-stage generation
- Available via CLI (
python -m ltx_pipelines.keyframe_interpolation) or Python API (KeyframeInterpolationPipelineclass)
Frequently Asked Questions
What makes KeyframeInterpolationPipeline different from image-to-video pipelines?
KeyframeInterpolationPipeline specifically handles multiple image inputs with temporal spacing, using additive guiding latents to interpolate smooth transitions between them. Standard image-to-video pipelines (like TI2VidTwoStagesPipeline) typically use latent replacement rather than addition, and don't optimize for multi-keyframe trajectories. The additive approach in image_conditionings_by_adding_guiding_latent preserves more of the base diffusion model's generation quality while still respecting keyframe constraints.
Why does the pipeline require distilled LoRA weights?
The distilled LoRA enables fast Stage-2 refinement without quality degradation. According to the LTX-2 source, Stage 2 uses SimpleDenoiser with few steps (stage_2_sigmas) because the upsampled video from Stage 1 already looks good—the LoRA provides the necessary adaptation for this lightweight refinement mode. The checkpoint ltx-2.5-22b-distilled-lora-450-bf16.safetensors was specifically trained for this purpose.
How does the pipeline handle GPU memory with long videos?
Memory management relies on self-contained block allocation and AUTO_TILING. Each heavy component (PromptEncoder, DiffusionStage, VideoDecoder, etc.) allocates GPU memory, runs its computation, and immediately frees it before returning. Additionally, the VAE decoder uses adaptive tiling based on the checkpoint's specifications, splitting large frames into manageable tiles without manual configuration.
Can I use more than three keyframes, and how are they temporally distributed?
Yes—provide any number of keyframes via the --images argument or images parameter. The pipeline distributes them evenly across num_frames and interpolates between consecutive keyframes. The guiding latents are computed per-frame in image_conditionings_by_adding_guiding_latent, ensuring each output frame receives appropriate conditioning based on its temporal proximity to the nearest keyframes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →