How the Keyframe Interpolation Pipeline Works in LTX-2: A Two-Stage Diffusion Approach

The KeyframeInterpolationPipeline in Lightricks/LTX-2 implements a two-stage diffusion workflow that generates high-quality video by first creating a low-resolution coarse draft at half resolution and then refining it with a distilled LoRA upsampler to produce smooth, temporally consistent output.

The keyframe interpolation pipeline in Lightricks/LTX-2 enables users to generate high-fidelity video sequences by interpolating between provided keyframes using a sophisticated diffusion-based architecture. Located in packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py, this pipeline orchestrates multiple modular components to transform text prompts and visual references into coherent video and audio outputs.

Pipeline Architecture and Component Initialization

The KeyframeInterpolationPipeline class is constructed by wiring together specialized components that enable two-stage generation. In the __init__ method (lines 46-104), the pipeline initializes a prompt encoder, an image conditioner, two DiffusionStage objects for sequential processing, a video upsampler, and separate decoders for video and audio modalities.

Notably, Stage 2 receives both regular LoRA weights and distilled LoRA weights, allowing the refinement phase to efficiently utilize compressed knowledge for high-resolution output without recomputing the full diffusion process from scratch.

The Two-Stage Generation Process

The __call__ method (starting at line 117) executes the generation workflow through a deterministic sequence of validation, encoding, denoising, and decoding phases.

Input Validation and Seed Configuration

The pipeline first validates compatibility between the requested resolution and the two-stage architecture using assert_resolution (line 123). It then initializes a torch.Generator seeded with the user-provided seed and wraps it in a GaussianNoiser (lines 125-127) to ensure reproducible noise sampling throughout the diffusion steps.

Prompt Encoding and Multimodal Conditioning

The PromptEncoder produces text-conditioned embeddings for both positive and negative prompts, generating separate video and audio context vectors (lines 129-136). These embeddings guide the diffusion process toward the semantic content described in the text prompts while avoiding unwanted artifacts specified in the negative prompts.

Stage 1: Low-Resolution Coarse Generation

Stage 1 operates at half the target resolution (e.g., 360p for a 720p output) to establish the initial temporal structure:

  1. Sigma Schedule Construction: The LTX2Scheduler (defined in packages/ltx-core/src/ltx_core/components/schedulers.py) builds a noise schedule governing the diffusion timesteps.
  2. Image Conditioning: Keyframe images are transformed into guidance latents via image_conditionings_by_adding_guiding_latent, enriching the conditioning signals to steer generation toward the supplied visual references.
  3. Guider Factory Instantiation: The create_multimodal_guider_factory (from packages/ltx-core/src/ltx_core/components/guiders.py) creates video and audio guiders with configurable CFG scales.
  4. Denoising Execution: A FactoryGuidedDenoiser processes the noisy latent through the first DiffusionStage using the computed stage_1_output_shape (lines 143-149) and stage_1_conditionings (lines 150-158), producing a coarse latent representation (lines 170-191).

Upsampling and Stage 2: High-Resolution Refinement

Following Stage 1, the pipeline upscales the coarse latent by a factor of 2 using the VideoUpsampler (line 194):

upscaled_video_latent = self.upsampler(video_state.latent[:1])

Stage 2 leverages distilled LoRA weights alongside the original LoRAs to refine the upsampled latent efficiently:

  • A new stage_2_sigmas schedule drives the refinement (line 196)
  • stage_2_conditionings are computed for the higher resolution (lines 198-207)
  • A simpler denoiser refines the visual details without the full computational cost of Stage 1 (lines 209-228)
  • Audio latents are simultaneously refined using the first sigma value as the noise scale

Decoding and Output Generation

The final decoding phase converts latents to pixel-space outputs:

  • The video latent is decoded by the VideoDecoder (optionally using tiled decoding for memory efficiency), located at line 230
  • The audio latent is processed by the AudioDecoder (line 231)
  • The method returns an iterator yielding video frames and the decoded audio waveform (line 232)

Programmatic Usage and CLI

You can instantiate the pipeline programmatically to integrate keyframe interpolation into custom workflows:

from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_core.loader import LoraPathStrengthAndSDOps
from ltx_core.types import MultiModalGuiderParams
import torch

# Initialize pipeline with checkpoint and LoRA configurations

pipeline = KeyframeInterpolationPipeline(
    checkpoint_path="/path/to/checkpoint",
    distilled_lora=[
        LoraPathStrengthAndSDOps(
            path="/path/to/distilled_lora.pt", 
            strength=0.7, 
            sd_ops=[]
        )
    ],
    spatial_upsampler_path="/path/to/spatial_upsampler.pt",
    gemma_root="/path/to/gemma",
    loras=[],  # Optional additional LoRAs

)

# Prepare keyframe images as list of (tensor, strength) tuples

images = []  # Populate with actual torch tensors

# Execute generation

video_iter, audio = pipeline(
    prompt="A sunrise over a mountain lake",
    negative_prompt="low quality",
    seed=42,
    height=512,
    width=512,
    num_frames=24,
    frame_rate=24.0,
    num_inference_steps=30,
    video_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
    audio_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
    images=images,
)

Alternatively, use the built-in CLI entry point (lines 335-393) for end-to-end generation:

python -m ltx_pipelines.keyframe_interpolation \
    --checkpoint_path /path/to/checkpoint \
    --distilled_lora /path/to/distilled_lora.pt \
    --spatial_upsampler_path /path/to/spatial_upsampler.pt \
    --gemma_root /path/to/gemma \
    --prompt "A futuristic cityscape at night" \
    --height 720 --width 1280 --num_frames 60 --frame_rate 30 \
    --output_path result.mp4

Summary

  • The KeyframeInterpolationPipeline employs a two-stage diffusion architecture: Stage 1 generates a coarse half-resolution draft, while Stage 2 refines the 2× upsampled output using distilled LoRA weights.
  • Modality-specific conditioning via image_conditionings_by_adding_guiding_latent and multimodal guider factories ensures precise alignment between keyframes, text prompts, and generated output.
  • Memory-efficient decoding supports tiled video decoding to handle high-resolution outputs without excessive GPU memory requirements.
  • The pipeline is fully modular, with components defined in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py and core utilities in packages/ltx-core, enabling reuse across different video generation tasks.

Frequently Asked Questions

What is the purpose of the two-stage design in the LTX-2 keyframe interpolation pipeline?

The two-stage design optimizes the quality-efficiency trade-off by first generating a temporally coherent coarse video at half resolution using the full diffusion model, then applying a computationally efficient refinement stage with distilled LoRAs to upscale and enhance details. This approach avoids the prohibitive cost of running full diffusion at the target resolution while maintaining high-fidelity output.

How does the pipeline handle memory efficiency during video decoding?

The VideoDecoder supports tiled decoding strategies that process the video latent in spatial chunks rather than loading the entire high-resolution frame into memory simultaneously. This tiling mechanism, referenced in the decoding call at line 230, enables the generation of high-resolution videos (e.g., 720p or 1080p) on consumer-grade hardware without out-of-memory errors.

What role do distilled LoRA weights play in Stage 2 of the interpolation process?

Distilled LoRA weights provide a compressed, efficient representation of the full model's knowledge for the refinement task. In Stage 2 (lines 196-228), these weights enable rapid denoising of the upsampled latent without reloading the full pretrained weights, significantly reducing computational overhead while preserving the visual quality established in Stage 1.

Can the keyframe interpolation pipeline generate audio alongside video?

Yes, the pipeline is architecturally multimodal. The PromptEncoder generates separate audio embeddings (lines 129-136), and the AudioDecoder (line 231) processes the refined audio latent into a waveform. The pipeline returns both a video frame iterator and the decoded audio array, synchronized through the shared conditioning and generation process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →