Configuring Generated Keyframes for Enhanced Temporal Resolution in LTX-2
To increase temporal resolution in LTX-2 video generation, use --num-generated-keyframes N to insert additional latent frames that are denoised independently, reducing the diffusion length by approximately N / num_latent_frames.
LTX-2 introduces generated keyframes—intermediate latent frames processed outside the main diffusion trajectory—to improve motion smoothness and fine detail. This article walks through the source code implementation in the Lightricks/LTX-2 repository, showing how to configure keyframes via CLI or Python API and how the pipeline validates checkpoint compatibility.
How Generated Keyframes Improve Temporal Resolution
Standard diffusion processes every frame in sequence. Generated keyframes break this linearity by designating specific pixel-frame indices as independent denoising targets. Each keyframe receives dedicated diffusion steps, effectively shortening the path between adjacent frames in latent space.
The performance trade-off is straightforward: more keyframes yield sharper motion boundaries but increase compute. The relationship is roughly linear—N keyframes reduce effective diffusion length by ~N / total_latent_frames.
CLI Configuration with --num-generated-keyframes
The unified argument parser exposes keyframe control via add_generated_keyframes_arg in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 826-842):
# Typical pipeline entry point
from ltx_pipelines.utils.args import add_generated_keyframes_arg, default_2_stage_arg_parser
from ltx_pipelines.ti2vid_params import Ti2VidParams
params = Ti2VidParams()
parser = add_generated_keyframes_arg(
default_2_stage_arg_parser(params=params, supports_auto_duration=True)
)
Command-line usage follows standard argparse patterns:
python -m ltx_pipelines.ti2vid_two_stages \
--input video.mp4 \
--output out.mp4 \
--num-generated-keyframes 4
Resolving Keyframe Positions from User Input
The helper resolve_generated_keyframes in packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py handles two input modes (lines 70-110):
- Integer input: Triggers
evenly_spaced_keyframe_positionsto distribute keyframes across interior frames, excluding first and last positions - Sequence input: Validates explicit frame indices fall within
[0, num_frames)
from ltx_pipelines.utils.helpers import resolve_generated_keyframes
# Mode 1: Integer for evenly-spaced keyframes
positions = resolve_generated_keyframes(3, num_frames=16) # [4, 8, 12]
# Mode 2: Explicit list of indices
positions = resolve_generated_keyframes([3, 7, 11], num_frames=16) # [3, 7, 11]
Building Keyframe Conditionings
The generated_keyframe_conditionings function (lines 14-26 in helpers.py) constructs pipeline-ready conditionings:
from ltx_pipelines.utils.helpers import generated_keyframe_conditionings
# Example: 2 keyframes across 16 frames
conditioning = generated_keyframe_conditionings(2, num_frames=16)
# Returns: [VideoGeneratedKeyframeSlots(pixel_frame_indices=[5, 10])]
# Example: Explicit positions
conditioning = generated_keyframe_conditionings([3, 7, 11], num_frames=16)
# Returns: [VideoGeneratedKeyframeSlots(pixel_frame_indices=[3, 7, 11])]
These conditionings integrate into the standard conditioning list passed to diffusion stages.
Checkpoint Compatibility Validation
Not all LTX-2 checkpoints support generated keyframes. The DiffusionStage class in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py enforces this via assert_generated_keyframes_supported (lines 395-418):
from ltx_pipelines.utils.blocks import DiffusionStage
stage = DiffusionStage(transformer_builder=my_builder)
stage.assert_generated_keyframes_supported() # Raises RuntimeError if unsupported
The check verifies the checkpoint includes keyframe absolute-position embeddings (use_keyframes_abs_pos_embedding). This prevents silent failures when keyframes are requested on incompatible models.
Mask Propagation Through the Pipeline
Once validated, keyframe information flows through the pipeline as keyframes_mask attached to LatentState objects. In packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py (lines 52-56), the denoiser consumes this mask:
# Inside denoiser forward pass
if state.keyframes_mask is not None:
keyframe_mask_expanded = _repeat(state.keyframes_mask)
# Applied to distinguish keyframe tokens during denoising
The mask ensures keyframe tokens receive separate processing while maintaining alignment with non-keyframe positions.
Complete Workflow: TI2Vid Two-Stage Pipeline
The ti2vid_two_stages.py pipeline demonstrates end-to-end keyframe integration:
- Parse:
add_generated_keyframes_argextracts--num-generated-keyframes - Resolve:
resolve_generated_keyframesnormalizes to pixel-frame indices - Construct:
generated_keyframe_conditioningsbuildsVideoGeneratedKeyframeSlots - Validate:
stage_1.assert_generated_keyframes_supported()checks checkpoint - Create State:
create_noised_stateattacheskeyframes_maskto latent state - Denoise: Denoiser processes keyframes via mask-aware attention
Summary
- Generated keyframes increase temporal resolution by adding independently-denoised latent frames
- Configure via
--num-generated-keyframes N(CLI) orgenerated_keyframe_conditionings()(Python) - Positions resolve as evenly-spaced integers or explicit index lists in
helpers.py - Checkpoints must support
use_keyframes_abs_pos_embedding; enforced byblocks.py keyframes_maskpropagates throughLatentStatefor mask-aware denoising
Frequently Asked Questions
How many generated keyframes should I use?
Start with 2-4 keyframes for 16-frame sequences, scaling with total frame count. Each keyframe improves temporal fidelity at roughly linear compute cost. For content with rapid motion, increase density; for static scenes, reduce or disable entirely.
Can I use generated keyframes with any LTX-2 checkpoint?
No. The checkpoint must include keyframe absolute-position embeddings. Call assert_generated_keyframes_supported() on your diffusion stage to verify—this raises a clear error if the checkpoint lacks required weights rather than producing garbled output.
What's the difference between integer and list input for keyframe positions?
An integer triggers automatic spacing via evenly_spaced_keyframe_positions, distributing keyframes uniformly while excluding first and last frames. A list gives precise control over which pixel-frame indices serve as keyframes, useful for targeting specific motion events.
Do generated keyframes affect inference speed?
Yes—each keyframe adds independent denoising steps. The overhead scales approximately with keyframe count relative to total latent frames. For real-time applications, profile with num_generated_keyframes=0 as baseline, then increment to find your quality-latency tradeoff.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →