How to Work with Generated Keyframes in Lightricks LTX-2: A Complete Technical Guide
The generated keyframes feature in LTX-2 produces additional anchor frames during video diffusion that relax temporal compression at specific positions, resulting in smoother motion and higher visual fidelity when enabled via CLI or Python.
LTX-2's generated keyframes capability marks a significant advancement for video diffusion quality. This optional feature inserts extra conditioning points—beyond your provided start/end frames—that the model treats as full keyframes with dedicated latent tokens. According to the Lightricks/LTX-2 source code, mastering this feature requires understanding its dual enablement path (CLI flags plus checkpoint compatibility) and its modular pipeline integration.
How Generated Keyframes Work in LTX-2
The architecture operates through three coordinated layers: user input parsing, position calculation, and diffusion-stage injection. Understanding this flow ensures you can debug issues and customize behavior when the default settings don't suffice.
The Three-Stage Processing Pipeline
| Stage | Component | Function |
|---|---|---|
| CLI Registration | add_generated_keyframes_arg in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py#L26 |
Exposes --num-generated-keyframes to users |
| Position Resolution | resolve_generated_keyframes and evenly_spaced_keyframe_positions in packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L70-L94 |
Converts integers to frame indices or validates custom lists |
| Conditioning Injection | generated_keyframe_conditionings in packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L14 |
Wraps positions in VideoGeneratedKeyframeSlots |
| Capability Guard | assert_generated_keyframes_supported in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L405 |
Validates checkpoint compatibility |
When KeyframeInterpolationPipeline.__call__ executes (lines 210-227 in packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py), it invokes generated_keyframe_conditionings to build the conditioning object. This object is appended to stage_1_conditionings before denoising begins.
Critical requirement: Your checkpoint must have use_keyframes_abs_pos_embedding enabled in its transformer config. The DiffusionStage.assert_generated_keyframes_supported method enforces this—attempting to use generated keyframes with incompatible checkpoints raises a clear runtime error rather than failing silently.
Enabling Generated Keyframes via CLI
The fastest path to working with generated keyframes uses the built-in argument parser.
python -m ltx_pipelines.ti2vid_two_stages \
--model-paths /path/to/models \
--prompt "A futuristic cityscape transitioning from day to night" \
--height 720 \
--width 1280 \
--num-frames 32 \
--frame-rate 30 \
--num-generated-keyframes 4 \
--image /imgs/day_scene.jpg 0 1.0 \
--image /imgs/night_scene.jpg 31 1.0 \
--output-path result.mp4
What happens internally:
--num-generated-keyframes 4triggersevenly_spaced_keyframe_positions(4, 32)- This computes interior positions approximately at frames [6, 12, 18, 24] for a 32-frame sequence
- These positions exclude start (0) and end frames to preserve your explicit conditioning
You can also pass explicit frame indices:
--num-generated-keyframes 8 16 24
The resolve_generated_keyframes helper handles both formats—integers trigger even spacing, while lists undergo validation against the total frame count.
Programmatic Usage in Python
For research or application integration, instantiate the pipeline directly with generated keyframe support.
from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_pipelines.utils.args import add_generated_keyframes_arg, default_2_stage_arg_parser
from ltx_pipelines.utils.types import ImageConditioningInput
from ltx_core.model.paths import ModelPaths
# Step 1: Build parser with keyframe capability
base_parser = default_2_stage_arg_parser(
params={"supports_auto_duration": True}
)
parser = add_generated_keyframes_arg(base_parser)
args = parser.parse_args()
# Step 2: Initialize pipeline with model paths
pipeline = KeyframeInterpolationPipeline(
model_paths=ModelPaths.from_str(args.model_paths),
distilled_lora=args.distilled_lora,
spatial_upsampler_path=args.spatial_upsampler_path,
loras=args.lora,
)
# Step 3: Execute with generated keyframes parameter
video, audio, tiling_cfg = pipeline(
prompt=args.prompt,
negative_prompt=args.negative_prompt,
seed=args.seed,
height=args.height,
width=args.width,
num_frames=args.num_frames,
frame_rate=args.frame_rate,
num_inference_steps=args.num_inference_steps,
video_guider_params=video_guider_params,
audio_guider_params=audio_guider_params,
images=args.images,
# Generated keyframes flow through pipeline context
)
Important implementation detail: The pipeline.__call__ signature doesn't expose num_generated_keyframes directly. Instead, the parameter passes through the surrounding execution context, where the internal generated_keyframes variable is resolved from parsed arguments.
Advanced: Manual Conditioning Construction
For non-uniform keyframe placement or dynamic position calculation, construct the conditioning object explicitly.
from ltx_pipelines.utils.helpers import generated_keyframe_conditionings
# Custom positions for dramatic motion changes
critical_motion_frames = [5, 12, 20, 28]
total_frames = 32
# Build the conditioning object
keyframe_cond = generated_keyframe_conditionings(
generated_keyframes=critical_motion_frames,
num_frames=total_frames
)
# Inject into stage 1 conditionings before diffusion
stage_1_conditionings = existing_image_conditionings.copy()
stage_1_conditionings.extend(keyframe_cond)
# Pass to diffusion stage as normal
This manual approach reveals the underlying data structure: VideoGeneratedKeyframeSlots instances that carry absolute position embeddings. The checkpoint's use_keyframes_abs_pos_embedding flag must still be active for these tokens to process correctly.
Checkpoint Compatibility and Error Handling
The assert_generated_keyframes_supported method in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L405 performs a rigorous capability check before computation begins.
Verification criteria:
- Transformer config contains
use_keyframes_abs_pos_embedding: True - The embedding layers are initialized in the state dict
- Position indices fall within valid temporal ranges
If validation fails, the error message identifies the specific incompatibility—typically indicating you need a checkpoint trained with keyframe-aware positional embeddings.
Performance and Quality Trade-offs
| Configuration | Temporal Fidelity | Compute Cost | Best For |
|---|---|---|---|
| 0 generated keyframes | Baseline | Standard | Fast prototyping, simple scenes |
| 2-4 evenly spaced | Moderate improvement | +15-25% | General video with moderate motion |
| 6-8+ or custom positions | Maximum smoothness | +40-60% | Complex transformations, high-framerate output |
Each generated keyframe consumes additional latent tokens during diffusion. The absolute position embedding mechanism prevents collision with standard frame conditioning, but GPU memory scales with total keyframe count.
Summary
- Enable via CLI with
--num-generated-keyframes Nfor even spacing, or pass explicit frame indices - Validate checkpoint compatibility—the
use_keyframes_abs_pos_embeddingflag must be present in transformer config - Understand the pipeline flow: argument parsing → position resolution →
VideoGeneratedKeyframeSlotsconstruction → stage 1 conditioning injection - Reference key source files:
args.pyfor CLI,helpers.pyfor position logic,blocks.pyfor capability guards,keyframe_interpolation.pyfor integration - Use manual conditioning construction when you need precise control over keyframe placement
Frequently Asked Questions
What happens if I use generated keyframes with an incompatible checkpoint?
The pipeline raises a runtime error from assert_generated_keyframes_supported before any diffusion begins. The error message specifies that your checkpoint lacks use_keyframes_abs_pos_embedding support. No silent degradation occurs—you must obtain or train a compatible checkpoint.
Can I mix user-provided keyframes with generated keyframes?
Yes. The images parameter accepts ImageConditioningInput objects with explicit frame indices, while num_generated_keyframes inserts additional anchor points. The two conditioning types coexist in stage_1_conditionings without conflict, each processed by appropriate embedding layers.
How are frame positions calculated for integer keyframe counts?
The evenly_spaced_keyframe_positions function divides the interior frame range (excluding boundaries 0 and num_frames-1) into N+1 segments, placing keyframes at segment boundaries. For 32 frames with 4 generated keyframes: positions land at frames 6, 12, 18, and 24 approximately.
Does increasing generated keyframes always improve quality?
Not universally. Beyond 6-8 keyframes for typical sequences, diminishing returns set in while compute costs rise substantially. Complex motion scenes benefit most; static or slow-moving content shows minimal improvement. Experiment with your specific content domain to find optimal settings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →