How to Work with Generated Keyframes in Lightricks LTX-2: A Complete Technical Guide

The generated keyframes feature in LTX-2 produces additional anchor frames during video diffusion that relax temporal compression at specific positions, resulting in smoother motion and higher visual fidelity when enabled via CLI or Python.

LTX-2's generated keyframes capability marks a significant advancement for video diffusion quality. This optional feature inserts extra conditioning points—beyond your provided start/end frames—that the model treats as full keyframes with dedicated latent tokens. According to the Lightricks/LTX-2 source code, mastering this feature requires understanding its dual enablement path (CLI flags plus checkpoint compatibility) and its modular pipeline integration.

How Generated Keyframes Work in LTX-2

The architecture operates through three coordinated layers: user input parsing, position calculation, and diffusion-stage injection. Understanding this flow ensures you can debug issues and customize behavior when the default settings don't suffice.

The Three-Stage Processing Pipeline

Stage Component Function
CLI Registration add_generated_keyframes_arg in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py#L26 Exposes --num-generated-keyframes to users
Position Resolution resolve_generated_keyframes and evenly_spaced_keyframe_positions in packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L70-L94 Converts integers to frame indices or validates custom lists
Conditioning Injection generated_keyframe_conditionings in packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L14 Wraps positions in VideoGeneratedKeyframeSlots
Capability Guard assert_generated_keyframes_supported in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L405 Validates checkpoint compatibility

When KeyframeInterpolationPipeline.__call__ executes (lines 210-227 in packages/ltx-pipelines/src/ltx_pipelines/keyframe_interpolation.py), it invokes generated_keyframe_conditionings to build the conditioning object. This object is appended to stage_1_conditionings before denoising begins.

Critical requirement: Your checkpoint must have use_keyframes_abs_pos_embedding enabled in its transformer config. The DiffusionStage.assert_generated_keyframes_supported method enforces this—attempting to use generated keyframes with incompatible checkpoints raises a clear runtime error rather than failing silently.

Enabling Generated Keyframes via CLI

The fastest path to working with generated keyframes uses the built-in argument parser.

python -m ltx_pipelines.ti2vid_two_stages \
  --model-paths /path/to/models \
  --prompt "A futuristic cityscape transitioning from day to night" \
  --height 720 \
  --width 1280 \
  --num-frames 32 \
  --frame-rate 30 \
  --num-generated-keyframes 4 \
  --image /imgs/day_scene.jpg 0 1.0 \
  --image /imgs/night_scene.jpg 31 1.0 \
  --output-path result.mp4

What happens internally:

  • --num-generated-keyframes 4 triggers evenly_spaced_keyframe_positions(4, 32)
  • This computes interior positions approximately at frames [6, 12, 18, 24] for a 32-frame sequence
  • These positions exclude start (0) and end frames to preserve your explicit conditioning

You can also pass explicit frame indices:

--num-generated-keyframes 8 16 24

The resolve_generated_keyframes helper handles both formats—integers trigger even spacing, while lists undergo validation against the total frame count.

Programmatic Usage in Python

For research or application integration, instantiate the pipeline directly with generated keyframe support.

from ltx_pipelines.keyframe_interpolation import KeyframeInterpolationPipeline
from ltx_pipelines.utils.args import add_generated_keyframes_arg, default_2_stage_arg_parser
from ltx_pipelines.utils.types import ImageConditioningInput
from ltx_core.model.paths import ModelPaths

# Step 1: Build parser with keyframe capability

base_parser = default_2_stage_arg_parser(
    params={"supports_auto_duration": True}
)
parser = add_generated_keyframes_arg(base_parser)
args = parser.parse_args()

# Step 2: Initialize pipeline with model paths

pipeline = KeyframeInterpolationPipeline(
    model_paths=ModelPaths.from_str(args.model_paths),
    distilled_lora=args.distilled_lora,
    spatial_upsampler_path=args.spatial_upsampler_path,
    loras=args.lora,
)

# Step 3: Execute with generated keyframes parameter

video, audio, tiling_cfg = pipeline(
    prompt=args.prompt,
    negative_prompt=args.negative_prompt,
    seed=args.seed,
    height=args.height,
    width=args.width,
    num_frames=args.num_frames,
    frame_rate=args.frame_rate,
    num_inference_steps=args.num_inference_steps,
    video_guider_params=video_guider_params,
    audio_guider_params=audio_guider_params,
    images=args.images,
    # Generated keyframes flow through pipeline context

)

Important implementation detail: The pipeline.__call__ signature doesn't expose num_generated_keyframes directly. Instead, the parameter passes through the surrounding execution context, where the internal generated_keyframes variable is resolved from parsed arguments.

Advanced: Manual Conditioning Construction

For non-uniform keyframe placement or dynamic position calculation, construct the conditioning object explicitly.

from ltx_pipelines.utils.helpers import generated_keyframe_conditionings

# Custom positions for dramatic motion changes

critical_motion_frames = [5, 12, 20, 28]
total_frames = 32

# Build the conditioning object

keyframe_cond = generated_keyframe_conditionings(
    generated_keyframes=critical_motion_frames,
    num_frames=total_frames
)

# Inject into stage 1 conditionings before diffusion

stage_1_conditionings = existing_image_conditionings.copy()
stage_1_conditionings.extend(keyframe_cond)

# Pass to diffusion stage as normal

This manual approach reveals the underlying data structure: VideoGeneratedKeyframeSlots instances that carry absolute position embeddings. The checkpoint's use_keyframes_abs_pos_embedding flag must still be active for these tokens to process correctly.

Checkpoint Compatibility and Error Handling

The assert_generated_keyframes_supported method in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L405 performs a rigorous capability check before computation begins.

Verification criteria:

  • Transformer config contains use_keyframes_abs_pos_embedding: True
  • The embedding layers are initialized in the state dict
  • Position indices fall within valid temporal ranges

If validation fails, the error message identifies the specific incompatibility—typically indicating you need a checkpoint trained with keyframe-aware positional embeddings.

Performance and Quality Trade-offs

Configuration Temporal Fidelity Compute Cost Best For
0 generated keyframes Baseline Standard Fast prototyping, simple scenes
2-4 evenly spaced Moderate improvement +15-25% General video with moderate motion
6-8+ or custom positions Maximum smoothness +40-60% Complex transformations, high-framerate output

Each generated keyframe consumes additional latent tokens during diffusion. The absolute position embedding mechanism prevents collision with standard frame conditioning, but GPU memory scales with total keyframe count.

Summary

  • Enable via CLI with --num-generated-keyframes N for even spacing, or pass explicit frame indices
  • Validate checkpoint compatibility—the use_keyframes_abs_pos_embedding flag must be present in transformer config
  • Understand the pipeline flow: argument parsing → position resolution → VideoGeneratedKeyframeSlots construction → stage 1 conditioning injection
  • Reference key source files: args.py for CLI, helpers.py for position logic, blocks.py for capability guards, keyframe_interpolation.py for integration
  • Use manual conditioning construction when you need precise control over keyframe placement

Frequently Asked Questions

What happens if I use generated keyframes with an incompatible checkpoint?

The pipeline raises a runtime error from assert_generated_keyframes_supported before any diffusion begins. The error message specifies that your checkpoint lacks use_keyframes_abs_pos_embedding support. No silent degradation occurs—you must obtain or train a compatible checkpoint.

Can I mix user-provided keyframes with generated keyframes?

Yes. The images parameter accepts ImageConditioningInput objects with explicit frame indices, while num_generated_keyframes inserts additional anchor points. The two conditioning types coexist in stage_1_conditionings without conflict, each processed by appropriate embedding layers.

How are frame positions calculated for integer keyframe counts?

The evenly_spaced_keyframe_positions function divides the interior frame range (excluding boundaries 0 and num_frames-1) into N+1 segments, placing keyframes at segment boundaries. For 32 frames with 4 generated keyframes: positions land at frames 6, 12, 18, and 24 approximately.

Does increasing generated keyframes always improve quality?

Not universally. Beyond 6-8 keyframes for typical sequences, diminishing returns set in while compute costs rise substantially. Complex motion scenes benefit most; static or slow-moving content shows minimal improvement. Experiment with your specific content domain to find optimal settings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →