How Generated Keyframes Work in LTX-2 and Their Token Cost Explained

LTX-2 inserts generated keyframes as extra latent tokens that provide conditioning frames for the diffusion model, with each keyframe adding approximately one token to the sequence, calculated as N / num_latent_frames.

LTX-2 supports generated keyframes—intermediate conditioning frames that guide the video diffusion model toward better temporal consistency. These keyframes are not part of the original video but are synthesized or specified positions in the latent sequence where the model receives additional guidance. Understanding how generated keyframes work in LTX-2 and their associated token cost is essential for optimizing memory usage and inference performance.

What Are Generated Keyframes in LTX-2?

Generated keyframes serve as anchor points during video generation. Unlike input keyframes derived from source video, generated keyframes are latent positions where the model applies extra conditioning signals. This allows users to influence motion patterns and scene transitions without providing actual frame images at those timestamps.

The pipeline accepts generated keyframes through two interface styles:

  1. Integer count – Automatically places N evenly-spaced interior keyframes
  2. Explicit list – Specifies exact frame indices for conditioning

Configuring Generated Keyframes

CLI Arguments

The --num-generated-keyframes argument in ltx_pipelines/utils/args.py provides the primary interface:

python -m ltx_pipelines.ti2vid_two_stages \
    --input video.mp4 \
    --output out.mp4 \
    --num-generated-keyframes 3

The help text in [args.py (lines 834-842)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py#L834-L842) explicitly documents the token cost relationship: each generated keyframe relaxes the effective temporal constraint by one pixel frame.

Python API

For programmatic control, use the helper utilities:

from ltx_pipelines.utils.helpers import (
    generated_keyframe_conditionings,
    resolve_generated_keyframes
)

# Resolve positions for 2 interior keyframes across 16 latent frames

positions = resolve_generated_keyframes(2, num_frames=16)

# Result: [5, 10] (evenly spaced, excluding first/last frames)

# Build conditioning objects for the pipeline

cond = generated_keyframe_conditionings(
    num_keyframes=2,
    num_frames=16
)

# Returns list of VideoConditionByKeyframeIndex objects

Position Normalization and Validation

The [helpers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) module handles all keyframe position logic through two primary paths.

Automatic Spacing (Integer Input)

When an integer N is provided, evenly_spaced_keyframe_positions computes interior positions:

  • Excludes frame 0 (first) and frame num_frames-1 (last)
  • Distributes remaining keyframes with equal spacing
  • Handles edge cases where N >= num_frames - 2

Explicit Positions (List Input)

When frame indices are provided directly, resolve_generated_keyframes validates and sorts the sequence, ensuring all indices fall within valid bounds.

Checkpoint Compatibility Requirements

Generated keyframes require specialized model support. The pipeline enforces this in [blocks.py (lines 395-418)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L395-L418):


# Simplified logic from blocks.py

if not model_config.get("use_keyframes_abs_pos_embedding"):
    raise ValueError(
        "Checkpoint does not support keyframe absolute position embeddings. "
        "Generated keyframes cannot be used with this model."
    )

The keyframe absolute-position embedding (use_keyframes_abs_pos_embedding) is a training-time flag that enables the transformer to distinguish keyframe positions from regular video tokens. Without this embedding, the model cannot interpret the injected conditioning signals.

Mask Propagation During Diffusion

During the denoising loop, [denoisers.py (lines 54-56)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py#L54-L56) manages how generated keyframes interact with the attention mechanism:


# From denoisers.py - keyframes_mask handling

if keyframes_mask is not None:
    # Repeats mask for generated keyframe slots

    extended_mask = repeat_keyframes_mask_for_generated(...)

The keyframes_mask (tracked in state.keyframes_mask) marks which tokens receive special attention treatment. For generated keyframes, this mask is extended to include the new latent slots, ensuring the model recognizes them as conditioning anchors rather than standard video content.

Token Cost Calculation

The token cost of generated keyframes follows a direct relationship with sequence length:

Formula: token_increase ≈ N / num_latent_frames

Where:

  • N = number of generated keyframes
  • num_latent_frames = total latent frames in the video (typically video_length / 16 for 16× VAE compression)

Practical Examples

Video Length Latent Frames Generated Keyframes Token Overhead Relative Cost
64 frames 4 1 0.25 +25%
128 frames 8 2 0.25 +25%
256 frames 16 4 0.25 +25%
256 frames 16 8 0.50 +50%

The cost scales linearly with keyframe count but inversely with video duration. Short videos experience higher relative overhead per keyframe.

Memory and Latency Impact

Each additional token affects:

  • GPU memory – Attention matrices grow quadratically with sequence length
  • Inference time – More tokens require more attention computations per diffusion step
  • Activation caching – Keyframe conditionings add persistent tensors to the forward pass

Advanced Configuration: Image-Conditioned Keyframes

The generated_keyframe_conditionings function constructs the actual conditioning tensors:


# Core structure from helpers.py

def generated_keyframe_conditionings(num_keyframes, num_frames, images=None):
    positions = resolve_generated_keyframes(num_keyframes, num_frames)
    conditionings = []
    for pos in positions:
        cond = VideoConditionByKeyframeIndex(
            frame_index=pos,
            image=images[i] if images else None,
            strength=1.0  # configurable conditioning strength

        )
        conditionings.append(cond)
    return conditionings

When images is provided, the keyframes become image-conditioned—the model receives actual pixel guidance at those positions. When omitted, they function as structural anchors without explicit image content.

Summary

  • Generated keyframes are latent conditioning slots that improve temporal consistency in LTX-2 video generation
  • Configuration occurs via --num-generated-keyframes (CLI) or generated_keyframe_conditionings() (Python)
  • Checkpoint compatibility requires use_keyframes_abs_pos_embedding=True in the model config
  • Token cost equals N / num_latent_frames, adding memory and compute overhead proportional to keyframe count
  • The keyframes_mask propagates through [denoisers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py) to mark conditioning positions during diffusion

Frequently Asked Questions

How do I know if my LTX-2 checkpoint supports generated keyframes?

Check the model configuration for use_keyframes_abs_pos_embedding: true. The pipeline in [blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L395-L418) validates this flag and raises a clear error if missing. Most official LTX-2 releases from mid-2024 onward include this embedding.

Can I mix generated keyframes with input video keyframes?

Yes. Generated keyframes operate independently from input frame conditioning. You can specify both --keyframes (from source video) and --num-generated-keyframes (synthetic positions) simultaneously. The pipeline merges both conditioning types into the final keyframes_mask.

Why does adding keyframes increase memory usage more than expected?

The token cost formula underestimates total impact because attention complexity grows quadratically with sequence length. While raw tokens add linearly, the transformer self-attention layers process O(n²) relationships. For memory-constrained inference, prefer fewer generated keyframes on shorter video segments.

What is the maximum number of generated keyframes I should use?

Practical limits depend on your GPU memory and video length. A conservative rule: keep generated keyframes below 25% of latent frames (e.g., maximum 4 keyframes for a 16-frame latent sequence). Beyond this ratio, diminishing returns in quality are outweighed by significant latency increases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →