How Duration Prediction Works to Auto-Determine Frame Count in LTX-2

LTX-2 uses a lightweight regression head called DurationHead that analyzes caption encoder embeddings to predict video duration in seconds, which is then converted to a frame count that respects the VAE's temporal grid constraints.

The DurationHead enables automatic frame count determination from text prompts, eliminating manual --num-frames specification. This mechanism, available in LTX-2.5 and later checkpoints, keeps the predictor small (few megabytes) and runs on frozen connector outputs without requiring full model inference.

The Auto-Duration Request Flow

User-Level Configuration with AutoDuration

Users trigger automatic duration prediction by passing AutoDuration instead of an integer to num_frames. This dataclass accepts optional min_seconds and max_seconds bounds to constrain the prediction.

from ltx_pipelines.utils.types import AutoDuration

# Default: allow 1-20 second range

num_frames = AutoDuration()

# Constrained: force 2-15 second range

num_frames = AutoDuration(min_seconds=2.0, max_seconds=15.0)

The AutoDuration type is defined in ltx_pipelines/utils/types.py【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py#L15-L24】.

Early Validation with require_num_frames_source

Every pipeline validates the request immediately. The require_num_frames_source function in ltx_pipelines/utils/blocks.py raises a clear error if AutoDuration is requested but the checkpoint lacks a trained DurationHead (pre-2.5 checkpoints)【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L94-L100】.

This guard prevents silent failures and guides users to either upgrade checkpoints or specify frames manually.

The Duration Prediction Architecture

Step 1: Caption Encoding

The PromptEncoder (Gemma-based) processes the text caption and produces connector token tensors: video_encoding and audio_encoding. These frozen embeddings feed directly into the duration predictor—no diffusion model components execute at this stage.

Step 2: DurationPredictor Execution

The DurationPredictor class wraps the DurationHead and is instantiated via DurationPredictor.from_checkpoint【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L3-L12】.

When called, it forwards connector tensors to the head:


# From ltx_pipelines/utils/blocks.py, lines 50-71

def __call__(
    self,
    *,
    video_encoding: Optional[torch.Tensor] = None,
    audio_encoding: Optional[torch.Tensor] = None,
) -> float:
    # ... validation and tensor preparation ...

    output = self.duration_head(
        video_encoding=video_encoding,
        audio_encoding=audio_encoding,
    )
    return output["seconds"].item()

The predictor handles both video-only and audio-video scenarios, concatenating available encodings before passing to the head.

Step 3: DurationHead Regression

The core DurationHead in ltx_core/duration_head/duration_head.py processes tokens through a learned attention pooler followed by a small MLP【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-core/src/ltx_core/duration_head/duration_head.py#L52-L118】.

Key implementation details:

  • Input: Pooled token streams from connector outputs
  • Regression target: Log-seconds (exponentiated to seconds for final output)
  • Architecture: Attention-based pooling + 2-layer MLP with hidden dimension 256
  • Output: Single scalar representing predicted duration

Log-space prediction stabilizes training across the wide range of human-perceived durations (1-20+ seconds).

Frame Count Conversion and Grid Alignment

seconds_to_clamped_num_frames

Raw seconds predictions undergo three transformations in ltx_pipelines/utils/helpers.py【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L65-L85】:

num_frames = seconds_to_clamped_num_frames(
    seconds=predicted_seconds,
    frame_rate=25.0,  # or user-specified

    min_frames=round(min_seconds * frame_rate),
    max_frames=round(max_seconds * frame_rate),
)

The conversion pipeline:

  1. Round: Convert seconds to raw frame count (seconds * frame_rate)
  2. Clamp: Enforce [min_frames, max_frames] bounds for memory safety
  3. Snap: Align to VAE causal temporal grid via snap_frames_to_grid

Grid Snapping with snap_frames_to_grid

The VAE requires frames satisfying (frames - 1) % 8k == 0 for temporal alignment. The snap_frames_to_grid function in ltx_pipelines/utils/helpers.py【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L54-L62】 rounds to valid grid points:

  • Finds nearest valid frame count above and below
  • Selects closest valid value
  • Guarantees VAE compatibility without user intervention

This ensures generated videos decode correctly without temporal artifacts.

Complete Usage Examples

Standard Pipeline Usage

from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from ltx_pipelines.utils.types import AutoDuration
import torch

pipeline = TI2VidOneStagePipeline(
    checkpoint_path="checkpoints/ltx_v2a_lora.ckpt",
    dtype=torch.float16,
    device="cuda",
    # Intentionally omit num_frames to trigger auto-duration

)

result = pipeline(
    prompt="A sunrise over a misty forest",
    num_frames=AutoDuration(min_seconds=2.0, max_seconds=15.0),
)

print(f"Generated {result.video.shape[1]} frames at 25 FPS")

# Output: Generated 151 frames at 25 FPS (~6 seconds)

Manual DurationPredictor Usage

from ltx_pipelines.utils.blocks import DurationPredictor, resolve_num_frames
from ltx_pipelines.utils.types import AutoDuration

# Build predictor from checkpoint containing DurationHead

predictor = DurationPredictor.from_checkpoint(
    checkpoint_path="checkpoints/ltx_v2a_lora.ckpt",
    dtype=torch.float16,
    device="cuda",
)

# Assume connector encodings from prior encoding pass

video_enc, audio_enc = get_connector_encodings(prompt)

# Resolve to concrete frame count

num_frames = resolve_num_frames(
    num_frames=AutoDuration(),
    duration_predictor=predictor,
    video_encoding=video_enc,
    audio_encoding=audio_enc,
    frame_rate=25.0,
)

print(f"Predicted {num_frames} frames")

Key Implementation Files

File Component Purpose
packages/ltx-pipelines/src/ltx_pipelines/utils/types.py AutoDuration dataclass User-facing API for auto-duration requests
packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py DurationPredictor, resolve_num_frames, require_num_frames_source Orchestration and validation layer
packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py seconds_to_clamped_num_frames, snap_frames_to_grid Conversion and grid alignment
packages/ltx-core/src/ltx_core/duration_head/duration_head.py DurationHead Core regression network predicting log-seconds
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py CLI argument parsing --auto-duration flag handling

Summary

  • AutoDuration triggers automatic frame count determination via the DurationPredictor subsystem
  • DurationHead performs lightweight regression on frozen caption encoder outputs to predict duration in log-seconds
  • seconds_to_clamped_num_frames converts predictions to concrete frame counts with memory-safe clamping
  • snap_frames_to_grid enforces VAE temporal alignment constraints without user intervention
  • Early validation prevents runtime errors when checkpoints lack trained duration heads

The design prioritizes efficiency: the predictor adds minimal memory overhead, runs once per generation request, and avoids loading or executing the full diffusion model for duration estimation.

Frequently Asked Questions

What LTX-2 checkpoints support auto-duration prediction?

Auto-duration requires checkpoints with a trained DurationHead, which debuted in LTX-2.5 releases. Pre-2.5 checkpoints lack this component, and the pipeline raises a clear error if AutoDuration is requested without an available predictor. Always use DurationPredictor.from_checkpoint to verify predictor availability or check checkpoint metadata for the duration_head key.

Why does the predicted frame count differ slightly from my requested duration?

The snap_frames_to_grid function aligns frame counts to the VAE's causal temporal requirements: (frames - 1) % 8k == 0. For LTX-2's 8-frame temporal compression, valid frame counts follow sequences like 1, 9, 17, 25, 33, etc. Your requested duration is first converted to frames, then rounded to the nearest valid grid point. This ensures artifact-free video decoding.

Can I use DurationPredictor independently of the full pipeline?

Yes. Instantiate DurationPredictor.from_checkpoint with any compatible checkpoint, then call it directly with video and/or audio connector encodings. This is useful for dataset analysis, duration-conditioned training, or building custom inference workflows. The predictor accepts both video_encoding and audio_encoding tensors and returns a float seconds value.

How does the min_seconds / max_seconds clamping interact with grid snapping?

Clamping occurs before grid snapping. The pipeline first converts your bounds to frame counts, clamps the raw prediction to this range, then snaps to valid grid points. If your bounds exclude all valid grid points (unlikely with default 1-20s ranges), the result boundaries to the nearest valid grid point within your constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →