How to Use the LTX-2 Duration Head for Automatic Frame Count Prediction from Text Prompts

The LTX-2 Duration Head is a lightweight regression model that predicts video duration in seconds directly from encoded text prompts, automatically converting that duration into a frame count that complies with the VAE's temporal grid constraints.

The DurationHead in Lightricks/LTX-2 enables zero-shot video length estimation. Instead of manually specifying how many frames to generate, you can let the model infer an appropriate duration from the semantic content of your prompt. This component operates on frozen Connector token streams from the PromptEncoder, making it efficient to load and run alongside the main generation pipeline.

How the Duration Head Works

Architecture Overview

The Duration Head follows a streamlined path from text to seconds:

Prompt → TextEncoder → Connector (video-tokens / audio-tokens) → DurationHead → seconds

The DurationHead class in ltx_core/duration_head/duration_head.py performs three key operations:

  1. Projects modality-specific token streams (video or audio) into a shared hidden space
  2. Prepends a learnable modality token to distinguish video from audio encoding
  3. Pools the sequence with a single-query AttentionPooler, then passes the result through a tiny two-layer MLP

The MLP predicts log-duration (seconds). The forward pass exponentiates this value, ensuring callers always receive a positive duration in seconds.

From Seconds to Valid Frame Counts

Raw predictions are rarely usable directly. The DurationPredictor wrapper handles:

  • Clamping to user-specified min_seconds and max_seconds bounds
  • Conversion to frame count using the target frame_rate
  • Snapping to the VAE's causal temporal grid via (frames-1) % time_scale == 0

This conversion logic lives in ltx_pipelines/utils/helpers.py within the seconds_to_clamped_num_frames function.

Core Components and File Locations

Role Module Key Class / Function
Regression head ltx_core/duration_head/duration_head.py DurationHead
Builder / loader ltx_core/duration_head/model_configurator.py DurationHeadConfigurator
Pipeline glue ltx_pipelines/utils/blocks.py DurationPredictor, require_num_frames_source, resolve_num_frames
Seconds-to-frames conversion ltx_pipelines/utils/helpers.py seconds_to_clamped_num_frames

Automatic Frame Count Prediction in Practice

High-Level Pipeline Usage

The simplest way to use automatic duration prediction is to omit the num_frames argument when calling a pipeline like TI2VidOneStagePipeline:

from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from pathlib import Path
import torch

# Build the pipeline; checkpoint must include a DurationHead

pipeline = TI2VidOneStagePipeline.from_checkpoint(
    checkpoint_path=Path("checkpoints/ltx_v2_gemma4.pt"),
    dtype=torch.float16,
    device=torch.device("cuda"),
)

# No num_frames specified → automatic prediction from prompt

video, audio = pipeline(
    prompt="A sunrise over a misty forest, with birdsong in the background",
)

Internally, the pipeline executes this flow in TI2VidOneStagePipeline.__call__:


# Duration predictor is built once during pipeline initialization

self.duration_predictor = DurationPredictor.from_checkpoint(...)

# Fast-fail if auto-duration requested but head unavailable

require_num_frames_source(num_frames, self.duration_predictor)

# Encode prompt to get connector tokens

video_enc, audio_enc = self.encode_prompt(prompt)

# Resolve final frame count

num_frames = resolve_num_frames(
    num_frames, self.duration_predictor,
    video_encoding=video_enc,
    audio_encoding=audio_enc,
    frame_rate=self.frame_rate,
)

Direct DurationPredictor Usage

For custom workflows, instantiate DurationPredictor directly from ltx_pipelines/utils/blocks.py:

import torch
from ltx_pipelines.utils.blocks import DurationPredictor

# Load from checkpoint containing DurationHead weights

predictor = DurationPredictor.from_checkpoint(
    checkpoint_path="checkpoints/ltx_v2_gemma4.pt",
    dtype=torch.float16,
    device=torch.device("cuda"),
)

# Connector outputs from your PromptEncoder

video_tokens = torch.randn(1, 128, 4096, device="cuda")   # (B, Tv, Dv)

audio_tokens = torch.randn(1, 64, 2048, device="cuda")    # (B, Ta, Da)

# Predict frame count with bounds

frames = predictor(
    video_encoding=video_tokens,
    audio_encoding=audio_tokens,
    frame_rate=30.0,
    min_seconds=2.0,
    max_seconds=15.0,
)

print(f"Predicted length: {frames} frames ({frames/30:.2f}s)")

Manual Duration Clamping and Grid Snapping

If you need fine-grained control over the conversion pipeline, use the helper functions directly:

from ltx_pipelines.utils.helpers import seconds_to_clamped_num_frames

seconds = 7.84  # Raw prediction from DurationHead

frames = seconds_to_clamped_num_frames(
    seconds,
    frame_rate=30.0,
    min_frames=30,    # 1 second minimum at 30 fps

    max_frames=480,   # 16 second maximum

)

print(frames)  # → 241 frames (snapped to 8-frame VAE grid)

Backward Compatibility and Error Handling

Checkpoints predating LTX-2.5 lack duration_head.* weights. In this case, DurationPredictor.from_checkpoint returns None, and require_num_frames_source raises an early error to prevent silent failures:

from ltx_pipelines.utils.blocks import require_num_frames_source

# This raises an error if predictor is None and num_frames is AutoDuration

require_num_frames_source(num_frames, duration_predictor)

Always verify your checkpoint includes duration head weights before attempting automatic frame count prediction.

Summary

  • The DurationHead is a lightweight regression model (~few MB) that predicts video duration from Connector token streams
  • DurationPredictor wraps the head with checkpoint loading, bounds clamping, and frame rate conversion
  • Omit num_frames in pipeline calls to trigger automatic prediction, or use DurationPredictor directly for custom workflows
  • Predictions are automatically clamped and snapped to the VAE's temporal grid constraints
  • Pre-LTX-2.5 checkpoints lack duration weights; DurationPredictor.from_checkpoint returns None in these cases

Frequently Asked Questions

What happens if my checkpoint doesn't have a Duration Head?

DurationPredictor.from_checkpoint returns None, and any pipeline request for automatic duration will raise an error via require_num_frames_source. You must either upgrade to a checkpoint with duration head weights or manually specify num_frames.

Can I use the Duration Head for audio-only generation?

Yes. The DurationHead accepts both video_encoding and audio_encoding tensors, and uses a learnable modality token to distinguish between them. Pass only audio_encoding if your generation target is audio.

How does the prediction snap to the VAE grid?

The seconds_to_clamped_num_frames function in ltx_pipelines/utils/helpers.py computes the raw frame count from seconds and frame rate, then adjusts it to satisfy (frames-1) % time_scale == 0, ensuring compatibility with the causal temporal structure of LTX-2's VAE.

What frame rate should I specify for the predictor?

Use the same frame rate your VAE will operate at during generation—typically 24, 25, or 30 fps. The predictor converts its seconds output to frames using this rate, so mismatches will cause desynchronization between predicted duration and actual output length.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →