How to Use the LTX-2 Duration Head for Automatic Frame Count Prediction from Text Prompts
The LTX-2 Duration Head is a lightweight regression model that predicts video duration in seconds directly from encoded text prompts, automatically converting that duration into a frame count that complies with the VAE's temporal grid constraints.
The DurationHead in Lightricks/LTX-2 enables zero-shot video length estimation. Instead of manually specifying how many frames to generate, you can let the model infer an appropriate duration from the semantic content of your prompt. This component operates on frozen Connector token streams from the PromptEncoder, making it efficient to load and run alongside the main generation pipeline.
How the Duration Head Works
Architecture Overview
The Duration Head follows a streamlined path from text to seconds:
Prompt → TextEncoder → Connector (video-tokens / audio-tokens) → DurationHead → seconds
The DurationHead class in ltx_core/duration_head/duration_head.py performs three key operations:
- Projects modality-specific token streams (video or audio) into a shared hidden space
- Prepends a learnable modality token to distinguish video from audio encoding
- Pools the sequence with a single-query AttentionPooler, then passes the result through a tiny two-layer MLP
The MLP predicts log-duration (seconds). The forward pass exponentiates this value, ensuring callers always receive a positive duration in seconds.
From Seconds to Valid Frame Counts
Raw predictions are rarely usable directly. The DurationPredictor wrapper handles:
- Clamping to user-specified
min_secondsandmax_secondsbounds - Conversion to frame count using the target
frame_rate - Snapping to the VAE's causal temporal grid via
(frames-1) % time_scale == 0
This conversion logic lives in ltx_pipelines/utils/helpers.py within the seconds_to_clamped_num_frames function.
Core Components and File Locations
| Role | Module | Key Class / Function |
|---|---|---|
| Regression head | ltx_core/duration_head/duration_head.py |
DurationHead |
| Builder / loader | ltx_core/duration_head/model_configurator.py |
DurationHeadConfigurator |
| Pipeline glue | ltx_pipelines/utils/blocks.py |
DurationPredictor, require_num_frames_source, resolve_num_frames |
| Seconds-to-frames conversion | ltx_pipelines/utils/helpers.py |
seconds_to_clamped_num_frames |
Automatic Frame Count Prediction in Practice
High-Level Pipeline Usage
The simplest way to use automatic duration prediction is to omit the num_frames argument when calling a pipeline like TI2VidOneStagePipeline:
from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from pathlib import Path
import torch
# Build the pipeline; checkpoint must include a DurationHead
pipeline = TI2VidOneStagePipeline.from_checkpoint(
checkpoint_path=Path("checkpoints/ltx_v2_gemma4.pt"),
dtype=torch.float16,
device=torch.device("cuda"),
)
# No num_frames specified → automatic prediction from prompt
video, audio = pipeline(
prompt="A sunrise over a misty forest, with birdsong in the background",
)
Internally, the pipeline executes this flow in TI2VidOneStagePipeline.__call__:
# Duration predictor is built once during pipeline initialization
self.duration_predictor = DurationPredictor.from_checkpoint(...)
# Fast-fail if auto-duration requested but head unavailable
require_num_frames_source(num_frames, self.duration_predictor)
# Encode prompt to get connector tokens
video_enc, audio_enc = self.encode_prompt(prompt)
# Resolve final frame count
num_frames = resolve_num_frames(
num_frames, self.duration_predictor,
video_encoding=video_enc,
audio_encoding=audio_enc,
frame_rate=self.frame_rate,
)
Direct DurationPredictor Usage
For custom workflows, instantiate DurationPredictor directly from ltx_pipelines/utils/blocks.py:
import torch
from ltx_pipelines.utils.blocks import DurationPredictor
# Load from checkpoint containing DurationHead weights
predictor = DurationPredictor.from_checkpoint(
checkpoint_path="checkpoints/ltx_v2_gemma4.pt",
dtype=torch.float16,
device=torch.device("cuda"),
)
# Connector outputs from your PromptEncoder
video_tokens = torch.randn(1, 128, 4096, device="cuda") # (B, Tv, Dv)
audio_tokens = torch.randn(1, 64, 2048, device="cuda") # (B, Ta, Da)
# Predict frame count with bounds
frames = predictor(
video_encoding=video_tokens,
audio_encoding=audio_tokens,
frame_rate=30.0,
min_seconds=2.0,
max_seconds=15.0,
)
print(f"Predicted length: {frames} frames ({frames/30:.2f}s)")
Manual Duration Clamping and Grid Snapping
If you need fine-grained control over the conversion pipeline, use the helper functions directly:
from ltx_pipelines.utils.helpers import seconds_to_clamped_num_frames
seconds = 7.84 # Raw prediction from DurationHead
frames = seconds_to_clamped_num_frames(
seconds,
frame_rate=30.0,
min_frames=30, # 1 second minimum at 30 fps
max_frames=480, # 16 second maximum
)
print(frames) # → 241 frames (snapped to 8-frame VAE grid)
Backward Compatibility and Error Handling
Checkpoints predating LTX-2.5 lack duration_head.* weights. In this case, DurationPredictor.from_checkpoint returns None, and require_num_frames_source raises an early error to prevent silent failures:
from ltx_pipelines.utils.blocks import require_num_frames_source
# This raises an error if predictor is None and num_frames is AutoDuration
require_num_frames_source(num_frames, duration_predictor)
Always verify your checkpoint includes duration head weights before attempting automatic frame count prediction.
Summary
- The DurationHead is a lightweight regression model (~few MB) that predicts video duration from Connector token streams
- DurationPredictor wraps the head with checkpoint loading, bounds clamping, and frame rate conversion
- Omit
num_framesin pipeline calls to trigger automatic prediction, or useDurationPredictordirectly for custom workflows - Predictions are automatically clamped and snapped to the VAE's temporal grid constraints
- Pre-LTX-2.5 checkpoints lack duration weights;
DurationPredictor.from_checkpointreturnsNonein these cases
Frequently Asked Questions
What happens if my checkpoint doesn't have a Duration Head?
DurationPredictor.from_checkpoint returns None, and any pipeline request for automatic duration will raise an error via require_num_frames_source. You must either upgrade to a checkpoint with duration head weights or manually specify num_frames.
Can I use the Duration Head for audio-only generation?
Yes. The DurationHead accepts both video_encoding and audio_encoding tensors, and uses a learnable modality token to distinguish between them. Pass only audio_encoding if your generation target is audio.
How does the prediction snap to the VAE grid?
The seconds_to_clamped_num_frames function in ltx_pipelines/utils/helpers.py computes the raw frame count from seconds and frame rate, then adjusts it to satisfy (frames-1) % time_scale == 0, ensuring compatibility with the causal temporal structure of LTX-2's VAE.
What frame rate should I specify for the predictor?
Use the same frame rate your VAE will operate at during generation—typically 24, 25, or 30 fps. The predictor converts its seconds output to frames using this rate, so mismatches will cause desynchronization between predicted duration and actual output length.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →