Automatic Duration Prediction in LTX-2: How to Bypass Manual Frame Counts
LTX-2 eliminates manual frame counting through a DurationHead neural network that predicts shot duration from prompt embeddings, automatically converting seconds into VAE-compatible frame counts.
Video generation pipelines typically force you to specify exactly how many frames to produce. The LTX-2 open-source repository from Lightricks solves this with automatic duration prediction — a lightweight system that infers optimal video length directly from your text prompt. This guide explains the complete implementation, from the AutoDuration API to the underlying DurationPredictor mechanics.
How Automatic Duration Prediction Works in LTX-2
The system centers on a small neural head that analyzes prompt encoder outputs and outputs a duration in seconds. A wrapper class then snaps this prediction to the model's required temporal grid.
The core flow spans three stages:
- Checkpoint inspection — Detects whether a
DurationHeadis present - Prompt encoding — Processes your text into embeddings
- Frame resolution — Converts predicted seconds to valid frame counts
The VAE Temporal Grid Constraint
LTX-2's VAE requires frame counts following the pattern 8k + 1 (9, 17, 25, 33... frames). The DurationPredictor automatically rounds predictions to satisfy this constraint, preventing runtime errors.
Core Components and Source Locations
AutoDuration Dataclass
The AutoDuration dataclass in ltx_pipelines/utils/types.py represents a request for automatic prediction:
from dataclasses import dataclass
@dataclass
class AutoDuration:
min_seconds: float = 1.0
max_seconds: float = 20.0
Source: [ltx_pipelines/utils/types.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py#L15-L24)
A ready-to-use instance is exported as DEFAULT_AUTO_DURATION (line 26).
DurationPredictor Wrapper
The DurationPredictor class in ltx_pipelines/utils/blocks.py orchestrates the actual prediction:
from ltx_pipelines.utils.blocks import DurationPredictor
# Loaded from checkpoint if DurationHead weights exist
predictor = DurationPredictor.from_checkpoint(
duration_head_path="path/to/duration_head.safetensors",
device="cuda"
)
Source: [ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L3-L30)
DurationHead Neural Network
The actual prediction happens in DurationHead, a tiny network that consumes prompt encoder outputs:
# From ltx_core/duration_head/duration_head.py
class DurationHead(nn.Module):
def forward(self, video_encoding, audio_encoding=None):
# Returns raw seconds prediction
...
Source: [ltx_core/duration_head/duration_head.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/duration_head/duration_head.py#L52-L89)
Complete Implementation Pipeline
Step 1: Checkpoint Loading and Availability Detection
When instantiating any LTX-2 pipeline, the system checks ModelPaths.duration_head_path:
from ltx_pipelines.utils.model_paths import ModelPaths
paths = ModelPaths.from_checkpoint(checkpoint_path)
# paths.duration_head_path is Optional[str]
# If None, automatic duration prediction is unavailable
Source: [ltx_pipelines/utils/model_paths.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py#L33-L51)
Step 2: Guard Against Missing DurationHead
The require_num_frames_source function provides fast-fail behavior before expensive computation:
from ltx_pipelines.utils.blocks import require_num_frames_source
require_num_frames_source(
num_frames=DEFAULT_AUTO_DURATION,
duration_predictor=duration_predictor # None if checkpoint lacks DurationHead
)
# Raises ValueError with clear message if predictor is None but auto-duration requested
Source: [ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L94-L100)
Step 3: Frame Count Resolution
After prompt encoding, resolve_num_frames produces the final integer:
from ltx_pipelines.utils.blocks import resolve_num_frames
num_frames = resolve_num_frames(
num_frames=DEFAULT_AUTO_DURATION,
duration_predictor=duration_predictor,
video_encoding=video_enc, # From prompt encoder
audio_encoding=audio_enc, # Optional
frame_rate=25.0,
)
# Returns int snapped to 8k+1 grid
Source: [ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L108-L128)
Inside DurationPredictor.__call__, the logic flows through seconds_to_clamped_num_frames (called at line 76), which:
- Runs
DurationHeadforward pass → raw seconds - Clamps to
[min_seconds, max_seconds] - Converts to frames and rounds to nearest
8k + 1value
Practical Usage Examples
Basic Pipeline With Automatic Duration
import torch
from ltx_pipelines.ti2vid_one_stage import Ti2VidOneStage
from ltx_pipelines.utils.types import DEFAULT_AUTO_DURATION
pipe = Ti2VidOneStage(
checkpoint_path="path/to/checkpoint",
dtype=torch.float16,
device="cuda",
num_frames=DEFAULT_AUTO_DURATION, # Triggers automatic prediction
)
output = pipe(prompt="A drone flying over a coastal cliff at sunset")
# Frame count determined automatically from prompt
Custom Duration Window
from ltx_pipelines.utils.types import AutoDuration
# Request 3-10 second range instead of default 1-20
custom_duration = AutoDuration(min_seconds=3.0, max_seconds=10.0)
pipe = Ti2VidOneStage(
checkpoint_path="path/to/checkpoint",
num_frames=custom_duration,
)
Graceful Degradation for Missing DurationHead
from ltx_pipelines.utils.blocks import require_num_frames_source, DEFAULT_AUTO_DURATION
def safe_auto_duration(checkpoint_path, num_frames):
from ltx_pipelines.utils.blocks import DurationPredictor
predictor = DurationPredictor.from_checkpoint(
checkpoint_path, device="cuda"
)
try:
require_num_frames_source(num_frames, predictor)
return predictor
except ValueError as e:
print(f"Falling back to manual 121 frames: {e}")
return None # Caller should use fixed num_frames
# Usage
predictor = safe_auto_duration("path/to/checkpoint", DEFAULT_AUTO_DURATION)
Integration Points in Pipeline Classes
Pipeline implementations like Ti2VidOneStage wire these components together. The typical call chain:
__init__storesnum_framesparameter (can beintorAutoDuration)__call__receives prompt, encodes itresolve_num_framesconverts to concrete frame count- Generation proceeds with determined length
Source: [ti2vid_one_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_one_stage.py) shows this pattern in the constructor (around line 144) and forward logic.
Summary
AutoDuration— Dataclass signaling automatic frame count selection with configurable min/max boundsDurationPredictor— Wrapper managingDurationHeadinference and VAE-grid alignmentDurationHead— Neural network predicting seconds from prompt encoder outputs- Guard and resolver functions —
require_num_frames_sourcefor validation,resolve_num_framesfor conversion - Checkpoint-based availability — Auto-duration only works when
duration_head_pathexists in the model checkpoint
Frequently Asked Questions
What happens if I request automatic duration but my checkpoint lacks a DurationHead?
The require_num_frames_source guard raises a ValueError with a clear message before any expensive computation begins. This fast-fail behavior prevents wasted GPU cycles. Always verify your checkpoint includes duration_head_path or handle the exception gracefully.
How does LTX-2 ensure predicted frame counts work with the VAE?
The DurationPredictor rounds all predictions to the nearest valid 8k + 1 frame count through seconds_to_clamped_num_frames. This guarantees compatibility with the causal temporal grid required by the VAE architecture, eliminating shape mismatches during encoding/decoding.
Can I constrain the predicted duration to a specific range?
Yes. Instantiate AutoDuration(min_seconds=2.0, max_seconds=8.0) with your desired bounds instead of using DEFAULT_AUTO_DURATION. The predictor clamps raw network outputs to this window before frame conversion, giving you control over minimum and maximum clip lengths.
Is the DurationHead prediction deterministic?
The DurationHead forward pass is deterministic given fixed prompt encoder outputs. However, if your pipeline uses stochastic prompt augmentation or varying random seeds, the input embeddings—and thus the predicted duration—may vary across runs with identical text prompts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →