How Duration Prediction Works to Auto-Determine Frame Count in LTX-2
LTX-2 uses a lightweight regression head called DurationHead that analyzes caption encoder embeddings to predict video duration in seconds, which is then converted to a frame count that respects the VAE's temporal grid constraints.
The DurationHead enables automatic frame count determination from text prompts, eliminating manual --num-frames specification. This mechanism, available in LTX-2.5 and later checkpoints, keeps the predictor small (few megabytes) and runs on frozen connector outputs without requiring full model inference.
The Auto-Duration Request Flow
User-Level Configuration with AutoDuration
Users trigger automatic duration prediction by passing AutoDuration instead of an integer to num_frames. This dataclass accepts optional min_seconds and max_seconds bounds to constrain the prediction.
from ltx_pipelines.utils.types import AutoDuration
# Default: allow 1-20 second range
num_frames = AutoDuration()
# Constrained: force 2-15 second range
num_frames = AutoDuration(min_seconds=2.0, max_seconds=15.0)
The AutoDuration type is defined in ltx_pipelines/utils/types.py【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py#L15-L24】.
Early Validation with require_num_frames_source
Every pipeline validates the request immediately. The require_num_frames_source function in ltx_pipelines/utils/blocks.py raises a clear error if AutoDuration is requested but the checkpoint lacks a trained DurationHead (pre-2.5 checkpoints)【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L94-L100】.
This guard prevents silent failures and guides users to either upgrade checkpoints or specify frames manually.
The Duration Prediction Architecture
Step 1: Caption Encoding
The PromptEncoder (Gemma-based) processes the text caption and produces connector token tensors: video_encoding and audio_encoding. These frozen embeddings feed directly into the duration predictor—no diffusion model components execute at this stage.
Step 2: DurationPredictor Execution
The DurationPredictor class wraps the DurationHead and is instantiated via DurationPredictor.from_checkpoint【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L3-L12】.
When called, it forwards connector tensors to the head:
# From ltx_pipelines/utils/blocks.py, lines 50-71
def __call__(
self,
*,
video_encoding: Optional[torch.Tensor] = None,
audio_encoding: Optional[torch.Tensor] = None,
) -> float:
# ... validation and tensor preparation ...
output = self.duration_head(
video_encoding=video_encoding,
audio_encoding=audio_encoding,
)
return output["seconds"].item()
The predictor handles both video-only and audio-video scenarios, concatenating available encodings before passing to the head.
Step 3: DurationHead Regression
The core DurationHead in ltx_core/duration_head/duration_head.py processes tokens through a learned attention pooler followed by a small MLP【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-core/src/ltx_core/duration_head/duration_head.py#L52-L118】.
Key implementation details:
- Input: Pooled token streams from connector outputs
- Regression target: Log-seconds (exponentiated to seconds for final output)
- Architecture: Attention-based pooling + 2-layer MLP with hidden dimension 256
- Output: Single scalar representing predicted duration
Log-space prediction stabilizes training across the wide range of human-perceived durations (1-20+ seconds).
Frame Count Conversion and Grid Alignment
seconds_to_clamped_num_frames
Raw seconds predictions undergo three transformations in ltx_pipelines/utils/helpers.py【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L65-L85】:
num_frames = seconds_to_clamped_num_frames(
seconds=predicted_seconds,
frame_rate=25.0, # or user-specified
min_frames=round(min_seconds * frame_rate),
max_frames=round(max_seconds * frame_rate),
)
The conversion pipeline:
- Round: Convert seconds to raw frame count (
seconds * frame_rate) - Clamp: Enforce
[min_frames, max_frames]bounds for memory safety - Snap: Align to VAE causal temporal grid via
snap_frames_to_grid
Grid Snapping with snap_frames_to_grid
The VAE requires frames satisfying (frames - 1) % 8k == 0 for temporal alignment. The snap_frames_to_grid function in ltx_pipelines/utils/helpers.py【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L54-L62】 rounds to valid grid points:
- Finds nearest valid frame count above and below
- Selects closest valid value
- Guarantees VAE compatibility without user intervention
This ensures generated videos decode correctly without temporal artifacts.
Complete Usage Examples
Standard Pipeline Usage
from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from ltx_pipelines.utils.types import AutoDuration
import torch
pipeline = TI2VidOneStagePipeline(
checkpoint_path="checkpoints/ltx_v2a_lora.ckpt",
dtype=torch.float16,
device="cuda",
# Intentionally omit num_frames to trigger auto-duration
)
result = pipeline(
prompt="A sunrise over a misty forest",
num_frames=AutoDuration(min_seconds=2.0, max_seconds=15.0),
)
print(f"Generated {result.video.shape[1]} frames at 25 FPS")
# Output: Generated 151 frames at 25 FPS (~6 seconds)
Manual DurationPredictor Usage
from ltx_pipelines.utils.blocks import DurationPredictor, resolve_num_frames
from ltx_pipelines.utils.types import AutoDuration
# Build predictor from checkpoint containing DurationHead
predictor = DurationPredictor.from_checkpoint(
checkpoint_path="checkpoints/ltx_v2a_lora.ckpt",
dtype=torch.float16,
device="cuda",
)
# Assume connector encodings from prior encoding pass
video_enc, audio_enc = get_connector_encodings(prompt)
# Resolve to concrete frame count
num_frames = resolve_num_frames(
num_frames=AutoDuration(),
duration_predictor=predictor,
video_encoding=video_enc,
audio_encoding=audio_enc,
frame_rate=25.0,
)
print(f"Predicted {num_frames} frames")
Key Implementation Files
| File | Component | Purpose |
|---|---|---|
packages/ltx-pipelines/src/ltx_pipelines/utils/types.py |
AutoDuration dataclass |
User-facing API for auto-duration requests |
packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py |
DurationPredictor, resolve_num_frames, require_num_frames_source |
Orchestration and validation layer |
packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py |
seconds_to_clamped_num_frames, snap_frames_to_grid |
Conversion and grid alignment |
packages/ltx-core/src/ltx_core/duration_head/duration_head.py |
DurationHead |
Core regression network predicting log-seconds |
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py |
CLI argument parsing | --auto-duration flag handling |
Summary
AutoDurationtriggers automatic frame count determination via theDurationPredictorsubsystemDurationHeadperforms lightweight regression on frozen caption encoder outputs to predict duration in log-secondsseconds_to_clamped_num_framesconverts predictions to concrete frame counts with memory-safe clampingsnap_frames_to_gridenforces VAE temporal alignment constraints without user intervention- Early validation prevents runtime errors when checkpoints lack trained duration heads
The design prioritizes efficiency: the predictor adds minimal memory overhead, runs once per generation request, and avoids loading or executing the full diffusion model for duration estimation.
Frequently Asked Questions
What LTX-2 checkpoints support auto-duration prediction?
Auto-duration requires checkpoints with a trained DurationHead, which debuted in LTX-2.5 releases. Pre-2.5 checkpoints lack this component, and the pipeline raises a clear error if AutoDuration is requested without an available predictor. Always use DurationPredictor.from_checkpoint to verify predictor availability or check checkpoint metadata for the duration_head key.
Why does the predicted frame count differ slightly from my requested duration?
The snap_frames_to_grid function aligns frame counts to the VAE's causal temporal requirements: (frames - 1) % 8k == 0. For LTX-2's 8-frame temporal compression, valid frame counts follow sequences like 1, 9, 17, 25, 33, etc. Your requested duration is first converted to frames, then rounded to the nearest valid grid point. This ensures artifact-free video decoding.
Can I use DurationPredictor independently of the full pipeline?
Yes. Instantiate DurationPredictor.from_checkpoint with any compatible checkpoint, then call it directly with video and/or audio connector encodings. This is useful for dataset analysis, duration-conditioned training, or building custom inference workflows. The predictor accepts both video_encoding and audio_encoding tensors and returns a float seconds value.
How does the min_seconds / max_seconds clamping interact with grid snapping?
Clamping occurs before grid snapping. The pipeline first converts your bounds to frame counts, clamps the raw prediction to this range, then snaps to valid grid points. If your bounds exclude all valid grid points (unlikely with default 1-20s ranges), the result boundaries to the nearest valid grid point within your constraints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →