Automatic Duration Prediction in LTX-2: How to Bypass Manual Frame Counts

LTX-2 eliminates manual frame counting through a DurationHead neural network that predicts shot duration from prompt embeddings, automatically converting seconds into VAE-compatible frame counts.

Video generation pipelines typically force you to specify exactly how many frames to produce. The LTX-2 open-source repository from Lightricks solves this with automatic duration prediction — a lightweight system that infers optimal video length directly from your text prompt. This guide explains the complete implementation, from the AutoDuration API to the underlying DurationPredictor mechanics.

How Automatic Duration Prediction Works in LTX-2

The system centers on a small neural head that analyzes prompt encoder outputs and outputs a duration in seconds. A wrapper class then snaps this prediction to the model's required temporal grid.

The core flow spans three stages:

  1. Checkpoint inspection — Detects whether a DurationHead is present
  2. Prompt encoding — Processes your text into embeddings
  3. Frame resolution — Converts predicted seconds to valid frame counts

The VAE Temporal Grid Constraint

LTX-2's VAE requires frame counts following the pattern 8k + 1 (9, 17, 25, 33... frames). The DurationPredictor automatically rounds predictions to satisfy this constraint, preventing runtime errors.

Core Components and Source Locations

AutoDuration Dataclass

The AutoDuration dataclass in ltx_pipelines/utils/types.py represents a request for automatic prediction:

from dataclasses import dataclass

@dataclass
class AutoDuration:
    min_seconds: float = 1.0
    max_seconds: float = 20.0

Source: [ltx_pipelines/utils/types.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py#L15-L24)

A ready-to-use instance is exported as DEFAULT_AUTO_DURATION (line 26).

DurationPredictor Wrapper

The DurationPredictor class in ltx_pipelines/utils/blocks.py orchestrates the actual prediction:

from ltx_pipelines.utils.blocks import DurationPredictor

# Loaded from checkpoint if DurationHead weights exist

predictor = DurationPredictor.from_checkpoint(
    duration_head_path="path/to/duration_head.safetensors",
    device="cuda"
)

Source: [ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L3-L30)

DurationHead Neural Network

The actual prediction happens in DurationHead, a tiny network that consumes prompt encoder outputs:


# From ltx_core/duration_head/duration_head.py

class DurationHead(nn.Module):
    def forward(self, video_encoding, audio_encoding=None):
        # Returns raw seconds prediction

        ...

Source: [ltx_core/duration_head/duration_head.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/duration_head/duration_head.py#L52-L89)

Complete Implementation Pipeline

Step 1: Checkpoint Loading and Availability Detection

When instantiating any LTX-2 pipeline, the system checks ModelPaths.duration_head_path:

from ltx_pipelines.utils.model_paths import ModelPaths

paths = ModelPaths.from_checkpoint(checkpoint_path)

# paths.duration_head_path is Optional[str]

# If None, automatic duration prediction is unavailable

Source: [ltx_pipelines/utils/model_paths.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py#L33-L51)

Step 2: Guard Against Missing DurationHead

The require_num_frames_source function provides fast-fail behavior before expensive computation:

from ltx_pipelines.utils.blocks import require_num_frames_source

require_num_frames_source(
    num_frames=DEFAULT_AUTO_DURATION,
    duration_predictor=duration_predictor  # None if checkpoint lacks DurationHead

)

# Raises ValueError with clear message if predictor is None but auto-duration requested

Source: [ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L94-L100)

Step 3: Frame Count Resolution

After prompt encoding, resolve_num_frames produces the final integer:

from ltx_pipelines.utils.blocks import resolve_num_frames

num_frames = resolve_num_frames(
    num_frames=DEFAULT_AUTO_DURATION,
    duration_predictor=duration_predictor,
    video_encoding=video_enc,      # From prompt encoder

    audio_encoding=audio_enc,      # Optional

    frame_rate=25.0,
)

# Returns int snapped to 8k+1 grid

Source: [ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L108-L128)

Inside DurationPredictor.__call__, the logic flows through seconds_to_clamped_num_frames (called at line 76), which:

  • Runs DurationHead forward pass → raw seconds
  • Clamps to [min_seconds, max_seconds]
  • Converts to frames and rounds to nearest 8k + 1 value

Practical Usage Examples

Basic Pipeline With Automatic Duration

import torch
from ltx_pipelines.ti2vid_one_stage import Ti2VidOneStage
from ltx_pipelines.utils.types import DEFAULT_AUTO_DURATION

pipe = Ti2VidOneStage(
    checkpoint_path="path/to/checkpoint",
    dtype=torch.float16,
    device="cuda",
    num_frames=DEFAULT_AUTO_DURATION,  # Triggers automatic prediction

)

output = pipe(prompt="A drone flying over a coastal cliff at sunset")

# Frame count determined automatically from prompt

Custom Duration Window

from ltx_pipelines.utils.types import AutoDuration

# Request 3-10 second range instead of default 1-20

custom_duration = AutoDuration(min_seconds=3.0, max_seconds=10.0)

pipe = Ti2VidOneStage(
    checkpoint_path="path/to/checkpoint",
    num_frames=custom_duration,
)

Graceful Degradation for Missing DurationHead

from ltx_pipelines.utils.blocks import require_num_frames_source, DEFAULT_AUTO_DURATION

def safe_auto_duration(checkpoint_path, num_frames):
    from ltx_pipelines.utils.blocks import DurationPredictor
    
    predictor = DurationPredictor.from_checkpoint(
        checkpoint_path, device="cuda"
    )
    
    try:
        require_num_frames_source(num_frames, predictor)
        return predictor
    except ValueError as e:
        print(f"Falling back to manual 121 frames: {e}")
        return None  # Caller should use fixed num_frames

# Usage

predictor = safe_auto_duration("path/to/checkpoint", DEFAULT_AUTO_DURATION)

Integration Points in Pipeline Classes

Pipeline implementations like Ti2VidOneStage wire these components together. The typical call chain:

  1. __init__ stores num_frames parameter (can be int or AutoDuration)
  2. __call__ receives prompt, encodes it
  3. resolve_num_frames converts to concrete frame count
  4. Generation proceeds with determined length

Source: [ti2vid_one_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_one_stage.py) shows this pattern in the constructor (around line 144) and forward logic.

Summary

  • AutoDuration — Dataclass signaling automatic frame count selection with configurable min/max bounds
  • DurationPredictor — Wrapper managing DurationHead inference and VAE-grid alignment
  • DurationHead — Neural network predicting seconds from prompt encoder outputs
  • Guard and resolver functions — require_num_frames_source for validation, resolve_num_frames for conversion
  • Checkpoint-based availability — Auto-duration only works when duration_head_path exists in the model checkpoint

Frequently Asked Questions

What happens if I request automatic duration but my checkpoint lacks a DurationHead?

The require_num_frames_source guard raises a ValueError with a clear message before any expensive computation begins. This fast-fail behavior prevents wasted GPU cycles. Always verify your checkpoint includes duration_head_path or handle the exception gracefully.

How does LTX-2 ensure predicted frame counts work with the VAE?

The DurationPredictor rounds all predictions to the nearest valid 8k + 1 frame count through seconds_to_clamped_num_frames. This guarantees compatibility with the causal temporal grid required by the VAE architecture, eliminating shape mismatches during encoding/decoding.

Can I constrain the predicted duration to a specific range?

Yes. Instantiate AutoDuration(min_seconds=2.0, max_seconds=8.0) with your desired bounds instead of using DEFAULT_AUTO_DURATION. The predictor clamps raw network outputs to this window before frame conversion, giving you control over minimum and maximum clip lengths.

Is the DurationHead prediction deterministic?

The DurationHead forward pass is deterministic given fixed prompt encoder outputs. However, if your pipeline uses stochastic prompt augmentation or varying random seeds, the input embeddings—and thus the predicted duration—may vary across runs with identical text prompts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →