# How to Use the LTX-2 Duration Head for Automatic Frame Count Prediction from Text Prompts

> Learn how the LTX-2 Duration Head automatically predicts video frame counts from text prompts. This lightweight model converts text duration into VAE-compliant frame counts.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-18

---

**The LTX-2 Duration Head is a lightweight regression model that predicts video duration in seconds directly from encoded text prompts, automatically converting that duration into a frame count that complies with the VAE's temporal grid constraints.**

The **DurationHead** in Lightricks/LTX-2 enables zero-shot video length estimation. Instead of manually specifying how many frames to generate, you can let the model infer an appropriate duration from the semantic content of your prompt. This component operates on frozen **Connector** token streams from the **PromptEncoder**, making it efficient to load and run alongside the main generation pipeline.

## How the Duration Head Works

### Architecture Overview

The Duration Head follows a streamlined path from text to seconds:

```text
Prompt → TextEncoder → Connector (video-tokens / audio-tokens) → DurationHead → seconds

```

The `DurationHead` class in [`ltx_core/duration_head/duration_head.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/duration_head/duration_head.py) performs three key operations:

1. **Projects** modality-specific token streams (video or audio) into a shared hidden space
2. **Prepends** a learnable modality token to distinguish video from audio encoding
3. **Pools** the sequence with a single-query **AttentionPooler**, then passes the result through a tiny two-layer MLP

The MLP predicts **log-duration (seconds)**. The forward pass exponentiates this value, ensuring callers always receive a positive duration in seconds.

### From Seconds to Valid Frame Counts

Raw predictions are rarely usable directly. The `DurationPredictor` wrapper handles:

- **Clamping** to user-specified `min_seconds` and `max_seconds` bounds
- **Conversion** to frame count using the target `frame_rate`
- **Snapping** to the VAE's causal temporal grid via `(frames-1) % time_scale == 0`

This conversion logic lives in [`ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/helpers.py) within the `seconds_to_clamped_num_frames` function.

## Core Components and File Locations

| Role | Module | Key Class / Function |
|------|--------|----------------------|
| Regression head | [`ltx_core/duration_head/duration_head.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/duration_head/duration_head.py) | `DurationHead` |
| Builder / loader | [`ltx_core/duration_head/model_configurator.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/duration_head/model_configurator.py) | `DurationHeadConfigurator` |
| Pipeline glue | [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py) | `DurationPredictor`, `require_num_frames_source`, `resolve_num_frames` |
| Seconds-to-frames conversion | [`ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/helpers.py) | `seconds_to_clamped_num_frames` |

## Automatic Frame Count Prediction in Practice

### High-Level Pipeline Usage

The simplest way to use automatic duration prediction is to omit the `num_frames` argument when calling a pipeline like `TI2VidOneStagePipeline`:

```python
from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from pathlib import Path
import torch

# Build the pipeline; checkpoint must include a DurationHead

pipeline = TI2VidOneStagePipeline.from_checkpoint(
    checkpoint_path=Path("checkpoints/ltx_v2_gemma4.pt"),
    dtype=torch.float16,
    device=torch.device("cuda"),
)

# No num_frames specified → automatic prediction from prompt

video, audio = pipeline(
    prompt="A sunrise over a misty forest, with birdsong in the background",
)

```

Internally, the pipeline executes this flow in `TI2VidOneStagePipeline.__call__`:

```python

# Duration predictor is built once during pipeline initialization

self.duration_predictor = DurationPredictor.from_checkpoint(...)

# Fast-fail if auto-duration requested but head unavailable

require_num_frames_source(num_frames, self.duration_predictor)

# Encode prompt to get connector tokens

video_enc, audio_enc = self.encode_prompt(prompt)

# Resolve final frame count

num_frames = resolve_num_frames(
    num_frames, self.duration_predictor,
    video_encoding=video_enc,
    audio_encoding=audio_enc,
    frame_rate=self.frame_rate,
)

```

### Direct DurationPredictor Usage

For custom workflows, instantiate `DurationPredictor` directly from [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py):

```python
import torch
from ltx_pipelines.utils.blocks import DurationPredictor

# Load from checkpoint containing DurationHead weights

predictor = DurationPredictor.from_checkpoint(
    checkpoint_path="checkpoints/ltx_v2_gemma4.pt",
    dtype=torch.float16,
    device=torch.device("cuda"),
)

# Connector outputs from your PromptEncoder

video_tokens = torch.randn(1, 128, 4096, device="cuda")   # (B, Tv, Dv)

audio_tokens = torch.randn(1, 64, 2048, device="cuda")    # (B, Ta, Da)

# Predict frame count with bounds

frames = predictor(
    video_encoding=video_tokens,
    audio_encoding=audio_tokens,
    frame_rate=30.0,
    min_seconds=2.0,
    max_seconds=15.0,
)

print(f"Predicted length: {frames} frames ({frames/30:.2f}s)")

```

### Manual Duration Clamping and Grid Snapping

If you need fine-grained control over the conversion pipeline, use the helper functions directly:

```python
from ltx_pipelines.utils.helpers import seconds_to_clamped_num_frames

seconds = 7.84  # Raw prediction from DurationHead

frames = seconds_to_clamped_num_frames(
    seconds,
    frame_rate=30.0,
    min_frames=30,    # 1 second minimum at 30 fps

    max_frames=480,   # 16 second maximum

)

print(frames)  # → 241 frames (snapped to 8-frame VAE grid)

```

## Backward Compatibility and Error Handling

Checkpoints predating LTX-2.5 lack `duration_head.*` weights. In this case, `DurationPredictor.from_checkpoint` returns `None`, and `require_num_frames_source` raises an early error to prevent silent failures:

```python
from ltx_pipelines.utils.blocks import require_num_frames_source

# This raises an error if predictor is None and num_frames is AutoDuration

require_num_frames_source(num_frames, duration_predictor)

```

Always verify your checkpoint includes duration head weights before attempting automatic frame count prediction.

## Summary

- The **DurationHead** is a lightweight regression model (~few MB) that predicts video duration from Connector token streams
- **DurationPredictor** wraps the head with checkpoint loading, bounds clamping, and frame rate conversion
- Omit `num_frames` in pipeline calls to trigger automatic prediction, or use `DurationPredictor` directly for custom workflows
- Predictions are automatically clamped and snapped to the VAE's temporal grid constraints
- Pre-LTX-2.5 checkpoints lack duration weights; `DurationPredictor.from_checkpoint` returns `None` in these cases

## Frequently Asked Questions

### What happens if my checkpoint doesn't have a Duration Head?

`DurationPredictor.from_checkpoint` returns `None`, and any pipeline request for automatic duration will raise an error via `require_num_frames_source`. You must either upgrade to a checkpoint with duration head weights or manually specify `num_frames`.

### Can I use the Duration Head for audio-only generation?

Yes. The `DurationHead` accepts both `video_encoding` and `audio_encoding` tensors, and uses a learnable modality token to distinguish between them. Pass only `audio_encoding` if your generation target is audio.

### How does the prediction snap to the VAE grid?

The `seconds_to_clamped_num_frames` function in [`ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/helpers.py) computes the raw frame count from seconds and frame rate, then adjusts it to satisfy `(frames-1) % time_scale == 0`, ensuring compatibility with the causal temporal structure of LTX-2's VAE.

### What frame rate should I specify for the predictor?

Use the same frame rate your VAE will operate at during generation—typically 24, 25, or 30 fps. The predictor converts its seconds output to frames using this rate, so mismatches will cause desynchronization between predicted duration and actual output length.