# Automatic Duration Prediction in LTX-2: How to Bypass Manual Frame Counts

> Discover automatic duration prediction in LTX-2. Bypass manual frame counts with the DurationHead neural network. Convert seconds to VAE-compatible frames effortlessly.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-15

---

**LTX-2 eliminates manual frame counting through a `DurationHead` neural network that predicts shot duration from prompt embeddings, automatically converting seconds into VAE-compatible frame counts.**

Video generation pipelines typically force you to specify exactly how many frames to produce. The LTX-2 open-source repository from Lightricks solves this with **automatic duration prediction** — a lightweight system that infers optimal video length directly from your text prompt. This guide explains the complete implementation, from the `AutoDuration` API to the underlying `DurationPredictor` mechanics.

## How Automatic Duration Prediction Works in LTX-2

The system centers on a small neural head that analyzes prompt encoder outputs and outputs a duration in seconds. A wrapper class then snaps this prediction to the model's required temporal grid.

The core flow spans three stages:

1. **Checkpoint inspection** — Detects whether a `DurationHead` is present
2. **Prompt encoding** — Processes your text into embeddings
3. **Frame resolution** — Converts predicted seconds to valid frame counts

### The VAE Temporal Grid Constraint

LTX-2's VAE requires frame counts following the pattern `8k + 1` (9, 17, 25, 33... frames). The `DurationPredictor` automatically rounds predictions to satisfy this constraint, preventing runtime errors.

## Core Components and Source Locations

### AutoDuration Dataclass

The `AutoDuration` dataclass in [`ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/types.py) represents a request for automatic prediction:

```python
from dataclasses import dataclass

@dataclass
class AutoDuration:
    min_seconds: float = 1.0
    max_seconds: float = 20.0

```

Source: [[`ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/types.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py#L15-L24)

A ready-to-use instance is exported as `DEFAULT_AUTO_DURATION` (line 26).

### DurationPredictor Wrapper

The `DurationPredictor` class in [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py) orchestrates the actual prediction:

```python
from ltx_pipelines.utils.blocks import DurationPredictor

# Loaded from checkpoint if DurationHead weights exist

predictor = DurationPredictor.from_checkpoint(
    duration_head_path="path/to/duration_head.safetensors",
    device="cuda"
)

```

Source: [[`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L3-L30)

### DurationHead Neural Network

The actual prediction happens in `DurationHead`, a tiny network that consumes prompt encoder outputs:

```python

# From ltx_core/duration_head/duration_head.py

class DurationHead(nn.Module):
    def forward(self, video_encoding, audio_encoding=None):
        # Returns raw seconds prediction

        ...

```

Source: [[`ltx_core/duration_head/duration_head.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/duration_head/duration_head.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/duration_head/duration_head.py#L52-L89)

## Complete Implementation Pipeline

### Step 1: Checkpoint Loading and Availability Detection

When instantiating any LTX-2 pipeline, the system checks `ModelPaths.duration_head_path`:

```python
from ltx_pipelines.utils.model_paths import ModelPaths

paths = ModelPaths.from_checkpoint(checkpoint_path)

# paths.duration_head_path is Optional[str]

# If None, automatic duration prediction is unavailable

```

Source: [[`ltx_pipelines/utils/model_paths.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/model_paths.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py#L33-L51)

### Step 2: Guard Against Missing DurationHead

The `require_num_frames_source` function provides fast-fail behavior before expensive computation:

```python
from ltx_pipelines.utils.blocks import require_num_frames_source

require_num_frames_source(
    num_frames=DEFAULT_AUTO_DURATION,
    duration_predictor=duration_predictor  # None if checkpoint lacks DurationHead

)

# Raises ValueError with clear message if predictor is None but auto-duration requested

```

Source: [[`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L94-L100)

### Step 3: Frame Count Resolution

After prompt encoding, `resolve_num_frames` produces the final integer:

```python
from ltx_pipelines.utils.blocks import resolve_num_frames

num_frames = resolve_num_frames(
    num_frames=DEFAULT_AUTO_DURATION,
    duration_predictor=duration_predictor,
    video_encoding=video_enc,      # From prompt encoder

    audio_encoding=audio_enc,      # Optional

    frame_rate=25.0,
)

# Returns int snapped to 8k+1 grid

```

Source: [[`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L108-L128)

Inside `DurationPredictor.__call__`, the logic flows through `seconds_to_clamped_num_frames` (called at line 76), which:
- Runs `DurationHead` forward pass → raw seconds
- Clamps to `[min_seconds, max_seconds]`
- Converts to frames and rounds to nearest `8k + 1` value

## Practical Usage Examples

### Basic Pipeline With Automatic Duration

```python
import torch
from ltx_pipelines.ti2vid_one_stage import Ti2VidOneStage
from ltx_pipelines.utils.types import DEFAULT_AUTO_DURATION

pipe = Ti2VidOneStage(
    checkpoint_path="path/to/checkpoint",
    dtype=torch.float16,
    device="cuda",
    num_frames=DEFAULT_AUTO_DURATION,  # Triggers automatic prediction

)

output = pipe(prompt="A drone flying over a coastal cliff at sunset")

# Frame count determined automatically from prompt

```

### Custom Duration Window

```python
from ltx_pipelines.utils.types import AutoDuration

# Request 3-10 second range instead of default 1-20

custom_duration = AutoDuration(min_seconds=3.0, max_seconds=10.0)

pipe = Ti2VidOneStage(
    checkpoint_path="path/to/checkpoint",
    num_frames=custom_duration,
)

```

### Graceful Degradation for Missing DurationHead

```python
from ltx_pipelines.utils.blocks import require_num_frames_source, DEFAULT_AUTO_DURATION

def safe_auto_duration(checkpoint_path, num_frames):
    from ltx_pipelines.utils.blocks import DurationPredictor
    
    predictor = DurationPredictor.from_checkpoint(
        checkpoint_path, device="cuda"
    )
    
    try:
        require_num_frames_source(num_frames, predictor)
        return predictor
    except ValueError as e:
        print(f"Falling back to manual 121 frames: {e}")
        return None  # Caller should use fixed num_frames

# Usage

predictor = safe_auto_duration("path/to/checkpoint", DEFAULT_AUTO_DURATION)

```

## Integration Points in Pipeline Classes

Pipeline implementations like `Ti2VidOneStage` wire these components together. The typical call chain:

1. `__init__` stores `num_frames` parameter (can be `int` or `AutoDuration`)
2. `__call__` receives prompt, encodes it
3. `resolve_num_frames` converts to concrete frame count
4. Generation proceeds with determined length

Source: [[`ti2vid_one_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_one_stage.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_one_stage.py) shows this pattern in the constructor (around line 144) and forward logic.

## Summary

- **`AutoDuration`** — Dataclass signaling automatic frame count selection with configurable min/max bounds
- **`DurationPredictor`** — Wrapper managing `DurationHead` inference and VAE-grid alignment
- **`DurationHead`** — Neural network predicting seconds from prompt encoder outputs
- **Guard and resolver functions** — `require_num_frames_source` for validation, `resolve_num_frames` for conversion
- **Checkpoint-based availability** — Auto-duration only works when `duration_head_path` exists in the model checkpoint

## Frequently Asked Questions

### What happens if I request automatic duration but my checkpoint lacks a DurationHead?

The `require_num_frames_source` guard raises a `ValueError` with a clear message before any expensive computation begins. This fast-fail behavior prevents wasted GPU cycles. Always verify your checkpoint includes `duration_head_path` or handle the exception gracefully.

### How does LTX-2 ensure predicted frame counts work with the VAE?

The `DurationPredictor` rounds all predictions to the nearest valid `8k + 1` frame count through `seconds_to_clamped_num_frames`. This guarantees compatibility with the causal temporal grid required by the VAE architecture, eliminating shape mismatches during encoding/decoding.

### Can I constrain the predicted duration to a specific range?

Yes. Instantiate `AutoDuration(min_seconds=2.0, max_seconds=8.0)` with your desired bounds instead of using `DEFAULT_AUTO_DURATION`. The predictor clamps raw network outputs to this window before frame conversion, giving you control over minimum and maximum clip lengths.

### Is the DurationHead prediction deterministic?

The `DurationHead` forward pass is deterministic given fixed prompt encoder outputs. However, if your pipeline uses stochastic prompt augmentation or varying random seeds, the input embeddings—and thus the predicted duration—may vary across runs with identical text prompts.