# How Duration Prediction Works to Auto-Determine Frame Count in LTX-2

> Discover how LTX-2's DurationHead predicts video duration in seconds, automatically setting frame count while respecting VAE temporal constraints for efficient video generation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-20

---

**LTX-2 uses a lightweight regression head called `DurationHead` that analyzes caption encoder embeddings to predict video duration in seconds, which is then converted to a frame count that respects the VAE's temporal grid constraints.**

The `DurationHead` enables automatic frame count determination from text prompts, eliminating manual `--num-frames` specification. This mechanism, available in LTX-2.5 and later checkpoints, keeps the predictor small (few megabytes) and runs on frozen connector outputs without requiring full model inference.

## The Auto-Duration Request Flow

### User-Level Configuration with `AutoDuration`

Users trigger automatic duration prediction by passing `AutoDuration` instead of an integer to `num_frames`. This dataclass accepts optional `min_seconds` and `max_seconds` bounds to constrain the prediction.

```python
from ltx_pipelines.utils.types import AutoDuration

# Default: allow 1-20 second range

num_frames = AutoDuration()

# Constrained: force 2-15 second range

num_frames = AutoDuration(min_seconds=2.0, max_seconds=15.0)

```

The `AutoDuration` type is defined in [`ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/types.py)【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py#L15-L24】.

### Early Validation with `require_num_frames_source`

Every pipeline validates the request immediately. The `require_num_frames_source` function in [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py) raises a clear error if `AutoDuration` is requested but the checkpoint lacks a trained `DurationHead` (pre-2.5 checkpoints)【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L94-L100】.

This guard prevents silent failures and guides users to either upgrade checkpoints or specify frames manually.

## The Duration Prediction Architecture

### Step 1: Caption Encoding

The `PromptEncoder` (Gemma-based) processes the text caption and produces connector token tensors: `video_encoding` and `audio_encoding`. These frozen embeddings feed directly into the duration predictor—no diffusion model components execute at this stage.

### Step 2: `DurationPredictor` Execution

The `DurationPredictor` class wraps the `DurationHead` and is instantiated via `DurationPredictor.from_checkpoint`【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L3-L12】.

When called, it forwards connector tensors to the head:

```python

# From ltx_pipelines/utils/blocks.py, lines 50-71

def __call__(
    self,
    *,
    video_encoding: Optional[torch.Tensor] = None,
    audio_encoding: Optional[torch.Tensor] = None,
) -> float:
    # ... validation and tensor preparation ...

    output = self.duration_head(
        video_encoding=video_encoding,
        audio_encoding=audio_encoding,
    )
    return output["seconds"].item()

```

The predictor handles both video-only and audio-video scenarios, concatenating available encodings before passing to the head.

### Step 3: `DurationHead` Regression

The core `DurationHead` in [`ltx_core/duration_head/duration_head.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/duration_head/duration_head.py) processes tokens through a learned attention pooler followed by a small MLP【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-core/src/ltx_core/duration_head/duration_head.py#L52-L118】.

Key implementation details:
- **Input**: Pooled token streams from connector outputs
- **Regression target**: **Log-seconds** (exponentiated to seconds for final output)
- **Architecture**: Attention-based pooling + 2-layer MLP with hidden dimension 256
- **Output**: Single scalar representing predicted duration

Log-space prediction stabilizes training across the wide range of human-perceived durations (1-20+ seconds).

## Frame Count Conversion and Grid Alignment

### `seconds_to_clamped_num_frames`

Raw seconds predictions undergo three transformations in [`ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/helpers.py)【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L65-L85】:

```python
num_frames = seconds_to_clamped_num_frames(
    seconds=predicted_seconds,
    frame_rate=25.0,  # or user-specified

    min_frames=round(min_seconds * frame_rate),
    max_frames=round(max_seconds * frame_rate),
)

```

The conversion pipeline:

1. **Round**: Convert seconds to raw frame count (`seconds * frame_rate`)
2. **Clamp**: Enforce `[min_frames, max_frames]` bounds for memory safety
3. **Snap**: Align to VAE causal temporal grid via `snap_frames_to_grid`

### Grid Snapping with `snap_frames_to_grid`

The VAE requires frames satisfying `(frames - 1) % 8k == 0` for temporal alignment. The `snap_frames_to_grid` function in [`ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/helpers.py)【/cache/repos/github.com/Lightricks/LTX-2/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py#L54-L62】 rounds to valid grid points:

- Finds nearest valid frame count above and below
- Selects closest valid value
- Guarantees VAE compatibility without user intervention

This ensures generated videos decode correctly without temporal artifacts.

## Complete Usage Examples

### Standard Pipeline Usage

```python
from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from ltx_pipelines.utils.types import AutoDuration
import torch

pipeline = TI2VidOneStagePipeline(
    checkpoint_path="checkpoints/ltx_v2a_lora.ckpt",
    dtype=torch.float16,
    device="cuda",
    # Intentionally omit num_frames to trigger auto-duration

)

result = pipeline(
    prompt="A sunrise over a misty forest",
    num_frames=AutoDuration(min_seconds=2.0, max_seconds=15.0),
)

print(f"Generated {result.video.shape[1]} frames at 25 FPS")

# Output: Generated 151 frames at 25 FPS (~6 seconds)

```

### Manual `DurationPredictor` Usage

```python
from ltx_pipelines.utils.blocks import DurationPredictor, resolve_num_frames
from ltx_pipelines.utils.types import AutoDuration

# Build predictor from checkpoint containing DurationHead

predictor = DurationPredictor.from_checkpoint(
    checkpoint_path="checkpoints/ltx_v2a_lora.ckpt",
    dtype=torch.float16,
    device="cuda",
)

# Assume connector encodings from prior encoding pass

video_enc, audio_enc = get_connector_encodings(prompt)

# Resolve to concrete frame count

num_frames = resolve_num_frames(
    num_frames=AutoDuration(),
    duration_predictor=predictor,
    video_encoding=video_enc,
    audio_encoding=audio_enc,
    frame_rate=25.0,
)

print(f"Predicted {num_frames} frames")

```

## Key Implementation Files

| File | Component | Purpose |
|------|-----------|---------|
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py) | `AutoDuration` dataclass | User-facing API for auto-duration requests |
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) | `DurationPredictor`, `resolve_num_frames`, `require_num_frames_source` | Orchestration and validation layer |
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) | `seconds_to_clamped_num_frames`, `snap_frames_to_grid` | Conversion and grid alignment |
| [`packages/ltx-core/src/ltx_core/duration_head/duration_head.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/duration_head/duration_head.py) | `DurationHead` | Core regression network predicting log-seconds |
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) | CLI argument parsing | `--auto-duration` flag handling |

## Summary

- **`AutoDuration`** triggers automatic frame count determination via the `DurationPredictor` subsystem
- **`DurationHead`** performs lightweight regression on frozen caption encoder outputs to predict duration in log-seconds
- **`seconds_to_clamped_num_frames`** converts predictions to concrete frame counts with memory-safe clamping
- **`snap_frames_to_grid`** enforces VAE temporal alignment constraints without user intervention
- **Early validation** prevents runtime errors when checkpoints lack trained duration heads

The design prioritizes efficiency: the predictor adds minimal memory overhead, runs once per generation request, and avoids loading or executing the full diffusion model for duration estimation.

## Frequently Asked Questions

### What LTX-2 checkpoints support auto-duration prediction?

Auto-duration requires checkpoints with a trained `DurationHead`, which debuted in LTX-2.5 releases. Pre-2.5 checkpoints lack this component, and the pipeline raises a clear error if `AutoDuration` is requested without an available predictor. Always use `DurationPredictor.from_checkpoint` to verify predictor availability or check checkpoint metadata for the `duration_head` key.

### Why does the predicted frame count differ slightly from my requested duration?

The `snap_frames_to_grid` function aligns frame counts to the VAE's causal temporal requirements: `(frames - 1) % 8k == 0`. For LTX-2's 8-frame temporal compression, valid frame counts follow sequences like 1, 9, 17, 25, 33, etc. Your requested duration is first converted to frames, then rounded to the nearest valid grid point. This ensures artifact-free video decoding.

### Can I use `DurationPredictor` independently of the full pipeline?

Yes. Instantiate `DurationPredictor.from_checkpoint` with any compatible checkpoint, then call it directly with video and/or audio connector encodings. This is useful for dataset analysis, duration-conditioned training, or building custom inference workflows. The predictor accepts both `video_encoding` and `audio_encoding` tensors and returns a float seconds value.

### How does the `min_seconds` / `max_seconds` clamping interact with grid snapping?

Clamping occurs before grid snapping. The pipeline first converts your bounds to frame counts, clamps the raw prediction to this range, then snaps to valid grid points. If your bounds exclude all valid grid points (unlikely with default 1-20s ranges), the result boundaries to the nearest valid grid point within your constraints.