# Rotary Position Embedding (RoPE) Configurations for Video and Audio Streams in LTX-2

> Explore Rotary Position Embedding RoPE configurations in LTX-2. Discover SPLIT variant for video 3-D embeddings and 1-D temporal for audio streams, optimizing performance.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-06-21

---

**LTX-2 utilizes the `SPLIT` RoPE variant by default, configuring 3-D positional embeddings with `[20, 2048, 2048]` dimensions for video streams and 1-D temporal embeddings for audio streams.**

The Lightricks LTX-2 video generation model employs **Rotary Position Embeddings (RoPE)** to encode positional information across transformer attention layers. Understanding how these **Rotary Position Embedding configurations** differ between **video** and **audio** modalities is essential for implementing custom inference pipelines or fine-tuning workflows. This article examines the actual source code implementation to explain how LTX-2 handles multi-dimensional positional encoding for heterogeneous media streams.

## RoPE Variants: INTERLEAVED vs SPLIT

LTX-2 supports two distinct RoPE implementations defined in the `LTXRopeType` enum within [`packages/ltx-core/src/ltx_core/model/transformer/rope.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/rope.py) (lines 11-14). The **`INTERLEAVED`** variant represents a legacy mode that repeats sin/cos patterns across the entire embedding vector, maintained solely for backward compatibility with older checkpoints. Modern LTX-2 implementations default to the **`SPLIT`** variant, which processes sine and cosine tensors separately through the `apply_split_rotary_emb` function to align with per-head broadcasting logic and match the training data layout.

## Video Stream RoPE Configuration (3-D)

Video streams in LTX-2 utilize a **3-D frequency grid** encompassing temporal, spatial X, and spatial Y dimensions. The default **`max_pos`** configuration vector `[20, 2048, 2048]`—defined in [`rope.py`](https://github.com/Lightricks/LTX-2/blob/main/rope.py) (lines 88-92)—specifies 20 timesteps with 2048-pixel resolution for both spatial axes. This 3-D grid is constructed via the `precompute_freqs_cis` function and passed to attention layers, enabling the model to capture volumetric spatial-temporal relationships in generated video content.

## Audio Stream RoPE Configuration (1-D)

Audio streams employ the same **`SPLIT`** RoPE type but consume only the **temporal dimension** of the frequency grid. While the global default `max_pos` remains `[20, 2048, 2048]` for configuration consistency, audio transformers index exclusively the first entry (20 timesteps) and ignore the spatial components. This 1-D approach reflects the sequential nature of audio waveforms compared to the volumetric structure of video data, with the decoder extracting only the temporal slice from the full frequency grid.

## Source Code Implementation

### Enum Definition and Dispatch Logic

The `LTXRopeType` enum is defined in [`rope.py`](https://github.com/Lightricks/LTX-2/blob/main/rope.py) (lines 11-14) with two members: `INTERLEAVED` and `SPLIT`. The dispatch function **`apply_rotary_emb`** (lines 16-27) routes calls to the appropriate implementation based on the `rope_type` parameter. For `SPLIT` operations, this invokes the separate tensor handling logic required for per-head frequency broadcasting.

### Frequency Grid Generation

The **`precompute_freqs_cis`** function (lines 98-104) generates sinusoidal frequency tensors, padding dimensions to match `dim // 2` when using `SPLIT` mode. The implementation processes the grid through `split_freqs_cis` to prepare tensors compatible with per-head broadcasting (lines 112-119, 124-132). For video inputs, this generates the full 3-D grid, while audio processing extracts only the temporal slice from the first dimension of the indices grid.

### Transformer Configuration Objects

Configuration propagation occurs through **`TransformerArgs`** ([`transformer_args.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer_args.py), lines 99-112) and **`ModelConfigurator`** ([`model_configurator.py`](https://github.com/Lightricks/LTX-2/blob/main/model_configurator.py), lines 66-118). These classes instantiate transformers with modality-specific RoPE settings, ensuring video layers receive the complete 3-D frequency grid while audio layers request 1-D temporal frequencies despite using the identical `LTXRopeType.SPLIT` enum value.

## Code Examples

### Configuring Video Transformers

```python
from ltx_core.model.transformer.transformer import LTXTransformer
from ltx_core.model.transformer.rope import LTXRopeType

video_transformer = LTXTransformer(
    dim=1024,
    n_heads=16,
    depth=24,
    rope_type=LTXRopeType.SPLIT,   # 3-D RoPE for video

    # additional configuration...

)

```

### Configuring Audio Transformers

```python
audio_transformer = LTXTransformer(
    dim=512,
    n_heads=8,
    depth=12,
    rope_type=LTXRopeType.SPLIT,   # 1-D temporal RoPE for audio

    # additional configuration...

)

```

### Low-Level RoPE API Usage

```python
import torch
from ltx_core.model.transformer.rope import (
    precompute_freqs_cis,
    LTXRopeType,
    apply_rotary_emb,
)

# 1-D audio frequencies (temporal only)

audio_cos, audio_sin = precompute_freqs_cis(
    indices_grid=torch.zeros(1, 20, 1, 1),
    dim=512,
    out_dtype=torch.float32,
    rope_type=LTXRopeType.SPLIT,
)

# 3-D video frequencies (temporal + X + Y)

video_cos, video_sin = precompute_freqs_cis(
    indices_grid=torch.zeros(1, 20, 2048, 2048),
    dim=1024,
    out_dtype=torch.float32,
    rope_type=LTXRopeType.SPLIT,
)

# Apply to query tensors within attention

q_rotated = apply_rotary_emb(q, (video_cos, video_sin), LTXRopeType.SPLIT)

```

## Summary

- LTX-2 implements two RoPE variants (`INTERLEAVED` and `SPLIT`), with `SPLIT` serving as the default and recommended configuration for current checkpoints.
- **Video streams** utilize **3-D RoPE** with dimensions `[20, 2048, 2048]` covering temporal and dual spatial axes.
- **Audio streams** use **1-D RoPE** leveraging only the temporal component (20 timesteps) from the shared frequency grid.
- The **`LTXRopeType.SPLIT`** configuration is modality-agnostic, with dimensional behavior controlled by the `indices_grid` shape passed to **`precompute_freqs_cis`**.
- Key implementation files include [`rope.py`](https://github.com/Lightricks/LTX-2/blob/main/rope.py), [`transformer_args.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer_args.py), and [`model_configurator.py`](https://github.com/Lightricks/LTX-2/blob/main/model_configurator.py) within the `ltx_core` package.

## Frequently Asked Questions

### What is the difference between INTERLEAVED and SPLIT RoPE in LTX-2?

The **`INTERLEAVED`** variant repeats sin/cos patterns across the entire embedding vector as a legacy implementation, while **`SPLIT`** treats sine and cosine tensors separately to support per-head broadcasting. Modern LTX-2 checkpoints exclusively use `SPLIT` RoPE for both video and audio processing, as implemented in the `apply_rotary_emb` dispatch function.

### How does LTX-2 handle positional encoding for audio-only streams?

Audio transformers use the same `LTXRopeType.SPLIT` configuration but receive a 1-D temporal frequency grid instead of the full 3-D grid used for video. The implementation extracts only the first dimension (20 timesteps) from the default `max_pos` vector in [`rope.py`](https://github.com/Lightricks/LTX-2/blob/main/rope.py), ignoring the spatial coordinates entirely.

### Can I use INTERLEAVED RoPE with current LTX-2 checkpoints?

No, the **`INTERLEAVED`** variant is retained solely for backward compatibility and is **not compatible** with current LTX-2 checkpoints. Attempting to use it with modern model weights will result in incorrect positional encoding and degraded output quality, as the current training regime exclusively utilizes the `SPLIT` variant.

### Where are the default max_pos values defined for video dimensions?

The default **`max_pos`** values `[20, 2048, 2048]` are defined in [`packages/ltx-core/src/ltx_core/model/transformer/rope.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/rope.py) (lines 88-92). These values represent the maximum supported timesteps (20) and spatial resolutions (2048×2048 pixels) for video generation tasks, though audio processing utilizes only the temporal component.