Rotary Position Embedding (RoPE) Configurations for Video and Audio Streams in LTX-2

LTX-2 utilizes the SPLIT RoPE variant by default, configuring 3-D positional embeddings with [20, 2048, 2048] dimensions for video streams and 1-D temporal embeddings for audio streams.

The Lightricks LTX-2 video generation model employs Rotary Position Embeddings (RoPE) to encode positional information across transformer attention layers. Understanding how these Rotary Position Embedding configurations differ between video and audio modalities is essential for implementing custom inference pipelines or fine-tuning workflows. This article examines the actual source code implementation to explain how LTX-2 handles multi-dimensional positional encoding for heterogeneous media streams.

RoPE Variants: INTERLEAVED vs SPLIT

LTX-2 supports two distinct RoPE implementations defined in the LTXRopeType enum within packages/ltx-core/src/ltx_core/model/transformer/rope.py (lines 11-14). The INTERLEAVED variant represents a legacy mode that repeats sin/cos patterns across the entire embedding vector, maintained solely for backward compatibility with older checkpoints. Modern LTX-2 implementations default to the SPLIT variant, which processes sine and cosine tensors separately through the apply_split_rotary_emb function to align with per-head broadcasting logic and match the training data layout.

Video Stream RoPE Configuration (3-D)

Video streams in LTX-2 utilize a 3-D frequency grid encompassing temporal, spatial X, and spatial Y dimensions. The default max_pos configuration vector [20, 2048, 2048]—defined in rope.py (lines 88-92)—specifies 20 timesteps with 2048-pixel resolution for both spatial axes. This 3-D grid is constructed via the precompute_freqs_cis function and passed to attention layers, enabling the model to capture volumetric spatial-temporal relationships in generated video content.

Audio Stream RoPE Configuration (1-D)

Audio streams employ the same SPLIT RoPE type but consume only the temporal dimension of the frequency grid. While the global default max_pos remains [20, 2048, 2048] for configuration consistency, audio transformers index exclusively the first entry (20 timesteps) and ignore the spatial components. This 1-D approach reflects the sequential nature of audio waveforms compared to the volumetric structure of video data, with the decoder extracting only the temporal slice from the full frequency grid.

Source Code Implementation

Enum Definition and Dispatch Logic

The LTXRopeType enum is defined in rope.py (lines 11-14) with two members: INTERLEAVED and SPLIT. The dispatch function apply_rotary_emb (lines 16-27) routes calls to the appropriate implementation based on the rope_type parameter. For SPLIT operations, this invokes the separate tensor handling logic required for per-head frequency broadcasting.

Frequency Grid Generation

The precompute_freqs_cis function (lines 98-104) generates sinusoidal frequency tensors, padding dimensions to match dim // 2 when using SPLIT mode. The implementation processes the grid through split_freqs_cis to prepare tensors compatible with per-head broadcasting (lines 112-119, 124-132). For video inputs, this generates the full 3-D grid, while audio processing extracts only the temporal slice from the first dimension of the indices grid.

Transformer Configuration Objects

Configuration propagation occurs through TransformerArgs (transformer_args.py, lines 99-112) and ModelConfigurator (model_configurator.py, lines 66-118). These classes instantiate transformers with modality-specific RoPE settings, ensuring video layers receive the complete 3-D frequency grid while audio layers request 1-D temporal frequencies despite using the identical LTXRopeType.SPLIT enum value.

Code Examples

Configuring Video Transformers

from ltx_core.model.transformer.transformer import LTXTransformer
from ltx_core.model.transformer.rope import LTXRopeType

video_transformer = LTXTransformer(
    dim=1024,
    n_heads=16,
    depth=24,
    rope_type=LTXRopeType.SPLIT,   # 3-D RoPE for video

    # additional configuration...

)

Configuring Audio Transformers

audio_transformer = LTXTransformer(
    dim=512,
    n_heads=8,
    depth=12,
    rope_type=LTXRopeType.SPLIT,   # 1-D temporal RoPE for audio

    # additional configuration...

)

Low-Level RoPE API Usage

import torch
from ltx_core.model.transformer.rope import (
    precompute_freqs_cis,
    LTXRopeType,
    apply_rotary_emb,
)

# 1-D audio frequencies (temporal only)

audio_cos, audio_sin = precompute_freqs_cis(
    indices_grid=torch.zeros(1, 20, 1, 1),
    dim=512,
    out_dtype=torch.float32,
    rope_type=LTXRopeType.SPLIT,
)

# 3-D video frequencies (temporal + X + Y)

video_cos, video_sin = precompute_freqs_cis(
    indices_grid=torch.zeros(1, 20, 2048, 2048),
    dim=1024,
    out_dtype=torch.float32,
    rope_type=LTXRopeType.SPLIT,
)

# Apply to query tensors within attention

q_rotated = apply_rotary_emb(q, (video_cos, video_sin), LTXRopeType.SPLIT)

Summary

  • LTX-2 implements two RoPE variants (INTERLEAVED and SPLIT), with SPLIT serving as the default and recommended configuration for current checkpoints.
  • Video streams utilize 3-D RoPE with dimensions [20, 2048, 2048] covering temporal and dual spatial axes.
  • Audio streams use 1-D RoPE leveraging only the temporal component (20 timesteps) from the shared frequency grid.
  • The LTXRopeType.SPLIT configuration is modality-agnostic, with dimensional behavior controlled by the indices_grid shape passed to precompute_freqs_cis.
  • Key implementation files include rope.py, transformer_args.py, and model_configurator.py within the ltx_core package.

Frequently Asked Questions

What is the difference between INTERLEAVED and SPLIT RoPE in LTX-2?

The INTERLEAVED variant repeats sin/cos patterns across the entire embedding vector as a legacy implementation, while SPLIT treats sine and cosine tensors separately to support per-head broadcasting. Modern LTX-2 checkpoints exclusively use SPLIT RoPE for both video and audio processing, as implemented in the apply_rotary_emb dispatch function.

How does LTX-2 handle positional encoding for audio-only streams?

Audio transformers use the same LTXRopeType.SPLIT configuration but receive a 1-D temporal frequency grid instead of the full 3-D grid used for video. The implementation extracts only the first dimension (20 timesteps) from the default max_pos vector in rope.py, ignoring the spatial coordinates entirely.

Can I use INTERLEAVED RoPE with current LTX-2 checkpoints?

No, the INTERLEAVED variant is retained solely for backward compatibility and is not compatible with current LTX-2 checkpoints. Attempting to use it with modern model weights will result in incorrect positional encoding and degraded output quality, as the current training regime exclusively utilizes the SPLIT variant.

Where are the default max_pos values defined for video dimensions?

The default max_pos values [20, 2048, 2048] are defined in packages/ltx-core/src/ltx_core/model/transformer/rope.py (lines 88-92). These values represent the maximum supported timesteps (20) and spatial resolutions (2048×2048 pixels) for video generation tasks, though audio processing utilizes only the temporal component.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →