How LTX-2 Achieves Sub-Frame Alignment Using Cross-Attention: 1-D Temporal RoPE Explained

LTX-2 achieves sub-frame alignment by injecting 1-D temporal Rotary Positional Encoding (RoPE) into bidirectional cross-modal attention layers, allowing the model to synchronize audio samples and video patches at a granularity finer than individual frames.

LTX-2 is an open-source video generation model developed by Lightricks that synchronizes audio and visual content with high precision. The model achieves sub-frame alignment through a specialized cross-attention mechanism that operates below the video frame rate, enabling precise lip-sync and sound-to-visual timing.

Bidirectional Cross-Modal Attention Architecture

LTX-2 implements a dual-path cross-attention system that enables information flow between audio and video modalities in both directions.

Dual Attention Modules

The transformer contains two dedicated cross-attention modules instantiated in packages/ltx-core/src/ltx_core/model/transformer/transformer.py (lines 151-174):

  • audio_to_video_attn: Queries originate from video features, while keys and values come from audio representations
  • video_to_audio_attn: Queries originate from audio features, while keys and values come from video representations

This bidirectional design ensures that both modalities can influence each other during the generation process, rather than treating one as a static condition for the other.

1-D Temporal RoPE for Sub-Frame Precision

Unlike the 3-D RoPE used for spatial video patches, the cross-modal attention paths employ 1-D RoPE that encodes only the temporal axis, enabling alignment at sub-frame resolutions.

Generating Temporal Frequency Grids

The positional embeddings are generated in packages/ltx-core/src/ltx_core/model/transformer/rope.py via the generate_freq_grid_* and precompute_freqs_cis functions. These utilities create frequency tensors that represent temporal positions at a granularity capable of resolving individual audio samples within a video frame's duration.

Embeddings1DConnector Implementation

The Embeddings1DConnector class (found in packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py, lines 72-86) builds 1-D RoPE frequencies for sequential embeddings. This connector processes audio or video token streams and outputs freqs_cis tensors that encode temporal positions linearly rather than spatially.

The resulting freqs_cis tensor is passed to the cross-attention call as the pe argument, supplying the temporal coordinate information necessary for sub-frame alignment.

Cross-Attention Forward Pass with RoPE

During the forward pass, the transformer applies cross-attention after self-attention, injecting the 1-D RoPE into the attention computation. In packages/ltx-core/src/ltx_core/model/transformer/transformer.py (lines 860-996), the video-to-audio cross-attention is called as follows:

vx = vx + self._apply_text_cross_attention(
    vx_normed,
    video.context,
    self.attn2,
    self.scale_shift_table,
    getattr(self, "prompt_scale_shift_table", None),
    video.timesteps,
    video.prompt_timestep,
    video.context_mask,
    cross_attention_adaln=self.cross_attention_adaln,
)

Here, self.attn2 represents the video-to-audio cross-attention module. The same pattern applies for the audio-to-video direction. The pe argument (containing the 1-D RoPE tensor) is supplied to the underlying Attention implementation in packages/ltx-core/src/ltx_core/model/transformer/attention.py, which rotates queries and keys before computing the dot-product.

This rotation encodes the relative temporal positions into the attention scores, allowing the model to learn alignments between audio samples and video patches that occur at any offset within a frame's duration.

Sub-Frame Alignment in Practice

Because the temporal RoPE can represent positions finer than a video frame (e.g., individual samples within a frame's duration), the attention mechanism learns to align lip-sync cues occurring at any audio sample with the corresponding visual patch. This achieves sub-frame synchronization across modalities without requiring explicit frame-rate conversion or temporal pooling.

Code Example: Implementing 1-D RoPE for Cross-Modal Attention

The following implementation demonstrates how to prepare 1-D RoPE frequencies and pass them through the transformer:


# 1️⃣ Prepare 1-D RoPE frequencies for an audio token stream

from ltx_core.text_encoders.gemma.embeddings_connector import Embeddings1DConnector

connector = Embeddings1DConnector(
    num_layers=2,
    rope_type="split",               # 1-D RoPE variant

    causal_temporal_positioning=False,
)
audio_embeddings, audio_mask = connector(
    audio_latents, 
    additive_attention_mask=None
)

# 2️⃣ Pass the embeddings into the transformer (cross-attention will use RoPE)

from ltx_core.model.transformer.transformer import Transformer

model = Transformer(
    video=video_args,               # video transformer args

    audio=audio_args,               # audio transformer args (includes the 1-D RoPE)

)
video_out, audio_out = model(video=video_args, audio=audio_args)

# 3️⃣ Inside the transformer, the cross-attention receives the RoPE tensor:

#    `self.attn2(..., pe=video.positional_embeddings, ...)`

Summary

  • Bidirectional cross-attention: LTX-2 uses audio_to_video_attn and video_to_audio_attn modules to enable mutual information flow between modalities.
  • 1-D temporal RoPE: The model applies 1-D Rotary Positional Encoding to cross-modal attention, encoding only the temporal axis rather than spatial dimensions.
  • Sub-frame resolution: By representing temporal positions at a granularity finer than video frames, the model aligns audio samples with visual content at sub-frame precision.
  • Implementation location: Key components reside in transformer.py (instantiation and forward pass), rope.py (frequency generation), and embeddings_connector.py (1-D embedding preparation).

Frequently Asked Questions

What is sub-frame alignment in video generation?

Sub-frame alignment refers to the ability to synchronize audio and video content at a temporal resolution finer than individual video frames. Standard video operates at 24-60 frames per second, while audio typically uses 44,100 samples per second. Sub-frame alignment allows models to match specific audio samples with precise moments within a frame's duration, enabling accurate lip-sync and sound-visual coordination.

How does 1-D RoPE differ from 3-D RoPE in LTX-2?

3-D RoPE encodes spatial and temporal dimensions (height, width, time) for video patches, treating each patch as a 3D volume. 1-D RoPE encodes only the temporal axis, making it suitable for sequential token streams like audio or flattened temporal sequences. In LTX-2, the cross-modal attention layers use 1-D RoPE to align modalities purely along the time dimension without spatial confounding.

Where are the cross-attention modules instantiated in the codebase?

The bidirectional cross-attention modules are instantiated in packages/ltx-core/src/ltx_core/model/transformer/transformer.py at lines 151-174. This section defines both audio_to_video_attn and video_to_audio_attn as distinct transformer layers that operate during the forward pass.

Why is bidirectional cross-attention necessary for audio-video synchronization?

Bidirectional cross-attention allows both modalities to update their representations based on the other. Unidirectional attention would only allow one modality to attend to the other (e.g., video conditioned on audio), but bidirectional flow enables the model to resolve ambiguities in both directions—such as when audio clarifies visual timing or vice versa—resulting in more coherent multimodal generation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →