# How LTX-2 Achieves Sub-Frame Alignment Using Cross-Attention: 1-D Temporal RoPE Explained

> LTX-2 achieves sub-frame alignment using 1-D temporal RoPE and cross-attention. This powerful technique synchronizes audio and video at a granularity finer than individual frames.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-06-21

---

**LTX-2 achieves sub-frame alignment by injecting 1-D temporal Rotary Positional Encoding (RoPE) into bidirectional cross-modal attention layers, allowing the model to synchronize audio samples and video patches at a granularity finer than individual frames.**

LTX-2 is an open-source video generation model developed by Lightricks that synchronizes audio and visual content with high precision. The model achieves sub-frame alignment through a specialized cross-attention mechanism that operates below the video frame rate, enabling precise lip-sync and sound-to-visual timing.

## Bidirectional Cross-Modal Attention Architecture

LTX-2 implements a dual-path cross-attention system that enables information flow between audio and video modalities in both directions.

### Dual Attention Modules

The transformer contains two dedicated cross-attention modules instantiated in [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py) (lines 151-174):

- **`audio_to_video_attn`**: Queries originate from video features, while keys and values come from audio representations
- **`video_to_audio_attn`**: Queries originate from audio features, while keys and values come from video representations

This bidirectional design ensures that both modalities can influence each other during the generation process, rather than treating one as a static condition for the other.

## 1-D Temporal RoPE for Sub-Frame Precision

Unlike the 3-D RoPE used for spatial video patches, the cross-modal attention paths employ **1-D RoPE** that encodes only the temporal axis, enabling alignment at sub-frame resolutions.

### Generating Temporal Frequency Grids

The positional embeddings are generated in [`packages/ltx-core/src/ltx_core/model/transformer/rope.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/rope.py) via the `generate_freq_grid_*` and `precompute_freqs_cis` functions. These utilities create frequency tensors that represent temporal positions at a granularity capable of resolving individual audio samples within a video frame's duration.

### Embeddings1DConnector Implementation

The `Embeddings1DConnector` class (found in [`packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py), lines 72-86) builds 1-D RoPE frequencies for sequential embeddings. This connector processes audio or video token streams and outputs `freqs_cis` tensors that encode temporal positions linearly rather than spatially.

The resulting `freqs_cis` tensor is passed to the cross-attention call as the `pe` argument, supplying the temporal coordinate information necessary for sub-frame alignment.

## Cross-Attention Forward Pass with RoPE

During the forward pass, the transformer applies cross-attention after self-attention, injecting the 1-D RoPE into the attention computation. In [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py) (lines 860-996), the video-to-audio cross-attention is called as follows:

```python
vx = vx + self._apply_text_cross_attention(
    vx_normed,
    video.context,
    self.attn2,
    self.scale_shift_table,
    getattr(self, "prompt_scale_shift_table", None),
    video.timesteps,
    video.prompt_timestep,
    video.context_mask,
    cross_attention_adaln=self.cross_attention_adaln,
)

```

Here, `self.attn2` represents the **video-to-audio** cross-attention module. The same pattern applies for the audio-to-video direction. The `pe` argument (containing the 1-D RoPE tensor) is supplied to the underlying `Attention` implementation in [`packages/ltx-core/src/ltx_core/model/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/attention.py), which rotates queries and keys before computing the dot-product.

This rotation encodes the relative temporal positions into the attention scores, allowing the model to learn alignments between audio samples and video patches that occur at any offset within a frame's duration.

## Sub-Frame Alignment in Practice

Because the temporal RoPE can represent positions finer than a video frame (e.g., individual samples within a frame's duration), the attention mechanism learns to align lip-sync cues occurring at any audio sample with the corresponding visual patch. This achieves sub-frame synchronization across modalities without requiring explicit frame-rate conversion or temporal pooling.

## Code Example: Implementing 1-D RoPE for Cross-Modal Attention

The following implementation demonstrates how to prepare 1-D RoPE frequencies and pass them through the transformer:

```python

# 1️⃣ Prepare 1-D RoPE frequencies for an audio token stream

from ltx_core.text_encoders.gemma.embeddings_connector import Embeddings1DConnector

connector = Embeddings1DConnector(
    num_layers=2,
    rope_type="split",               # 1-D RoPE variant

    causal_temporal_positioning=False,
)
audio_embeddings, audio_mask = connector(
    audio_latents, 
    additive_attention_mask=None
)

# 2️⃣ Pass the embeddings into the transformer (cross-attention will use RoPE)

from ltx_core.model.transformer.transformer import Transformer

model = Transformer(
    video=video_args,               # video transformer args

    audio=audio_args,               # audio transformer args (includes the 1-D RoPE)

)
video_out, audio_out = model(video=video_args, audio=audio_args)

# 3️⃣ Inside the transformer, the cross-attention receives the RoPE tensor:

#    `self.attn2(..., pe=video.positional_embeddings, ...)`

```

## Summary

- **Bidirectional cross-attention**: LTX-2 uses `audio_to_video_attn` and `video_to_audio_attn` modules to enable mutual information flow between modalities.
- **1-D temporal RoPE**: The model applies 1-D Rotary Positional Encoding to cross-modal attention, encoding only the temporal axis rather than spatial dimensions.
- **Sub-frame resolution**: By representing temporal positions at a granularity finer than video frames, the model aligns audio samples with visual content at sub-frame precision.
- **Implementation location**: Key components reside in [`transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer.py) (instantiation and forward pass), [`rope.py`](https://github.com/Lightricks/LTX-2/blob/main/rope.py) (frequency generation), and [`embeddings_connector.py`](https://github.com/Lightricks/LTX-2/blob/main/embeddings_connector.py) (1-D embedding preparation).

## Frequently Asked Questions

### What is sub-frame alignment in video generation?

Sub-frame alignment refers to the ability to synchronize audio and video content at a temporal resolution finer than individual video frames. Standard video operates at 24-60 frames per second, while audio typically uses 44,100 samples per second. Sub-frame alignment allows models to match specific audio samples with precise moments within a frame's duration, enabling accurate lip-sync and sound-visual coordination.

### How does 1-D RoPE differ from 3-D RoPE in LTX-2?

**3-D RoPE** encodes spatial and temporal dimensions (height, width, time) for video patches, treating each patch as a 3D volume. **1-D RoPE** encodes only the temporal axis, making it suitable for sequential token streams like audio or flattened temporal sequences. In LTX-2, the cross-modal attention layers use 1-D RoPE to align modalities purely along the time dimension without spatial confounding.

### Where are the cross-attention modules instantiated in the codebase?

The bidirectional cross-attention modules are instantiated in [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py) at lines 151-174. This section defines both `audio_to_video_attn` and `video_to_audio_attn` as distinct transformer layers that operate during the forward pass.

### Why is bidirectional cross-attention necessary for audio-video synchronization?

Bidirectional cross-attention allows both modalities to update their representations based on the other. Unidirectional attention would only allow one modality to attend to the other (e.g., video conditioned on audio), but bidirectional flow enables the model to resolve ambiguities in both directions—such as when audio clarifies visual timing or vice versa—resulting in more coherent multimodal generation.