# LTX-2 Transformer Block: Complete Sequence of Operations Explained

> Uncover the LTX-2 transformer block operations. Explore the sequence of self-attention, cross-attention, and feed-forward layers in this multimodal processing pipeline.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-06-21

---

**Each LTX-2 transformer block (`BasicAVTransformerBlock`) processes multimodal tokens through a rigid pipeline of AdaLN-conditioned self-attention, text cross-attention, optional audio-video cross-attention, and gated feed-forward layers before returning updated representations.**

The `BasicAVTransformerBlock` class in the Lightricks/LTX-2 repository defines the core computational unit for video and audio generation. An LTX-2 transformer block handles both unimodal and multimodal inputs using adaptive layer normalization (AdaLN) and specialized attention mechanisms configured via `TransformerConfig` objects. Understanding the exact sequence of operations within these blocks is essential for modifying generation behavior or optimizing inference pipelines.

## LTX-2 Transformer Block Construction and Initialization

Before the forward pass begins, the block initializes its components based on provided modality configurations.

### Configuration and Layer Instantiation

The constructor receives optional *video* and *audio* `TransformerConfig` objects, a rotary-embedding type (`LTXRopeType`), an epsilon value for RMS normalization, and pluggable operations (`TransformerOpsConfig`). For each modality present, the block instantiates self-attention (`attn1`), cross-attention (`attn2`), and a feed-forward network (`ff`) according to the configuration dimensions and attention heads. This initialization occurs in lines 86–109 and 111–122 of the transformer module.

### AdaLN Scale-Shift Table Allocation

For every enabled modality, the block allocates learnable `scale_shift_table` parameters. These tables store per-timestep affine transformation parameters that modulate token representations throughout the forward pass. When cross-attention AdaLN is enabled, additional prompt-specific tables are also initialized. This setup appears in lines 123–129 and 136–147 of [`transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer.py).

## Forward Pass Operations in LTX-2 Transformer Blocks

The `forward` method (starting at line 53) executes an eight-step sequence for each active modality. The operations proceed in a fixed order: self-attention, text cross-attention, audio-video cross-attention (when both modalities are present), and finally feed-forward processing.

### Step 1: Self-Attention with AdaZero Modulation

The first transformation applied to incoming tokens is multi-head self-attention (MSA) with adaptive normalization.

- **AdaLN Parameter Extraction**: The block calls `get_ada_values` to extract three parameters—`shift`, `scale`, and `gate`—from the modality’s `scale_shift_table` based on the current timestep.
- **RMS Normalization and AdaZero**: Input tokens undergo RMS normalization, then modulation via the `ada_zero_function`, which applies the extracted shift and scale parameters.
- **Attention Computation**: The normalized tensor passes through `attn1` (the first attention layer), utilizing optional positional embeddings and self-attention masks.
- **Gating**: The `post_sa_function` applies RMS normalization again and multiplies the output by the gate parameter before the residual connection.

For video inputs, this flow occupies lines 70–85; for audio inputs, lines 99–114.

### Step 2: Text Cross-Attention Implementation

Following self-attention, the block processes text conditioning through `_apply_text_cross_attention`.

This method receives the normalized representation (`vx_normed` for video, `ax_normed` for audio) and performs three sub-operations: it optionally extracts additional AdaLN parameters for the query (lines 38–40), applies `apply_cross_attention_adaln` to modulate both the query and context tensors, and executes the second attention layer (`attn2` or `audio_attn2`). The implementation spans lines 22–34 (definition) and lines 86–96 and 115–125 (usage).

### Step 3: Audio-Video Cross-Attention (A2V and V2A)

When both video and audio configurations are provided, the LTX-2 transformer block performs bidirectional cross-attention between modalities.

- **Snapshot Preservation**: The code saves pre-cross-attention tensors (`vx_pre_av`, `ax_pre_av`) to prevent order-dependent bias in the updates.
- **Audio-to-Video (A2V)**: The block extracts dedicated AdaLN parameters (`scale_ca_video_a2v`, `shift_ca_video_a2v`) from a specific table, applies AdaZero scaling, and runs `audio_to_video_attn` using audio tokens as keys and values while attending to video queries.
- **Video-to-Audio (V2A)**: The symmetric operation uses `video_to_audio_attn` with corresponding AdaLN parameters to update audio representations based on video content.

Both directions apply gating before adding the residual update. This cross-modal interaction is implemented in lines 28–66 of [`transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer.py).

### Step 4: Feed-Forward Network Processing

The final operation within each LTX-2 transformer block is the feed-forward (MLP) sub-layer.

After all attention computations complete, the block fetches MLP-specific AdaLN parameters (`shift_mlp`, `scale_mlp`, `gate_mlp`) from lines 98–101. Tokens undergo AdaZero scaling, pass through the modality-specific feed-forward network (`ff` for video, `audio_ff` for audio), receive gate multiplication, and merge back via residual connection. The video feed-forward logic appears in lines 98–108, while audio processing occupies lines 111–121.

### Output Generation

The `forward` method concludes at line 416 by returning updated `TransformerArgs` objects containing the transformed tensors. If a specific modality was not enabled in the configuration, the method returns `None` for that argument.

## Code Example: Running a Single LTX-2 Transformer Block

The following example demonstrates how to instantiate and execute one LTX-2 transformer block with both video and audio pathways enabled.

```python
import torch
from ltx_core.model.transformer.transformer import BasicAVTransformerBlock, TransformerConfig
from ltx_core.model.transformer.transformer_args import TransformerArgs

# Configure per-modality parameters

video_cfg = TransformerConfig(
    dim=512, heads=8, d_head=64,
    context_dim=768, apply_gated_attention=False,
    cross_attention_adaln=True
)

audio_cfg = TransformerConfig(
    dim=256, heads=4, d_head=64,
    context_dim=384, apply_gated_attention=False,
    cross_attention_adaln=False
)

# Initialize the transformer block

block = BasicAVTransformerBlock(video=video_cfg, audio=audio_cfg)

# Prepare dummy input tensors

batch, seq_len = 2, 16

video_args = TransformerArgs(
    x=torch.randn(batch, seq_len, video_cfg.dim),
    timesteps=torch.randn(batch, 1, 1),
    positional_embeddings=torch.randn(batch, seq_len, video_cfg.dim),
    self_attention_mask=None,
    self_attn_perturbation_mask=None,
    self_attn_all_perturbed=False,
    context=None,
    context_mask=None,
    cross_scale_shift_timestep=torch.randn(batch, 1, 1),
    cross_gate_timestep=torch.randn(batch, 1, 1),
    cross_attn_skip_all=False,
    enabled=True,
)

audio_args = TransformerArgs(
    x=torch.randn(batch, seq_len, audio_cfg.dim),
    timesteps=torch.randn(batch, 1, 1),
    positional_embeddings=torch.randn(batch, seq_len, audio_cfg.dim),
    self_attention_mask=None,
    self_attn_perturbation_mask=None,
    self_attn_all_perturbed=False,
    context=None,
    context_mask=None,
    cross_scale_shift_timestep=torch.randn(batch, 1, 1),
    cross_gate_timestep=torch.randn(batch, 1, 1),
    cross_attn_skip_all=False,
    enabled=True,
)

# Execute forward pass

new_video, new_audio = block(video_args, audio_args)

print(new_video.x.shape)   # Output: torch.Size([2, 16, 512])

print(new_audio.x.shape)   # Output: torch.Size([2, 16, 256])

```

## Key Source Files and Implementation Details

Understanding the LTX-2 transformer block requires familiarity with these specific source files in the repository:

- **[`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py)** – Contains the `BasicAVTransformerBlock` class definition and the complete forward pipeline logic referenced in lines 53–416.

- **[`packages/ltx-core/src/ltx_core/model/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/attention.py)** – Implements the `Attention` class used by `attn1` and `attn2` for both self-attention and cross-attention operations.

- **[`packages/ltx-core/src/ltx_core/model/transformer/feed_forward.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/feed_forward.py)** – Defines the MLP architecture applied in the final step of each block.

- **[`packages/ltx-core/src/ltx_core/model/transformer/adaln.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/adaln.py)** – Provides utility functions for computing AdaLN coefficients and the `ada_zero_function` used throughout the normalization steps.

- **[`packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py)** – Defines the `TransformerArgs` dataclass that encapsulates input tensors, masks, timesteps, and configuration flags passed between blocks.

## Summary

An LTX-2 transformer block processes multimodal data through a strictly ordered sequence of operations:

- **Initialization**: Constructs attention and feed-forward layers per modality, plus learnable AdaLN scale-shift tables.
- **Self-Attention**: Applies RMS normalization, AdaZero modulation, and gated self-attention using `attn1`.
- **Text Cross-Attention**: Modulates queries via `_apply_text_cross_attention` and processes text conditioning through `attn2`.
- **Cross-Modal Attention**: When both modalities exist, performs bidirectional A2V and V2A attention with dedicated AdaLN parameters.
- **Feed-Forward**: Concludes with gated MLP processing using modality-specific parameters before returning updated `TransformerArgs`.

This architecture enables the LTX-2 model to handle complex audio-video generation tasks through carefully coordinated normalization and attention mechanisms.

## Frequently Asked Questions

### What is the purpose of the AdaLN scale-shift table in an LTX-2 transformer block?

The `scale_shift_table` stores learnable parameters that generate timestep-conditioned shift, scale, and gate values for adaptive layer normalization. According to the source code in lines 123–147 of [`transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer.py), these tables allow each block to modulate token representations based on diffusion timesteps, enabling the model to adapt its behavior across different noise levels during the generation process.

### How does the LTX-2 transformer block handle single-modality versus multimodal inputs?

The block dynamically constructs processing pathways based on the provided `TransformerConfig` objects during initialization. If only video configuration is provided, the block skips audio-specific layers and bypasses the A2V/V2A cross-attention logic entirely. As implemented in the forward method starting at line 53, the code checks for the existence of each modality before applying the corresponding attention and feed-forward operations.

### Why does the block save snapshots (`vx_pre_av`, `ax_pre_av`) before audio-video cross-attention?

These snapshots prevent order-dependent bias when performing bidirectional cross-attention between audio and video streams. By preserving the original tensor states before either A2V or V2A attention occurs (as seen in lines 28–66), the block ensures that both modalities attend to the same pre-cross-attention representations rather than having the second operation influenced by the first operation's updates.

### What differentiates `attn1` from `attn2` in the LTX-2 transformer block architecture?

`attn1` performs self-attention within a single modality, processing tokens as queries, keys, and values drawn from the same representation. In contrast, `attn2` performs cross-attention where queries come from the video or audio tokens while keys and values come from external text context. This distinction is evident in the constructor (lines 86–122) where `attn2` receives a `context_dim` parameter for text conditioning, whereas `attn1` uses the modality's own dimension.