LTX-2 Transformer Block: Complete Sequence of Operations Explained

Each LTX-2 transformer block (BasicAVTransformerBlock) processes multimodal tokens through a rigid pipeline of AdaLN-conditioned self-attention, text cross-attention, optional audio-video cross-attention, and gated feed-forward layers before returning updated representations.

The BasicAVTransformerBlock class in the Lightricks/LTX-2 repository defines the core computational unit for video and audio generation. An LTX-2 transformer block handles both unimodal and multimodal inputs using adaptive layer normalization (AdaLN) and specialized attention mechanisms configured via TransformerConfig objects. Understanding the exact sequence of operations within these blocks is essential for modifying generation behavior or optimizing inference pipelines.

LTX-2 Transformer Block Construction and Initialization

Before the forward pass begins, the block initializes its components based on provided modality configurations.

Configuration and Layer Instantiation

The constructor receives optional video and audio TransformerConfig objects, a rotary-embedding type (LTXRopeType), an epsilon value for RMS normalization, and pluggable operations (TransformerOpsConfig). For each modality present, the block instantiates self-attention (attn1), cross-attention (attn2), and a feed-forward network (ff) according to the configuration dimensions and attention heads. This initialization occurs in lines 86–109 and 111–122 of the transformer module.

AdaLN Scale-Shift Table Allocation

For every enabled modality, the block allocates learnable scale_shift_table parameters. These tables store per-timestep affine transformation parameters that modulate token representations throughout the forward pass. When cross-attention AdaLN is enabled, additional prompt-specific tables are also initialized. This setup appears in lines 123–129 and 136–147 of transformer.py.

Forward Pass Operations in LTX-2 Transformer Blocks

The forward method (starting at line 53) executes an eight-step sequence for each active modality. The operations proceed in a fixed order: self-attention, text cross-attention, audio-video cross-attention (when both modalities are present), and finally feed-forward processing.

Step 1: Self-Attention with AdaZero Modulation

The first transformation applied to incoming tokens is multi-head self-attention (MSA) with adaptive normalization.

  • AdaLN Parameter Extraction: The block calls get_ada_values to extract three parameters—shift, scale, and gate—from the modality’s scale_shift_table based on the current timestep.
  • RMS Normalization and AdaZero: Input tokens undergo RMS normalization, then modulation via the ada_zero_function, which applies the extracted shift and scale parameters.
  • Attention Computation: The normalized tensor passes through attn1 (the first attention layer), utilizing optional positional embeddings and self-attention masks.
  • Gating: The post_sa_function applies RMS normalization again and multiplies the output by the gate parameter before the residual connection.

For video inputs, this flow occupies lines 70–85; for audio inputs, lines 99–114.

Step 2: Text Cross-Attention Implementation

Following self-attention, the block processes text conditioning through _apply_text_cross_attention.

This method receives the normalized representation (vx_normed for video, ax_normed for audio) and performs three sub-operations: it optionally extracts additional AdaLN parameters for the query (lines 38–40), applies apply_cross_attention_adaln to modulate both the query and context tensors, and executes the second attention layer (attn2 or audio_attn2). The implementation spans lines 22–34 (definition) and lines 86–96 and 115–125 (usage).

Step 3: Audio-Video Cross-Attention (A2V and V2A)

When both video and audio configurations are provided, the LTX-2 transformer block performs bidirectional cross-attention between modalities.

  • Snapshot Preservation: The code saves pre-cross-attention tensors (vx_pre_av, ax_pre_av) to prevent order-dependent bias in the updates.
  • Audio-to-Video (A2V): The block extracts dedicated AdaLN parameters (scale_ca_video_a2v, shift_ca_video_a2v) from a specific table, applies AdaZero scaling, and runs audio_to_video_attn using audio tokens as keys and values while attending to video queries.
  • Video-to-Audio (V2A): The symmetric operation uses video_to_audio_attn with corresponding AdaLN parameters to update audio representations based on video content.

Both directions apply gating before adding the residual update. This cross-modal interaction is implemented in lines 28–66 of transformer.py.

Step 4: Feed-Forward Network Processing

The final operation within each LTX-2 transformer block is the feed-forward (MLP) sub-layer.

After all attention computations complete, the block fetches MLP-specific AdaLN parameters (shift_mlp, scale_mlp, gate_mlp) from lines 98–101. Tokens undergo AdaZero scaling, pass through the modality-specific feed-forward network (ff for video, audio_ff for audio), receive gate multiplication, and merge back via residual connection. The video feed-forward logic appears in lines 98–108, while audio processing occupies lines 111–121.

Output Generation

The forward method concludes at line 416 by returning updated TransformerArgs objects containing the transformed tensors. If a specific modality was not enabled in the configuration, the method returns None for that argument.

Code Example: Running a Single LTX-2 Transformer Block

The following example demonstrates how to instantiate and execute one LTX-2 transformer block with both video and audio pathways enabled.

import torch
from ltx_core.model.transformer.transformer import BasicAVTransformerBlock, TransformerConfig
from ltx_core.model.transformer.transformer_args import TransformerArgs

# Configure per-modality parameters

video_cfg = TransformerConfig(
    dim=512, heads=8, d_head=64,
    context_dim=768, apply_gated_attention=False,
    cross_attention_adaln=True
)

audio_cfg = TransformerConfig(
    dim=256, heads=4, d_head=64,
    context_dim=384, apply_gated_attention=False,
    cross_attention_adaln=False
)

# Initialize the transformer block

block = BasicAVTransformerBlock(video=video_cfg, audio=audio_cfg)

# Prepare dummy input tensors

batch, seq_len = 2, 16

video_args = TransformerArgs(
    x=torch.randn(batch, seq_len, video_cfg.dim),
    timesteps=torch.randn(batch, 1, 1),
    positional_embeddings=torch.randn(batch, seq_len, video_cfg.dim),
    self_attention_mask=None,
    self_attn_perturbation_mask=None,
    self_attn_all_perturbed=False,
    context=None,
    context_mask=None,
    cross_scale_shift_timestep=torch.randn(batch, 1, 1),
    cross_gate_timestep=torch.randn(batch, 1, 1),
    cross_attn_skip_all=False,
    enabled=True,
)

audio_args = TransformerArgs(
    x=torch.randn(batch, seq_len, audio_cfg.dim),
    timesteps=torch.randn(batch, 1, 1),
    positional_embeddings=torch.randn(batch, seq_len, audio_cfg.dim),
    self_attention_mask=None,
    self_attn_perturbation_mask=None,
    self_attn_all_perturbed=False,
    context=None,
    context_mask=None,
    cross_scale_shift_timestep=torch.randn(batch, 1, 1),
    cross_gate_timestep=torch.randn(batch, 1, 1),
    cross_attn_skip_all=False,
    enabled=True,
)

# Execute forward pass

new_video, new_audio = block(video_args, audio_args)

print(new_video.x.shape)   # Output: torch.Size([2, 16, 512])

print(new_audio.x.shape)   # Output: torch.Size([2, 16, 256])

Key Source Files and Implementation Details

Understanding the LTX-2 transformer block requires familiarity with these specific source files in the repository:

Summary

An LTX-2 transformer block processes multimodal data through a strictly ordered sequence of operations:

  • Initialization: Constructs attention and feed-forward layers per modality, plus learnable AdaLN scale-shift tables.
  • Self-Attention: Applies RMS normalization, AdaZero modulation, and gated self-attention using attn1.
  • Text Cross-Attention: Modulates queries via _apply_text_cross_attention and processes text conditioning through attn2.
  • Cross-Modal Attention: When both modalities exist, performs bidirectional A2V and V2A attention with dedicated AdaLN parameters.
  • Feed-Forward: Concludes with gated MLP processing using modality-specific parameters before returning updated TransformerArgs.

This architecture enables the LTX-2 model to handle complex audio-video generation tasks through carefully coordinated normalization and attention mechanisms.

Frequently Asked Questions

What is the purpose of the AdaLN scale-shift table in an LTX-2 transformer block?

The scale_shift_table stores learnable parameters that generate timestep-conditioned shift, scale, and gate values for adaptive layer normalization. According to the source code in lines 123–147 of transformer.py, these tables allow each block to modulate token representations based on diffusion timesteps, enabling the model to adapt its behavior across different noise levels during the generation process.

How does the LTX-2 transformer block handle single-modality versus multimodal inputs?

The block dynamically constructs processing pathways based on the provided TransformerConfig objects during initialization. If only video configuration is provided, the block skips audio-specific layers and bypasses the A2V/V2A cross-attention logic entirely. As implemented in the forward method starting at line 53, the code checks for the existence of each modality before applying the corresponding attention and feed-forward operations.

Why does the block save snapshots (vx_pre_av, ax_pre_av) before audio-video cross-attention?

These snapshots prevent order-dependent bias when performing bidirectional cross-attention between audio and video streams. By preserving the original tensor states before either A2V or V2A attention occurs (as seen in lines 28–66), the block ensures that both modalities attend to the same pre-cross-attention representations rather than having the second operation influenced by the first operation's updates.

What differentiates attn1 from attn2 in the LTX-2 transformer block architecture?

attn1 performs self-attention within a single modality, processing tokens as queries, keys, and values drawn from the same representation. In contrast, attn2 performs cross-attention where queries come from the video or audio tokens while keys and values come from external text context. This distinction is evident in the constructor (lines 86–122) where attn2 receives a context_dim parameter for text conditioning, whereas attn1 uses the modality's own dimension.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →