LTX-2 Transformer Blocks: Architecture, Count, and Operations Explained

LTX-2 utilizes a stack of 48 transformer blocks by default, with each BasicAVTransformerBlock performing self-attention, cross-attention, feed-forward computations, cross-modality attention between audio and video, and AdaLN modulation to enable joint video-audio generation.

The LTX-2 model from the Lightricks/LTX-2 repository processes video and audio streams through a deep transformer architecture designed for synchronized generative tasks. At the heart of this system lies a configurable stack of identical transformer blocks that handle both intra-modality and cross-modality interactions. Understanding the exact count and operational composition of these LTX-2 transformer blocks is essential for researchers fine-tuning the model or adapting it for custom diffusion-based generation.

How Many Transformer Blocks Does LTX-2 Use?

The default LTX-2 configuration employs 48 transformer blocks, controlled by the num_layers parameter in the LTXModel class constructor. According to the source code in packages/ltx-core/src/ltx_core/model/transformer/model.py, the initialization accepts num_layers: int = 48 and constructs the stack using a torch.nn.ModuleList:


# packages/ltx-core/src/ltx_core/model/transformer/model.py

def __init__(..., num_layers: int = 48, ...):
    ...
    self._init_transformer_blocks(num_layers=num_layers, ...)

During construction, the model instantiates 48 identical BasicAVTransformerBlock modules:


# packages/ltx-core/src/ltx_core/model/transformer/model.py

self.transformer_blocks = torch.nn.ModuleList(
    [BasicAVTransformerBlock(video=video_config,
                            audio=audio_config,
                            rope_type=self.rope_type,
                            norm_eps=norm_eps,
                            ops=ops)
     for _ in range(num_layers)]
)

This produces a vanilla LTX-2 model containing 48 sequential layers, each capable of processing both video and audio tokens simultaneously.

Core Operations Performed by Each Block

Every BasicAVTransformerBlock in the LTX-2 transformer stack executes a comprehensive set of operations that enable both independent modality processing and cross-modal information exchange. These operations are implemented in packages/ltx-core/src/ltx_core/model/transformer/transformer.py.

Intra-Modality Attention and Feed-Forward Networks

Each block maintains separate processing pathways for video and audio tokens:

  • Self-attention: self.attn1 processes video tokens, while self.audio_attn1 handles audio tokens, performing standard multi-head attention within each modality.
  • Cross-attention to text: self.attn2 (video) and self.audio_attn2 (audio) attend to conditioning text embeddings supplied via context_dim.
  • Feed-forward networks: self.ff and self.audio_ff apply position-wise MLP transformations after attention operations.

Cross-Modality Information Flow

To synchronize audio and video generation, each block implements bidirectional cross-attention:

  • Audio-to-video attention: self.audio_to_video_attn allows video queries to attend to audio keys and values.
  • Video-to-audio attention: self.video_to_audio_attn enables audio queries to attend to video representations.

These components ensure that lip movements, sound effects, and musical elements remain temporally aligned throughout the generation process.

Conditioning and Normalization Mechanisms

Each transformer block incorporates advanced conditioning techniques:

  • AdaLN modulation: Learnable scale_shift_table and audio_scale_shift_table parameters combine with timestep embeddings through AdaLayerNorm (defined in packages/ltx-core/src/ltx_core/model/transformer/adaln.py) to modulate activations based on the diffusion denoising step.
  • Rotary Position Embeddings (RoPE): Configured via the rope_type parameter (defaulting to SPLIT), these provide relative position information to attention heads.
  • Gated attention: When enabled via apply_gated_attention, a learned scalar gate modulates attention scores for finer control over information flow.

Implementation in the LTX-2 Codebase

The transformer architecture spans several key files in the repository:

File Role
packages/ltx-core/src/ltx_core/model/transformer/model.py Defines LTXModel, configures the number of blocks via num_layers, and initializes the ModuleList.
packages/ltx-core/src/ltx_core/model/transformer/transformer.py Implements BasicAVTransformerBlock containing all attention and feed-forward operations.
packages/ltx-core/src/ltx_core/model/transformer/attention.py Provides the Attention module used by both self-attention and cross-attention paths.
packages/ltx-core/src/ltx_core/model/transformer/adaln.py Implements the adaptive layer normalization mechanism.

Configuring the Transformer Stack

You can inspect or modify the LTX-2 transformer block count programmatically. The following example demonstrates creating a model and examining its architecture:

import torch
from ltx_core.model.transformer.model import LTXModel, LTXModelType

# Create a vanilla LTX-2 model with 48 blocks

model = LTXModel(
    model_type=LTXModelType.AudioVideo,   # Enable both video & audio streams

    num_layers=48,                        # Default depth, customizable

    num_attention_heads=32,
    attention_head_dim=128,
    in_channels=128,
    out_channels=128,
)

# Verify the number of transformer blocks

print("Number of transformer blocks:", len(model.transformer_blocks))

# Inspect the sub-modules within the first block

first_block = model.transformer_blocks[0]
print("Sub-modules in a block:", list(first_block._modules.keys()))

Typical output confirms the 48-layer structure and dual-modality design:


Number of transformer blocks: 48
Sub-modules in a block: ['attn1', 'attn2', 'ff',
                         'audio_attn1', 'audio_attn2', 'audio_ff',
                         'audio_to_video_attn', 'video_to_audio_attn',
                         'scale_shift_table', 'audio_scale_shift_table',
                         ...]

Summary

Frequently Asked Questions

Can I change the number of transformer blocks in LTX-2?

Yes. When initializing LTXModel in packages/ltx-core/src/ltx_core/model/transformer/model.py, pass a custom integer to the num_layers argument. While the default is 48 for the standard model, you can reduce this for faster inference or increase it for potentially higher-quality generation at the cost of computational resources.

What is the difference between self-attention and cross-modality attention in LTX-2?

Self-attention (via attn1 and audio_attn1) operates within a single modality, allowing video tokens to attend only to other video tokens (or audio to audio). Cross-modality attention (via audio_to_video_attn and video_to_audio_attn) enables information flow between modalities, allowing video features to influence audio generation and vice versa, which is crucial for synchronized audio-video output.

How does AdaLN modulation work in LTX-2 transformer blocks?

Each block contains learnable parameters in scale_shift_table and audio_scale_shift_table that are combined with diffusion timestep embeddings through the AdaLayerNorm mechanism. This adaptive normalization scales and shifts the layer inputs based on the current denoising step, allowing the model to modulate its behavior throughout the reverse diffusion process.

Where are the transformer blocks defined in the LTX-2 codebase?

The block count configuration resides in packages/ltx-core/src/ltx_core/model/transformer/model.py within the LTXModel.__init__ method. The actual implementation of BasicAVTransformerBlock, including all attention mechanisms and feed-forward networks, is located in packages/ltx-core/src/ltx_core/model/transformer/transformer.py. Support utilities like attention mechanisms and AdaLN are found in the adjacent attention.py and adaln.py files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →