How LTX-2's Asymmetric Dual-Stream Transformer Architecture Processes Video and Audio Jointly
LTX-2 uses an asymmetric dual-stream transformer architecture that processes video and audio through separate transformer blocks with independent hidden dimensions and attention heads, while enabling controlled cross-modal interaction via directional cross-attention layers.
The LTX-2 model from Lightricks introduces a novel approach to multimodal generation through its asymmetric dual-stream transformer architecture. This design allows the model to handle video and audio modalities with specialized parameters while maintaining efficient information flow between them. Unlike symmetric architectures that force identical configurations across modalities, LTX-2's implementation treats each stream independently while preserving the ability to fuse representations through gated cross-attention mechanisms.
Core Components and File Structure
BasicAVTransformerBlock Implementation
Located in packages/ltx-core/src/ltx_core/model/transformer/transformer.py at line 86, the BasicAVTransformerBlock class serves as the fundamental building block. It hosts separate video and audio sub-blocks within a single transformer block, enabling parallel processing of both modalities while allowing controlled interaction between them.
TransformerConfig and Asymmetric Parameters
The TransformerConfig class (defined at line 30 of transformer.py) holds modality-specific hyperparameters including dimension sizes, head counts, and context dimensions. This configuration enables the architectural asymmetry where video might use 32 heads of 128 dimensions each (4096 inner dimension) while audio uses 32 heads of 64 dimensions (2048 inner dimension). Cross-attention dimensions and gating timesteps remain independent, providing fine-grained control over modality interaction.
LTXModel Assembly
The top-level LTXModel class in packages/ltx-core/src/ltx_core/model/transformer/model.py (line 41) constructs a stack of 48 BasicAVTransformerBlock instances by default. During initialization, it creates separate patchify projections (self.patchify_proj for video, self.audio_patchify_proj for audio) and instantiates per-modality AdaLN layers with distinct scale-shift tables.
Dual-Stream Data Flow and Mechanisms
Self-Attention Paths
Each stream processes independently through dedicated attention mechanisms. Video flows through self.attn1 (self-attention) and self.attn2 (text-cross-attention), while audio uses self.audio_attn1 and self.audio_attn2. Both paths apply AdaLayerNormSingle modulation using per-modality scale-shift tables (e.g., self.scale_shift_table) before attention operations.
Directional Cross-Attention
Cross-attention operates bidirectionally but asymmetrically within each block. The self.audio_to_video_attn module queries video tokens using audio keys and values, while self.video_to_audio_attn queries audio tokens using video keys and values. Each direction maintains separate scale-shift tables (scale_shift_table_a2v_ca_video, scale_shift_table_a2v_ca_audio) and gating parameters to control the influence each modality has on the other.
Feed-Forward and Output Projection
Following attention, each stream passes through independent feed-forward networks (self.ff for video, self.audio_ff for audio). Final outputs are projected via self.proj_out (video) and self.audio_proj_out (audio), preserving the dimensional asymmetry throughout the entire forward pass.
Modality Modes and Runtime Configuration
The TransformerArgs dataclass in packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py (line 16) carries runtime tensors, masks, and timesteps through the blocks. The LTXModel supports three operational modes via the LTXModelType enum:
- Audio-Video: Both streams active with full cross-attention enabled
- Video-Only: Audio config omitted; blocks behave as standard video transformers
- Audio-Only: Video config omitted; blocks behave as standard audio transformers
This flexibility allows the same architecture to handle unimodal generation without requiring separate model implementations.
Implementation Example
The following example demonstrates creating a dual-stream LTX-2 model with asymmetric dimensions and processing dummy video and audio tensors:
import torch
from ltx_core.model.transformer.model import LTXModel, LTXModelType
# Dummy video tensor: (batch, frames, height, width, channels) → flattened to (B, N, C)
video_tensor = torch.randn(2, 16, 64, 64, 128) # B=2, T=16, C=128
video_tensor = video_tensor.view(2, 16, -1) # (B, T, C)
# Dummy audio tensor: (batch, time, channels)
audio_tensor = torch.randn(2, 200, 128) # B=2, T=200, C=128
# Build a dual‑stream LTX model (default settings)
model = LTXModel(
model_type=LTXModelType.AudioVideo,
num_attention_heads=32,
attention_head_dim=128,
audio_num_attention_heads=32,
audio_attention_head_dim=64,
in_channels=128,
audio_in_channels=128,
)
# Forward pass – returns (video_output, audio_output)
video_out, audio_out = model(video=video_tensor, audio=audio_tensor)
print("Video latent shape:", video_out.shape) # e.g. (2, 16, 128)
print("Audio latent shape:", audio_out.shape) # e.g. (2, 200, 128)
The example demonstrates the symmetric API (model(video=…, audio=…)) while the internal block layout remains asymmetric, with video and audio streams maintaining their distinct dimensional configurations throughout processing.
Summary
- LTX-2's asymmetric dual-stream transformer architecture processes video and audio through separate transformer blocks with independent configurations, allowing each modality to maintain optimal hyperparameters for its specific characteristics.
- Cross-attention mechanisms operate bidirectionally with separate scale-shift tables and gating parameters, enabling controlled information flow between modalities without forcing dimensional alignment.
- Three operational modes (Audio-Video, Video-Only, Audio-Only) allow the same architecture to handle unimodal or multimodal inputs by selectively enabling stream-specific components.
- Modality-specific AdaLN layers apply timestep-dependent scaling and shifting through separate parameters, ensuring proper diffusion model conditioning for each stream independently.
Frequently Asked Questions
What makes LTX-2's dual-stream architecture "asymmetric"?
The architecture is asymmetric because the video and audio streams can have different hidden dimensions, attention head counts, and positional embedding configurations. For example, video might use 4096 dimensions (32 heads × 128 dimensions) while audio uses 2048 dimensions (32 heads × 64 dimensions), with independent cross-attention dimensions and gating parameters for each direction.
How does cross-attention work between the video and audio streams?
Cross-attention operates through dedicated modules within BasicAVTransformerBlock: audio_to_video_attn allows audio tokens to query video representations, while video_to_audio_attn allows video tokens to query audio representations. Each direction uses separate scale-shift tables (scale_shift_table_a2v_ca_video, scale_shift_table_a2v_ca_audio) to control the gating and influence of the cross-modal interaction.
Can LTX-2 operate with only video or only audio input?
Yes, the LTXModel supports three modes via the LTXModelType enum. When initialized as Video-Only, the audio configuration is omitted and blocks behave as standard video transformers. In Audio-Only mode, the video configuration is omitted. This flexibility allows the same architecture to handle unimodal generation without requiring separate model implementations.
Where are the AdaLN parameters defined for each modality?
AdaLayerNorm parameters are defined in packages/ltx-core/src/ltx_core/model/transformer/adaln.py and instantiated separately for each stream within LTXModel. Each modality maintains its own scale-shift tables (e.g., self.scale_shift_table for video, audio_scale_shift_table for audio) and additional tables for cross-attention gates, enabling timestep-dependent modulation that respects the asymmetric dimensional settings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →