How LTX-2 Distributes Model Parameters Between Video and Audio Streams
LTX-2 allocates separate transformer parameters for video and audio streams, using distinct self-attention, feed-forward, and AdaLN modulation layers for each modality, plus bidirectional cross-attention modules that enable interaction without sharing weights.
The LTX-2 architecture from Lightricks implements a dual-branch transformer design that processes video and audio through parallel yet independent parameter sets. Unlike unified multimodal models that share weights across modalities, LTX-2 maintains modality-specific layers while enabling cross-modal attention through dedicated projection matrices. This article examines exactly how parameters are distributed between the video and audio streams based on the official source code in packages/ltx-core/src/ltx_core/model/transformer/transformer.py.
Parallel Transformer Branches with Separate Parameter Sets
LTX-2 instantiates two parallel transformer branches that share architectural logic but maintain completely separate learnable parameters. Each modality has its own self-attention mechanisms and feed-forward networks, ensuring that video and audio representations develop independently until explicit cross-modal fusion occurs.
Self-Attention Layers
For the video stream, the transformer defines self.attn1 and self.attn2 as the primary self-attention modules. The audio stream receives dedicated counterparts named self.audio_attn1 and self.audio_attn2, initialized at lines 151-166 of the transformer implementation.
Because these modules instantiate separate Attention classes, each maintains independent query, key, value, and output projection matrices. The video and audio branches do not share attention weights, allowing each modality to learn modality-specific attention patterns.
Feed-Forward Networks
Each branch contains its own feed-forward network. The video path uses self.ff (defined in the shared code path), while the audio path uses self.audio_ff (lines 147-148). These are distinct FeedForward instances with separate weight matrices and bias terms, doubling the FFN parameter count compared to a single shared network.
Cross-Modal Attention Mechanisms
Cross-modal interaction occurs through two unidirectional attention modules, each contributing a full set of projection parameters.
Audio-to-Video Attention
The self.audio_to_video_attn module (lines 152-162) enables the video stream to attend to audio features. This module creates query projections from video tokens and key/value projections from audio tokens, adding a complete set of Q, K, V, and output linear layers distinct from the self-attention parameters.
Video-to-Audio Attention
Conversely, self.video_to_audio_attn (lines 164-174) allows audio queries to attend to video keys and values. This symmetrical design means cross-modal parameters effectively double compared to a single-direction implementation, as each direction requires independent projection matrices.
AdaLN Modulation Tables
Adaptive Layer Normalization (AdaLN) parameters follow the same separation pattern. The video stream uses self.scale_shift_table, while audio uses self.audio_scale_shift_table.
For cross-modal attention, dedicated modulation tables condition the layer outputs: self.scale_shift_table_a2v_ca_video and self.scale_shift_table_a2v_ca_audio (lines 176-177). These tables provide learnable scale and shift parameters that modulate the cross-attention outputs independently for each direction.
Configuration-Driven Dimensionality
Parameter counts scale according to configuration values defined in training YAML files such as video_suffix_lora.yaml and audio_suffix_lora.yaml. The TransformerArgs dataclass in packages/ltx-core/src/ltx_core/types.py holds these values:
video:
dim: 768 # hidden size for video tokens
heads: 12 # number of attention heads
d_head: 64 # per-head dimension
audio:
dim: 512 # hidden size for audio tokens
heads: 8
d_head: 64
These dimensions determine the sizes of linear layers inside Attention and FeedForward modules. A video token with dim=768 and 12 heads produces larger projection matrices than an audio token with dim=512 and 8 heads, creating an asymmetric parameter distribution that favors the video stream when using these default values.
Dynamic Parameter Budget Adjustment
The architecture supports dynamic parameter budgets through conditional initialization. Lines 58-60 in the forward method check whether audio is enabled, allowing the model to run as video-only when audio=None.
# Example: constructing a transformer with both video and audio streams
from ltx_core.model.transformer.transformer import Transformer
from ltx_core.types import TransformerArgs
video_cfg = TransformerArgs(
dim=768,
heads=12,
d_head=64,
# … other video-specific settings …
)
audio_cfg = TransformerArgs(
dim=512,
heads=8,
d_head=64,
# … other audio-specific settings …
)
model = Transformer(video=video_cfg, audio=audio_cfg)
# Forward pass – the model will route video tokens through the video branch,
# audio tokens through the audio branch, and then perform cross-attention.
video_out, audio_out = model(video=video_input, audio=audio_input)
When training video-only models, setting audio=None eliminates all audio-specific parameters (audio_attn1, audio_attn2, audio_ff, audio_scale_shift_table, and both cross-attention modules), reducing the total parameter count proportionally.
# Example: disabling the audio branch (video-only model)
model = Transformer(video=video_cfg, audio=None)
video_out, _ = model(video=video_input, audio=None) # only video path runs
Inspecting Parameter Distribution
You can verify the parameter allocation programmatically:
# Example: inspecting parameter counts per branch
total_params = sum(p.numel() for p in model.parameters())
video_params = sum(p.numel() for n, p in model.named_parameters() if "audio" not in n)
audio_params = sum(p.numel() for n, p in model.named_parameters() if "video" not in n and "audio" in n)
print(f"Total: {total_params:,} | Video: {video_params:,} | Audio: {audio_params:,}")
Summary
- Separate linear projections – Video and audio maintain independent Q, K, V, and output projection matrices with no shared weights.
- Independent feed-forward networks – Each modality has its own
FeedForwardinstance (self.ffvsself.audio_ff). - Bidirectional cross-attention – Two distinct modules (
audio_to_video_attnandvideo_to_audio_attn) each contribute a full attention block's worth of parameters. - Modality-specific AdaLN tables – Scale-shift parameters are separated into video, audio, and cross-modal-specific tables.
- Configurable dimensions – Parameter counts scale according to
dim,heads, andd_headvalues in configuration files. - Runtime flexibility – The audio branch can be disabled entirely by passing
audio=None, dynamically reducing the active parameter count.
Frequently Asked Questions
Do video and audio streams share any parameters in LTX-2?
No, the video and audio streams maintain completely separate parameter sets according to the source code in transformer.py. Each modality has its own self-attention projections (attn1/attn2 vs audio_attn1/audio_attn2), feed-forward networks (ff vs audio_ff), and AdaLN modulation tables. The only interaction occurs through cross-attention modules, which still use distinct projection matrices for each direction rather than sharing weights.
How does cross-attention affect the total parameter count?
Cross-attention adds a full self-attention block for each direction. Because LTX-2 implements bidirectional cross-attention (audio_to_video_attn and video_to_audio_attn), it doubles the cross-modal parameter count compared to a single-direction design. Each cross-attention module includes independent query, key, value, and output projections, plus dedicated AdaLN modulation tables (scale_shift_table_a2v_ca_video and scale_shift_table_a2v_ca_audio).
Can I train a video-only model using the LTX-2 architecture?
Yes, the architecture supports video-only training by setting audio=None when instantiating the Transformer class. The forward pass at lines 58-60 includes conditional checks that skip audio processing when the audio configuration is missing. This eliminates all audio-specific parameters (self-attention, feed-forward, and cross-attention modules) from the active computation graph and parameter count.
Where are the dimensionality settings configured for each stream?
Dimensionality settings are defined in YAML configuration files such as video_suffix_lora.yaml and audio_suffix_lora.yaml within packages/ltx-trainer/configs/. These values populate the TransformerArgs dataclass from packages/ltx-core/src/ltx_core/types.py, which then drives the construction of linear layers in the transformer. The dim, heads, and d_head parameters determine the width and depth of projection matrices for each modality independently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →