How LTX-2's Dual-Stream Transformer Handles Audio-Video Synchronization
LTX-2 synchronizes audio and video through a single transformer block containing two parallel streams that process each modality independently before fusing them via temporally-gated cross-attention, ensuring frame-level alignment through per-timestep AdaLN modulation.
The LTX-2 video generation model by Lightricks employs a sophisticated dual-stream transformer architecture to maintain tight synchronization between audio and visual content. This design processes video and audio in parallel streams within a single transformer block, using bidirectional cross-attention mechanisms and adaptive layer normalization to align temporal features at the architecture level. Understanding this implementation reveals how the open-source model achieves precise audio-video coherence without sacrificing modality-specific representations.
Architecture Overview of the Dual-Stream Transformer
At the core of LTX-2's synchronization capability is the BasicAVTransformerBlock class defined in packages/ltx-core/src/ltx_core/model/transformer/transformer.py. This block maintains separate processing paths for video and audio while enabling controlled information exchange between modalities.
Separate Modality Configurations
The block receives distinct TransformerConfig objects for video and audio (lines 86-94), defining dimension sizes, attention heads, and Ada-LN cross-attention settings. This configuration allows the architecture to handle unimodal inputs gracefully—if a modality is absent, its config is set to None, and the block processes only the available stream.
Independent Self-Attention Processing
Each modality first undergoes dedicated self-attention to preserve temporal ordering. According to lines 70-88 in transformer.py, video features pass through self.attn1 and self.attn2, while audio features use self.audio_attn1 and self.audio_attn2. This separation ensures that frame sequences and audio samples retain their internal structure before any cross-modal mixing occurs.
Cross-Modal Synchronization Mechanisms
Following independent processing, the dual-stream transformer aligns modalities through specialized cross-attention layers that operate bidirectionally.
Bidirectional Cross-Attention
The block implements two distinct cross-attention mechanisms (lines 51-74):
- Audio-to-Video (
self.audio_to_video_attn): Video queries attend to audio keys and values, allowing visual features to incorporate relevant sound context. - Video-to-Audio (
self.video_to_audio_attn): Audio queries attend to video keys and values, enabling sound generation to synchronize with visual events.
These layers are instantiated conditionally only when both modalities are present, using the same Attention implementation but swapping modality dimensions to ensure direct temporal correspondence.
Temporally-Gated AdaLN Modulation
Before attention operations, hidden states undergo AdaZero normalization and modulation via learned scale-shift tables (self.scale_shift_table and self.audio_scale_shift_table). As implemented in lines 115-138, the method self.get_av_ca_ada_values extracts per-timestep gate values (gate_out_a2v and gate_out_v2a) from dedicated tables (self.scale_shift_table_a2v_ca_*).
This time-aligned gating mechanism allows the model to amplify or attenuate cross-attention strength based on the exact alignment between audio and video timesteps, creating precise synchronization points.
Residual Integration and Output Processing
After cross-attention, outputs merge back into their respective streams through residual connections (lines 140-166). The cross-attention results are multiplied by gate tensors and optional perturbation masks before addition, ensuring that synchronized information integrates smoothly without overwriting modality-specific features.
Finally, each stream passes through dedicated feed-forward networks (self.ff for video, self.audio_ff for audio) with additional Ada-LN scaling (lines 168-190). This final processing step ensures that residual cross-modal information is properly normalized before the block returns updated TransformerArgs objects for both modalities.
Implementation Example
The following code demonstrates how to instantiate and use the dual-stream transformer block for audio-video synchronization:
# Example: Building a dual‑stream transformer block for a video‑to‑audio task
from ltx_core.model.transformer.transformer import (
BasicAVTransformerBlock,
TransformerConfig,
TransformerOpsConfig,
)
# Video config: 768‑dim, 12 heads, cross‑attention enabled
video_cfg = TransformerConfig(
dim=768,
heads=12,
d_head=64,
context_dim=768, # video‑to‑audio cross‑attention dim
apply_gated_attention=False,
cross_attention_adaln=True,
)
# Audio config: 512‑dim, 8 heads, cross‑attention enabled
audio_cfg = TransformerConfig(
dim=512,
heads=8,
d_head=64,
context_dim=512,
apply_gated_attention=False,
cross_attention_adaln=True,
)
# Optional custom ops (default PyTorch ops are used if omitted)
ops = TransformerOpsConfig()
block = BasicAVTransformerBlock(
video=video_cfg,
audio=audio_cfg,
rope_type="SPLIT", # LTXRopeType.SPLIT
norm_eps=1e-6,
ops=ops,
)
# Dummy inputs (B=2, T=16 for video; B=2, T=64 for audio)
import torch
video_tensor = torch.randn(2, 16, 768) # [batch, frames, dim]
audio_tensor = torch.randn(2, 64, 512) # [batch, samples, dim]
# Wrap tensors in TransformerArgs (simplified version)
from ltx_core.model.transformer.transformer_args import TransformerArgs
video_args = TransformerArgs(
x=video_tensor,
enabled=True,
positional_embeddings=None,
self_attention_mask=None,
cross_positional_embeddings=None,
cross_attn_skip_all=False,
cross_scale_shift_timestep=torch.zeros(2, 1, 1), # placeholder
cross_gate_timestep=torch.zeros(2, 1, 1),
cross_attn_perturbation_mask=1.0,
)
audio_args = TransformerArgs(
x=audio_tensor,
enabled=True,
positional_embeddings=None,
self_attention_mask=None,
cross_positional_embeddings=None,
cross_attn_skip_all=False,
cross_scale_shift_timestep=torch.zeros(2, 1, 1),
cross_gate_timestep=torch.zeros(2, 1, 1),
cross_attn_perturbation_mask=1.0,
)
# Forward pass – returns synchronized video and audio tensors
synced_video, synced_audio = block(video_args, audio_args)
print(synced_video.x.shape, synced_audio.x.shape)
The key source files implementing this architecture include:
packages/ltx-core/src/ltx_core/model/transformer/transformer.py– definesBasicAVTransformerBlockand the core dual-stream logicpackages/ltx-core/src/ltx_core/model/transformer/model.py– constructs the full model usingBasicAVTransformerBlockinstancespackages/ltx-core/src/ltx_core/model/transformer/attention.py– implements the underlyingAttentionclasspackages/ltx-core/src/ltx_core/model/transformer/transformer_args.py– containers for per-layer inputs
Summary
- Dual-stream architecture maintains separate processing paths for video and audio during self-attention, preserving modality-specific temporal structures.
- Bidirectional cross-attention enables direct information flow between streams, with video attending to audio and vice versa.
- Temporally-gated AdaLN provides per-timestep modulation through
get_av_ca_ada_values, allowing precise alignment of audio chunks with video frames. - Conditional instantiation allows the block to handle unimodal inputs by omitting cross-attention when one modality is absent.
- Residual merging with gating mechanisms ensures synchronized information integrates without overwriting original features.
Frequently Asked Questions
What is the primary advantage of a dual-stream transformer over single-stream architectures?
The dual-stream design preserves modality-specific representations through independent self-attention while enabling precise temporal alignment via cross-attention. This prevents early fusion from diluting distinct audio and video features, as implemented in the separate TransformerConfig processing paths of BasicAVTransformerBlock.
How does LTX-2 handle inputs with only video or only audio?
The BasicAVTransformerBlock conditionally instantiates cross-attention layers only when both modalities are present (around line 51 in transformer.py). When processing unimodal input, the unused configuration is set to None, and the block skips the cross-attention stage, processing only the available stream through its dedicated self-attention and feed-forward layers.
What role does AdaLN play in audio-video synchronization?
AdaLN (Adaptive Layer Normalization) provides per-timestep modulation through learned scale-shift tables (self.scale_shift_table_a2v_ca_*). The get_av_ca_ada_values method generates gate tensors (gate_out_a2v, gate_out_v2a) that scale cross-attention based on temporal alignment, allowing the model to emphasize or de-emphasize specific audio-video correspondences at specific timesteps.
Where is the cross-attention mechanism implemented in the codebase?
The bidirectional cross-attention is implemented in packages/ltx-core/src/ltx_core/model/transformer/transformer.py within the BasicAVTransformerBlock class, specifically around lines 51-74. The audio-to-video attention uses self.audio_to_video_attn, while video-to-audio uses self.video_to_audio_attn, both leveraging the Attention class from attention.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →