How Cross-Modality AdaLN Synchronizes Audio and Video Streams in LTX-2
Cross-modality AdaLN synchronizes audio and video in LTX-2 by extracting shared timestep-specific scale, shift, and gate parameters that apply identical temporal conditioning to both modalities before cross-attention, ensuring their hidden states remain normalized to the same generative timestep.
The LTX-2 video generation model developed by Lightricks processes audio and video through a unified transformer architecture that enables bidirectional information flow. By implementing cross-modality Adaptive Layer Normalization (AdaLN), the system ensures that audio-to-video and video-to-audio cross-attention operations occur with temporally aligned normalization parameters, preventing drift between the heterogeneous streams during generation.
Parameter Extraction and Computation
Computing AdaLN Dimensions in adaln.py
According to the Lightricks/LTX-2 source code, each transformer block maintains a table of scale and shift parameters managed by adaln_embedding_coefficient(). When cross-attention AdaLN is enabled, this function increases the parameter count by three additional values per block to accommodate the cross-modality modulation tensors. In packages/ltx-core/src/ltx_core/model/transformer/adaln.py, the AdaLayerNormSingle class serves as the base implementation for these adaptive normalization layers, providing the infrastructure for timestep-conditioned scaling.
Retrieving Cross-Modality Values
Before executing cross-attention, the model extracts timestep-specific normalization parameters using get_av_ca_ada_values(). This function, defined in packages/ltx-core/src/ltx_core/model/transformer/transformer.py, returns scale, shift, and gate tensors for the current timestep (cross_scale_shift_timestep and cross_gate_timestep). Because both audio and video streams reference the same temporal indices, the extracted values (scale_ca_video_a2v, shift_ca_video_a2v, gate_out_a2v for video-to-audio attention, and corresponding audio-side parameters) remain synchronized across modalities.
Applying Synchronous Modulation
Affine Transformation with ada_zero_function
The raw hidden states undergo modulation through ada_zero_function, which applies an affine transformation (shift plus scale) while preserving the residual connection. This operation ensures that queries, keys, and values entering the cross-attention mechanism are normalized against the same temporal reference point. The function applies the scale and shift tensors extracted via get_av_ca_ada_values() to the source modality's hidden states immediately before cross-attention.
Bidirectional Cross-Attention (A2V and V2A)
LTX-2 implements bidirectional cross-modality attention through audio_to_video_attn and video_to_audio_attn operations. For audio-to-video (A2V) cross-attention, the video stream receives modulation parameters (scale_ca_video_a2v, shift_ca_video_a2v) while the audio stream receives corresponding audio-side parameters. Both streams share the identical cross_scale_shift_timestep value, guaranteeing that the normalization applied to video keys aligns temporally with the normalization applied to audio queries.
# Extracting AdaLN parameters for A2V cross-attention with shared timestep
scale_ca_video_a2v, shift_ca_video_a2v, gate_out_a2v = self.get_av_ca_ada_values(
self.scale_shift_table_a2v_ca_video, # table for video side
vx.shape[0], # batch size
video.cross_scale_shift_timestep, # shared temporal index
video.cross_gate_timestep, # shared gate timestep
slice(0, 2), # indices for video scale/shift
)
# Apply AdaLN to video hidden state before cross-attention
a2v_vx_scaled = self.ada_zero_function(
vx_pre_av, self.norm_eps, scale_ca_video_a2v, shift_ca_video_a2v
)
# Cross-attention with temporal gating
vx = vx + (
self.audio_to_video_attn(
a2v_vx_scaled,
context=a2v_ax_scaled,
pe=video.cross_positional_embeddings,
k_pe=audio.cross_positional_embeddings,
)
* gate_out_a2v
* video.cross_attn_perturbation_mask
)
Gate-Controlled Blending and Temporal Lockstep
The synchronization mechanism relies on gate-controlled blending to determine fusion strength. The gate_out_a2v (or gate_out_v2a) scalar, derived from the shared cross_gate_timestep, multiplies the cross-attention result before it is added to the target modality's residual stream. This gating mechanism allows the model to learn when to fuse audio-driven visual updates and when to preserve modality independence, while ensuring that the decision occurs at the same temporal step for both streams.
Because the AdaLN parameters are timestep-dependent, the identical temporal index used for both audio and video guarantees that modulation applied to the video side aligns precisely with modulation applied to the audio side. Consequently, the cross-attention operations receive synchronously normalized queries, keys, and values, enabling temporally coherent blending of audio and video representations throughout the generation pipeline.
Summary
- Cross-modality AdaLN in LTX-2 synchronizes audio and video through shared timestep-dependent scale, shift, and gate parameters.
- The
adaln_embedding_coefficient()function inadaln.pyreserves three additional parameters per block when cross-attention AdaLN is enabled. get_av_ca_ada_values()extracts synchronized normalization tensors using the sharedcross_scale_shift_timestepfor both modalities.ada_zero_functionapplies affine transformations to hidden states before cross-attention, ensuring both streams use the same temporal reference.- Gate scalars (
gate_out_a2v,gate_out_v2a) control the strength of cross-modal fusion at each timestep, maintaining temporal lockstep between audio and video updates.
Frequently Asked Questions
What is cross-modality AdaLN in LTX-2?
Cross-modality AdaLN is an adaptive normalization mechanism that conditions both audio and video transformer streams on shared timestep embeddings. It extracts scale, shift, and gate parameters from the same temporal indices to ensure that cross-attention between modalities occurs with synchronized normalization, preventing temporal misalignment during joint generation.
How does LTX-2 ensure audio and video remain temporally aligned?
LTX-2 maintains temporal alignment by using identical cross_scale_shift_timestep and cross_gate_timestep values when extracting AdaLN parameters for both modalities. This ensures that the affine transformations applied to audio and video hidden states reference the same generative timestep, so their cross-attention interactions remain temporally coherent.
What role does the gate parameter play in cross-modality synchronization?
The gate parameter (extracted as gate_out_a2v or gate_out_v2a) scales the cross-attention output before it is added to the target modality's residual stream. Because this gate is derived from the shared timestep, it controls how much cross-modal information flows between audio and video at each specific generation step, ensuring that fusion strength remains synchronized across both streams.
Where is cross-modality AdaLN implemented in the codebase?
The core implementation resides in packages/ltx-core/src/ltx_core/model/transformer/adaln.py (parameter computation and AdaLayerNormSingle), with the extraction logic in packages/ltx-core/src/ltx_core/model/transformer/transformer.py (specifically get_av_ca_ada_values() and the cross-attention forward pass). The affine transformation utility ada_zero_function is typically located in the utilities module or within the transformer implementation file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →