# How Cross-Modality AdaLN Synchronizes Audio and Video Streams in LTX-2

> Discover how cross-modality AdaLN synchronizes audio and video in LTX-2. Learn how shared parameters ensure identical temporal conditioning for normalized generative timesteps.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: internals
- Published: 2026-06-21

---

**Cross-modality AdaLN synchronizes audio and video in LTX-2 by extracting shared timestep-specific scale, shift, and gate parameters that apply identical temporal conditioning to both modalities before cross-attention, ensuring their hidden states remain normalized to the same generative timestep.**

The LTX-2 video generation model developed by Lightricks processes audio and video through a unified transformer architecture that enables bidirectional information flow. By implementing **cross-modality Adaptive Layer Normalization (AdaLN)**, the system ensures that audio-to-video and video-to-audio cross-attention operations occur with temporally aligned normalization parameters, preventing drift between the heterogeneous streams during generation.

## Parameter Extraction and Computation

### Computing AdaLN Dimensions in adaln.py

According to the Lightricks/LTX-2 source code, each transformer block maintains a table of scale and shift parameters managed by `adaln_embedding_coefficient()`. When cross-attention AdaLN is enabled, this function increases the parameter count by three additional values per block to accommodate the cross-modality modulation tensors. In [`packages/ltx-core/src/ltx_core/model/transformer/adaln.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/adaln.py), the `AdaLayerNormSingle` class serves as the base implementation for these adaptive normalization layers, providing the infrastructure for timestep-conditioned scaling.

### Retrieving Cross-Modality Values

Before executing cross-attention, the model extracts timestep-specific normalization parameters using `get_av_ca_ada_values()`. This function, defined in [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py), returns **scale**, **shift**, and **gate** tensors for the current timestep (`cross_scale_shift_timestep` and `cross_gate_timestep`). Because both audio and video streams reference the same temporal indices, the extracted values (`scale_ca_video_a2v`, `shift_ca_video_a2v`, `gate_out_a2v` for video-to-audio attention, and corresponding audio-side parameters) remain synchronized across modalities.

## Applying Synchronous Modulation

### Affine Transformation with ada_zero_function

The raw hidden states undergo modulation through `ada_zero_function`, which applies an affine transformation (shift plus scale) while preserving the residual connection. This operation ensures that queries, keys, and values entering the cross-attention mechanism are normalized against the same temporal reference point. The function applies the scale and shift tensors extracted via `get_av_ca_ada_values()` to the source modality's hidden states immediately before cross-attention.

### Bidirectional Cross-Attention (A2V and V2A)

LTX-2 implements bidirectional cross-modality attention through `audio_to_video_attn` and `video_to_audio_attn` operations. For audio-to-video (A2V) cross-attention, the video stream receives modulation parameters (`scale_ca_video_a2v`, `shift_ca_video_a2v`) while the audio stream receives corresponding audio-side parameters. Both streams share the identical `cross_scale_shift_timestep` value, guaranteeing that the normalization applied to video keys aligns temporally with the normalization applied to audio queries.

```python

# Extracting AdaLN parameters for A2V cross-attention with shared timestep

scale_ca_video_a2v, shift_ca_video_a2v, gate_out_a2v = self.get_av_ca_ada_values(
    self.scale_shift_table_a2v_ca_video,          # table for video side

    vx.shape[0],                                  # batch size

    video.cross_scale_shift_timestep,             # shared temporal index

    video.cross_gate_timestep,                    # shared gate timestep

    slice(0, 2),                                  # indices for video scale/shift

)

# Apply AdaLN to video hidden state before cross-attention

a2v_vx_scaled = self.ada_zero_function(
    vx_pre_av, self.norm_eps, scale_ca_video_a2v, shift_ca_video_a2v
)

# Cross-attention with temporal gating

vx = vx + (
    self.audio_to_video_attn(
        a2v_vx_scaled,
        context=a2v_ax_scaled,
        pe=video.cross_positional_embeddings,
        k_pe=audio.cross_positional_embeddings,
    )
    * gate_out_a2v
    * video.cross_attn_perturbation_mask
)

```

## Gate-Controlled Blending and Temporal Lockstep

The synchronization mechanism relies on **gate-controlled blending** to determine fusion strength. The `gate_out_a2v` (or `gate_out_v2a`) scalar, derived from the shared `cross_gate_timestep`, multiplies the cross-attention result before it is added to the target modality's residual stream. This gating mechanism allows the model to learn *when* to fuse audio-driven visual updates and *when* to preserve modality independence, while ensuring that the decision occurs at the same temporal step for both streams.

Because the AdaLN parameters are **timestep-dependent**, the identical temporal index used for both audio and video guarantees that modulation applied to the video side aligns precisely with modulation applied to the audio side. Consequently, the cross-attention operations receive **synchronously normalized** queries, keys, and values, enabling temporally coherent blending of audio and video representations throughout the generation pipeline.

## Summary

- **Cross-modality AdaLN** in LTX-2 synchronizes audio and video through shared timestep-dependent scale, shift, and gate parameters.
- The `adaln_embedding_coefficient()` function in [`adaln.py`](https://github.com/Lightricks/LTX-2/blob/main/adaln.py) reserves three additional parameters per block when cross-attention AdaLN is enabled.
- `get_av_ca_ada_values()` extracts synchronized normalization tensors using the shared `cross_scale_shift_timestep` for both modalities.
- `ada_zero_function` applies affine transformations to hidden states before cross-attention, ensuring both streams use the same temporal reference.
- Gate scalars (`gate_out_a2v`, `gate_out_v2a`) control the strength of cross-modal fusion at each timestep, maintaining temporal lockstep between audio and video updates.

## Frequently Asked Questions

### What is cross-modality AdaLN in LTX-2?

Cross-modality AdaLN is an adaptive normalization mechanism that conditions both audio and video transformer streams on shared timestep embeddings. It extracts scale, shift, and gate parameters from the same temporal indices to ensure that cross-attention between modalities occurs with synchronized normalization, preventing temporal misalignment during joint generation.

### How does LTX-2 ensure audio and video remain temporally aligned?

LTX-2 maintains temporal alignment by using identical `cross_scale_shift_timestep` and `cross_gate_timestep` values when extracting AdaLN parameters for both modalities. This ensures that the affine transformations applied to audio and video hidden states reference the same generative timestep, so their cross-attention interactions remain temporally coherent.

### What role does the gate parameter play in cross-modality synchronization?

The gate parameter (extracted as `gate_out_a2v` or `gate_out_v2a`) scales the cross-attention output before it is added to the target modality's residual stream. Because this gate is derived from the shared timestep, it controls *how much* cross-modal information flows between audio and video at each specific generation step, ensuring that fusion strength remains synchronized across both streams.

### Where is cross-modality AdaLN implemented in the codebase?

The core implementation resides in [`packages/ltx-core/src/ltx_core/model/transformer/adaln.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/adaln.py) (parameter computation and `AdaLayerNormSingle`), with the extraction logic in [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py) (specifically `get_av_ca_ada_values()` and the cross-attention forward pass). The affine transformation utility `ada_zero_function` is typically located in the utilities module or within the transformer implementation file.