How to Configure Cross-Modal Attention Scaling for Modality Guidance in LTX-2
Set the av_ca_timestep_scale_multiplier parameter in your LTX-2 model configuration YAML to control how strongly secondary modalities (e.g., audio) influence the primary generation (e.g., video) through scaled cross-attention gates.
LTX-2 is an open-source multimodal generative model developed by Lightricks that enables cross-modal generation tasks such as audio-guided video inpainting. To fine-tune the intensity of modality guidance during inference or training, you must configure the cross-modal attention scaling mechanism via the av_ca_timestep_scale_multiplier argument in the transformer arguments preprocessor.
Understanding Cross-Modal Attention Scaling
The cross-modal attention scaling factor determines the strength of influence that a secondary modality exerts on the primary modality during the diffusion process. This scaling is applied to the gate noise timestep that modulates the AdaLayerNorm gates used for cross-attention.
The Role of av_ca_timestep_scale_multiplier
In packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py, the MultiModalTransformerArgsPreprocessor class implements the scaling logic. The parameter av_ca_timestep_scale_multiplier (audio-video cross-attention timestep scale multiplier) acts as a multiplier on the primary modality's timestep scale, specifically adjusting the timestep fed into the cross_gate_adaln module.
A larger value increases the cross-modality influence, while a smaller value weakens it. Setting this to 0 effectively nullifies the cross-attention contribution.
Implementation in the Source Code
The scaling calculation occurs in the _prepare_cross_attention_timestep method of MultiModalTransformerArgsPreprocessor. The code calculates a scaling factor by dividing av_ca_timestep_scale_multiplier by the primary timestep_scale_multiplier, then applies this to the cross-modality sigma:
# In MultiModalTransformerArgsPreprocessor._prepare_cross_attention_timestep
av_ca_factor = self.av_ca_timestep_scale_multiplier / timestep_scale_multiplier
gate_noise_timestep, _ = self.cross_gate_adaln(
(cross_modality_sigma * timestep_scale_multiplier * av_ca_factor).flatten(),
hidden_dtype=hidden_dtype,
)
This results in the gate_noise_timestep that controls the cross-attention gates, directly modulating how much the secondary modality affects the hidden states.
Configuring the Scaling Factor
To adjust cross-modal attention scaling in your LTX-2 workflow, modify the model configuration YAML file (e.g., video_inpainting_lora.yaml or audio_inpainting_lora.yaml).
-
Locate your configuration file in the repository's config directory.
-
Add or modify the
av_ca_timestep_scale_multiplierkey alongside the primarytimestep_scale_multiplier:
# Example configuration snippet
timestep_scale_multiplier: 1000 # Primary modality scale
av_ca_timestep_scale_multiplier: 500 # Cross-modal (audio-video) scale
- Reload the model using the trainer scripts (such as
train.pyorprocess_dataset.py), which automatically read the YAML configuration. TheMultiModalTransformerArgsPreprocessorwill instantiate with the new scaling factor on the next run.
Disabling Cross-Modal Guidance
To completely remove cross-modal influence, use one of two methods:
-
Set the multiplier to zero: Configure
av_ca_timestep_scale_multiplier: 0in your YAML. This produces a zero-scaled timestep for the gate, nullifying the cross-attention contribution. -
Provide a zero-sigma tensor: Supply a cross-modality tensor with
sigmaequal to0in theModalitydefinition. The preprocessing logic checkscross_modality.sigma.numel() > 1, and when the condition fails or sigma is zero, the cross-attention tensors are omitted from the computation.
Summary
- Cross-modal attention scaling in LTX-2 is controlled by the
av_ca_timestep_scale_multiplierparameter in your YAML configuration. - The scaling logic resides in
MultiModalTransformerArgsPreprocessorwithinpackages/ltx-core/src/ltx_core/model/transformer/transformer_args.py. - The factor scales the gate noise timestep for
cross_gate_adaln, directly controlling cross-attention strength. - Set the value to
0to disable cross-modal guidance entirely. - Changes take effect immediately upon reloading the model configuration in trainer scripts.
Frequently Asked Questions
What does av_ca_timestep_scale_multiplier stand for?
The parameter name stands for "audio-video cross-attention timestep scale multiplier." It specifically controls the scaling factor applied to the timestep when preparing cross-modal attention between audio and video modalities, though the mechanism applies to any cross-modal pair in the LTX-2 architecture.
How do I completely disable audio influence on video generation?
Set av_ca_timestep_scale_multiplier to 0 in your configuration YAML, or provide a cross-modality Modality tensor with sigma equal to 0. Both methods cause the MultiModalTransformerArgsPreprocessor to either zero-out the gate contributions or skip the cross-attention tensors entirely.
Where is the cross-modal attention scaling applied in the architecture?
The scaling is applied in the MultiModalTransformerArgsPreprocessor._prepare_cross_attention_timestep method located in packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py. It modifies the timestep passed to the cross_gate_adaln AdaLayerNorm gates that control the cross-attention layers in the transformer.
What values should I use for av_ca_timestep_scale_multiplier?
Typical values range from 0 (no influence) to 1000 (full influence), often configured relative to your timestep_scale_multiplier. For balanced guidance, set it equal to or lower than the primary modality scale (e.g., 500 when the primary is 1000). Adjust based on your specific inpainting or generation task requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →