How to Configure Cross-Modal Attention Scaling for Modality Guidance in LTX-2

Set the av_ca_timestep_scale_multiplier parameter in your LTX-2 model configuration YAML to control how strongly secondary modalities (e.g., audio) influence the primary generation (e.g., video) through scaled cross-attention gates.

LTX-2 is an open-source multimodal generative model developed by Lightricks that enables cross-modal generation tasks such as audio-guided video inpainting. To fine-tune the intensity of modality guidance during inference or training, you must configure the cross-modal attention scaling mechanism via the av_ca_timestep_scale_multiplier argument in the transformer arguments preprocessor.

Understanding Cross-Modal Attention Scaling

The cross-modal attention scaling factor determines the strength of influence that a secondary modality exerts on the primary modality during the diffusion process. This scaling is applied to the gate noise timestep that modulates the AdaLayerNorm gates used for cross-attention.

The Role of av_ca_timestep_scale_multiplier

In packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py, the MultiModalTransformerArgsPreprocessor class implements the scaling logic. The parameter av_ca_timestep_scale_multiplier (audio-video cross-attention timestep scale multiplier) acts as a multiplier on the primary modality's timestep scale, specifically adjusting the timestep fed into the cross_gate_adaln module.

A larger value increases the cross-modality influence, while a smaller value weakens it. Setting this to 0 effectively nullifies the cross-attention contribution.

Implementation in the Source Code

The scaling calculation occurs in the _prepare_cross_attention_timestep method of MultiModalTransformerArgsPreprocessor. The code calculates a scaling factor by dividing av_ca_timestep_scale_multiplier by the primary timestep_scale_multiplier, then applies this to the cross-modality sigma:


# In MultiModalTransformerArgsPreprocessor._prepare_cross_attention_timestep

av_ca_factor = self.av_ca_timestep_scale_multiplier / timestep_scale_multiplier
gate_noise_timestep, _ = self.cross_gate_adaln(
    (cross_modality_sigma * timestep_scale_multiplier * av_ca_factor).flatten(),
    hidden_dtype=hidden_dtype,
)

This results in the gate_noise_timestep that controls the cross-attention gates, directly modulating how much the secondary modality affects the hidden states.

Configuring the Scaling Factor

To adjust cross-modal attention scaling in your LTX-2 workflow, modify the model configuration YAML file (e.g., video_inpainting_lora.yaml or audio_inpainting_lora.yaml).

  1. Locate your configuration file in the repository's config directory.

  2. Add or modify the av_ca_timestep_scale_multiplier key alongside the primary timestep_scale_multiplier:


# Example configuration snippet

timestep_scale_multiplier: 1000          # Primary modality scale

av_ca_timestep_scale_multiplier: 500     # Cross-modal (audio-video) scale
  1. Reload the model using the trainer scripts (such as train.py or process_dataset.py), which automatically read the YAML configuration. The MultiModalTransformerArgsPreprocessor will instantiate with the new scaling factor on the next run.

Disabling Cross-Modal Guidance

To completely remove cross-modal influence, use one of two methods:

  • Set the multiplier to zero: Configure av_ca_timestep_scale_multiplier: 0 in your YAML. This produces a zero-scaled timestep for the gate, nullifying the cross-attention contribution.

  • Provide a zero-sigma tensor: Supply a cross-modality tensor with sigma equal to 0 in the Modality definition. The preprocessing logic checks cross_modality.sigma.numel() > 1, and when the condition fails or sigma is zero, the cross-attention tensors are omitted from the computation.

Summary

  • Cross-modal attention scaling in LTX-2 is controlled by the av_ca_timestep_scale_multiplier parameter in your YAML configuration.
  • The scaling logic resides in MultiModalTransformerArgsPreprocessor within packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py.
  • The factor scales the gate noise timestep for cross_gate_adaln, directly controlling cross-attention strength.
  • Set the value to 0 to disable cross-modal guidance entirely.
  • Changes take effect immediately upon reloading the model configuration in trainer scripts.

Frequently Asked Questions

What does av_ca_timestep_scale_multiplier stand for?

The parameter name stands for "audio-video cross-attention timestep scale multiplier." It specifically controls the scaling factor applied to the timestep when preparing cross-modal attention between audio and video modalities, though the mechanism applies to any cross-modal pair in the LTX-2 architecture.

How do I completely disable audio influence on video generation?

Set av_ca_timestep_scale_multiplier to 0 in your configuration YAML, or provide a cross-modality Modality tensor with sigma equal to 0. Both methods cause the MultiModalTransformerArgsPreprocessor to either zero-out the gate contributions or skip the cross-attention tensors entirely.

Where is the cross-modal attention scaling applied in the architecture?

The scaling is applied in the MultiModalTransformerArgsPreprocessor._prepare_cross_attention_timestep method located in packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py. It modifies the timestep passed to the cross_gate_adaln AdaLayerNorm gates that control the cross-attention layers in the transformer.

What values should I use for av_ca_timestep_scale_multiplier?

Typical values range from 0 (no influence) to 1000 (full influence), often configured relative to your timestep_scale_multiplier. For balanced guidance, set it equal to or lower than the primary modality scale (e.g., 500 when the primary is 1000). Adjust based on your specific inpainting or generation task requirements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →