# How to Configure Cross-Modal Attention Scaling for Modality Guidance in LTX-2

> Configure LTX-2 cross-modal attention scaling for modality guidance. Use av_ca_timestep_scale_multiplier to control secondary modality influence on primary generation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-06-20

---

**Set the `av_ca_timestep_scale_multiplier` parameter in your LTX-2 model configuration YAML to control how strongly secondary modalities (e.g., audio) influence the primary generation (e.g., video) through scaled cross-attention gates.**

LTX-2 is an open-source multimodal generative model developed by Lightricks that enables cross-modal generation tasks such as audio-guided video inpainting. To fine-tune the intensity of **modality guidance** during inference or training, you must configure the **cross-modal attention scaling** mechanism via the `av_ca_timestep_scale_multiplier` argument in the transformer arguments preprocessor.

## Understanding Cross-Modal Attention Scaling

The cross-modal attention scaling factor determines the strength of influence that a secondary modality exerts on the primary modality during the diffusion process. This scaling is applied to the gate noise timestep that modulates the AdaLayerNorm gates used for cross-attention.

### The Role of av_ca_timestep_scale_multiplier

In [`packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py), the `MultiModalTransformerArgsPreprocessor` class implements the scaling logic. The parameter `av_ca_timestep_scale_multiplier` (audio-video cross-attention timestep scale multiplier) acts as a multiplier on the primary modality's timestep scale, specifically adjusting the timestep fed into the `cross_gate_adaln` module.

A larger value increases the cross-modality influence, while a smaller value weakens it. Setting this to `0` effectively nullifies the cross-attention contribution.

## Implementation in the Source Code

The scaling calculation occurs in the `_prepare_cross_attention_timestep` method of `MultiModalTransformerArgsPreprocessor`. The code calculates a scaling factor by dividing `av_ca_timestep_scale_multiplier` by the primary `timestep_scale_multiplier`, then applies this to the cross-modality sigma:

```python

# In MultiModalTransformerArgsPreprocessor._prepare_cross_attention_timestep

av_ca_factor = self.av_ca_timestep_scale_multiplier / timestep_scale_multiplier
gate_noise_timestep, _ = self.cross_gate_adaln(
    (cross_modality_sigma * timestep_scale_multiplier * av_ca_factor).flatten(),
    hidden_dtype=hidden_dtype,
)

```

This results in the `gate_noise_timestep` that controls the cross-attention gates, directly modulating how much the secondary modality affects the hidden states.

## Configuring the Scaling Factor

To adjust **cross-modal attention scaling** in your LTX-2 workflow, modify the model configuration YAML file (e.g., [`video_inpainting_lora.yaml`](https://github.com/Lightricks/LTX-2/blob/main/video_inpainting_lora.yaml) or [`audio_inpainting_lora.yaml`](https://github.com/Lightricks/LTX-2/blob/main/audio_inpainting_lora.yaml)).

1. **Locate your configuration file** in the repository's config directory.

2. **Add or modify** the `av_ca_timestep_scale_multiplier` key alongside the primary `timestep_scale_multiplier`:

```yaml

# Example configuration snippet

timestep_scale_multiplier: 1000          # Primary modality scale

av_ca_timestep_scale_multiplier: 500     # Cross-modal (audio-video) scale

```

3. **Reload the model** using the trainer scripts (such as [`train.py`](https://github.com/Lightricks/LTX-2/blob/main/train.py) or [`process_dataset.py`](https://github.com/Lightricks/LTX-2/blob/main/process_dataset.py)), which automatically read the YAML configuration. The `MultiModalTransformerArgsPreprocessor` will instantiate with the new scaling factor on the next run.

## Disabling Cross-Modal Guidance

To completely remove cross-modal influence, use one of two methods:

- **Set the multiplier to zero:** Configure `av_ca_timestep_scale_multiplier: 0` in your YAML. This produces a zero-scaled timestep for the gate, nullifying the cross-attention contribution.

- **Provide a zero-sigma tensor:** Supply a cross-modality tensor with `sigma` equal to `0` in the `Modality` definition. The preprocessing logic checks `cross_modality.sigma.numel() > 1`, and when the condition fails or sigma is zero, the cross-attention tensors are omitted from the computation.

## Summary

- **Cross-modal attention scaling** in LTX-2 is controlled by the `av_ca_timestep_scale_multiplier` parameter in your YAML configuration.
- The scaling logic resides in `MultiModalTransformerArgsPreprocessor` within [`packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py).
- The factor scales the gate noise timestep for `cross_gate_adaln`, directly controlling cross-attention strength.
- Set the value to `0` to disable cross-modal guidance entirely.
- Changes take effect immediately upon reloading the model configuration in trainer scripts.

## Frequently Asked Questions

### What does av_ca_timestep_scale_multiplier stand for?

The parameter name stands for "audio-video cross-attention timestep scale multiplier." It specifically controls the scaling factor applied to the timestep when preparing cross-modal attention between audio and video modalities, though the mechanism applies to any cross-modal pair in the LTX-2 architecture.

### How do I completely disable audio influence on video generation?

Set `av_ca_timestep_scale_multiplier` to `0` in your configuration YAML, or provide a cross-modality `Modality` tensor with `sigma` equal to `0`. Both methods cause the `MultiModalTransformerArgsPreprocessor` to either zero-out the gate contributions or skip the cross-attention tensors entirely.

### Where is the cross-modal attention scaling applied in the architecture?

The scaling is applied in the `MultiModalTransformerArgsPreprocessor._prepare_cross_attention_timestep` method located in [`packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py). It modifies the timestep passed to the `cross_gate_adaln` AdaLayerNorm gates that control the cross-attention layers in the transformer.

### What values should I use for av_ca_timestep_scale_multiplier?

Typical values range from `0` (no influence) to `1000` (full influence), often configured relative to your `timestep_scale_multiplier`. For balanced guidance, set it equal to or lower than the primary modality scale (e.g., `500` when the primary is `1000`). Adjust based on your specific inpainting or generation task requirements.