# How LTX-2 Distributes Model Parameters Between Video and Audio Streams

> Discover how LTX-2 distributes model parameters for video and audio streams with dedicated layers and bidirectional cross-attention. Learn about its unique approach to multimodal AI.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: architecture
- Published: 2026-06-21

---

**LTX-2 allocates separate transformer parameters for video and audio streams, using distinct self-attention, feed-forward, and AdaLN modulation layers for each modality, plus bidirectional cross-attention modules that enable interaction without sharing weights.**

The LTX-2 architecture from Lightricks implements a dual-branch transformer design that processes video and audio through parallel yet independent parameter sets. Unlike unified multimodal models that share weights across modalities, LTX-2 maintains modality-specific layers while enabling cross-modal attention through dedicated projection matrices. This article examines exactly how parameters are distributed between the video and audio streams based on the official source code in [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py).

## Parallel Transformer Branches with Separate Parameter Sets

LTX-2 instantiates two parallel transformer branches that share architectural logic but maintain completely separate learnable parameters. Each modality has its own self-attention mechanisms and feed-forward networks, ensuring that video and audio representations develop independently until explicit cross-modal fusion occurs.

### Self-Attention Layers

For the video stream, the transformer defines `self.attn1` and `self.attn2` as the primary self-attention modules. The audio stream receives dedicated counterparts named `self.audio_attn1` and `self.audio_attn2`, initialized at lines 151-166 of the transformer implementation.

Because these modules instantiate separate `Attention` classes, each maintains independent query, key, value, and output projection matrices. The video and audio branches do not share attention weights, allowing each modality to learn modality-specific attention patterns.

### Feed-Forward Networks

Each branch contains its own feed-forward network. The video path uses `self.ff` (defined in the shared code path), while the audio path uses `self.audio_ff` (lines 147-148). These are distinct `FeedForward` instances with separate weight matrices and bias terms, doubling the FFN parameter count compared to a single shared network.

## Cross-Modal Attention Mechanisms

Cross-modal interaction occurs through two unidirectional attention modules, each contributing a full set of projection parameters.

### Audio-to-Video Attention

The `self.audio_to_video_attn` module (lines 152-162) enables the video stream to attend to audio features. This module creates query projections from video tokens and key/value projections from audio tokens, adding a complete set of Q, K, V, and output linear layers distinct from the self-attention parameters.

### Video-to-Audio Attention

Conversely, `self.video_to_audio_attn` (lines 164-174) allows audio queries to attend to video keys and values. This symmetrical design means cross-modal parameters effectively double compared to a single-direction implementation, as each direction requires independent projection matrices.

## AdaLN Modulation Tables

Adaptive Layer Normalization (AdaLN) parameters follow the same separation pattern. The video stream uses `self.scale_shift_table`, while audio uses `self.audio_scale_shift_table`.

For cross-modal attention, dedicated modulation tables condition the layer outputs: `self.scale_shift_table_a2v_ca_video` and `self.scale_shift_table_a2v_ca_audio` (lines 176-177). These tables provide learnable scale and shift parameters that modulate the cross-attention outputs independently for each direction.

## Configuration-Driven Dimensionality

Parameter counts scale according to configuration values defined in training YAML files such as [`video_suffix_lora.yaml`](https://github.com/Lightricks/LTX-2/blob/main/video_suffix_lora.yaml) and [`audio_suffix_lora.yaml`](https://github.com/Lightricks/LTX-2/blob/main/audio_suffix_lora.yaml). The `TransformerArgs` dataclass in [`packages/ltx-core/src/ltx_core/types.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/types.py) holds these values:

```yaml
video:
  dim: 768                # hidden size for video tokens

  heads: 12               # number of attention heads

  d_head: 64              # per-head dimension

audio:
  dim: 512                # hidden size for audio tokens

  heads: 8
  d_head: 64

```

These dimensions determine the sizes of linear layers inside `Attention` and `FeedForward` modules. A video token with `dim=768` and 12 heads produces larger projection matrices than an audio token with `dim=512` and 8 heads, creating an asymmetric parameter distribution that favors the video stream when using these default values.

## Dynamic Parameter Budget Adjustment

The architecture supports dynamic parameter budgets through conditional initialization. Lines 58-60 in the `forward` method check whether audio is enabled, allowing the model to run as video-only when `audio=None`.

```python

# Example: constructing a transformer with both video and audio streams

from ltx_core.model.transformer.transformer import Transformer
from ltx_core.types import TransformerArgs

video_cfg = TransformerArgs(
    dim=768,
    heads=12,
    d_head=64,
    # … other video-specific settings …

)

audio_cfg = TransformerArgs(
    dim=512,
    heads=8,
    d_head=64,
    # … other audio-specific settings …

)

model = Transformer(video=video_cfg, audio=audio_cfg)

# Forward pass – the model will route video tokens through the video branch,

# audio tokens through the audio branch, and then perform cross-attention.

video_out, audio_out = model(video=video_input, audio=audio_input)

```

When training video-only models, setting `audio=None` eliminates all audio-specific parameters (`audio_attn1`, `audio_attn2`, `audio_ff`, `audio_scale_shift_table`, and both cross-attention modules), reducing the total parameter count proportionally.

```python

# Example: disabling the audio branch (video-only model)

model = Transformer(video=video_cfg, audio=None)

video_out, _ = model(video=video_input, audio=None)   # only video path runs

```

### Inspecting Parameter Distribution

You can verify the parameter allocation programmatically:

```python

# Example: inspecting parameter counts per branch

total_params = sum(p.numel() for p in model.parameters())
video_params = sum(p.numel() for n, p in model.named_parameters() if "audio" not in n)
audio_params = sum(p.numel() for n, p in model.named_parameters() if "video" not in n and "audio" in n)

print(f"Total: {total_params:,} | Video: {video_params:,} | Audio: {audio_params:,}")

```

## Summary

- **Separate linear projections** – Video and audio maintain independent Q, K, V, and output projection matrices with no shared weights.
- **Independent feed-forward networks** – Each modality has its own `FeedForward` instance (`self.ff` vs `self.audio_ff`).
- **Bidirectional cross-attention** – Two distinct modules (`audio_to_video_attn` and `video_to_audio_attn`) each contribute a full attention block's worth of parameters.
- **Modality-specific AdaLN tables** – Scale-shift parameters are separated into video, audio, and cross-modal-specific tables.
- **Configurable dimensions** – Parameter counts scale according to `dim`, `heads`, and `d_head` values in configuration files.
- **Runtime flexibility** – The audio branch can be disabled entirely by passing `audio=None`, dynamically reducing the active parameter count.

## Frequently Asked Questions

### Do video and audio streams share any parameters in LTX-2?

No, the video and audio streams maintain completely separate parameter sets according to the source code in [`transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer.py). Each modality has its own self-attention projections (`attn1`/`attn2` vs `audio_attn1`/`audio_attn2`), feed-forward networks (`ff` vs `audio_ff`), and AdaLN modulation tables. The only interaction occurs through cross-attention modules, which still use distinct projection matrices for each direction rather than sharing weights.

### How does cross-attention affect the total parameter count?

Cross-attention adds a full self-attention block for each direction. Because LTX-2 implements bidirectional cross-attention (`audio_to_video_attn` and `video_to_audio_attn`), it doubles the cross-modal parameter count compared to a single-direction design. Each cross-attention module includes independent query, key, value, and output projections, plus dedicated AdaLN modulation tables (`scale_shift_table_a2v_ca_video` and `scale_shift_table_a2v_ca_audio`).

### Can I train a video-only model using the LTX-2 architecture?

Yes, the architecture supports video-only training by setting `audio=None` when instantiating the `Transformer` class. The forward pass at lines 58-60 includes conditional checks that skip audio processing when the audio configuration is missing. This eliminates all audio-specific parameters (self-attention, feed-forward, and cross-attention modules) from the active computation graph and parameter count.

### Where are the dimensionality settings configured for each stream?

Dimensionality settings are defined in YAML configuration files such as [`video_suffix_lora.yaml`](https://github.com/Lightricks/LTX-2/blob/main/video_suffix_lora.yaml) and [`audio_suffix_lora.yaml`](https://github.com/Lightricks/LTX-2/blob/main/audio_suffix_lora.yaml) within `packages/ltx-trainer/configs/`. These values populate the `TransformerArgs` dataclass from [`packages/ltx-core/src/ltx_core/types.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/types.py), which then drives the construction of linear layers in the transformer. The `dim`, `heads`, and `d_head` parameters determine the width and depth of projection matrices for each modality independently.