# How LTX-2's Asymmetric Dual-Stream Transformer Architecture Processes Video and Audio Jointly

> Discover how LTX-2's asymmetric dual-stream transformer architecture processes video and audio together. Explore its unique approach to cross-modal interaction for enhanced AI functionality.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-06-21

---

**LTX-2 uses an asymmetric dual-stream transformer architecture that processes video and audio through separate transformer blocks with independent hidden dimensions and attention heads, while enabling controlled cross-modal interaction via directional cross-attention layers.**

The LTX-2 model from Lightricks introduces a novel approach to multimodal generation through its asymmetric dual-stream transformer architecture. This design allows the model to handle video and audio modalities with specialized parameters while maintaining efficient information flow between them. Unlike symmetric architectures that force identical configurations across modalities, LTX-2's implementation treats each stream independently while preserving the ability to fuse representations through gated cross-attention mechanisms.

## Core Components and File Structure

### BasicAVTransformerBlock Implementation

Located in [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py) at line 86, the `BasicAVTransformerBlock` class serves as the fundamental building block. It hosts separate video and audio sub-blocks within a single transformer block, enabling parallel processing of both modalities while allowing controlled interaction between them.

### TransformerConfig and Asymmetric Parameters

The `TransformerConfig` class (defined at line 30 of transformer.py) holds modality-specific hyperparameters including dimension sizes, head counts, and context dimensions. This configuration enables the architectural asymmetry where video might use 32 heads of 128 dimensions each (4096 inner dimension) while audio uses 32 heads of 64 dimensions (2048 inner dimension). Cross-attention dimensions and gating timesteps remain independent, providing fine-grained control over modality interaction.

### LTXModel Assembly

The top-level `LTXModel` class in [`packages/ltx-core/src/ltx_core/model/transformer/model.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/model.py) (line 41) constructs a stack of 48 `BasicAVTransformerBlock` instances by default. During initialization, it creates separate patchify projections (`self.patchify_proj` for video, `self.audio_patchify_proj` for audio) and instantiates per-modality AdaLN layers with distinct scale-shift tables.

## Dual-Stream Data Flow and Mechanisms

### Self-Attention Paths

Each stream processes independently through dedicated attention mechanisms. Video flows through `self.attn1` (self-attention) and `self.attn2` (text-cross-attention), while audio uses `self.audio_attn1` and `self.audio_attn2`. Both paths apply `AdaLayerNormSingle` modulation using per-modality scale-shift tables (e.g., `self.scale_shift_table`) before attention operations.

### Directional Cross-Attention

Cross-attention operates bidirectionally but asymmetrically within each block. The `self.audio_to_video_attn` module queries video tokens using audio keys and values, while `self.video_to_audio_attn` queries audio tokens using video keys and values. Each direction maintains separate scale-shift tables (`scale_shift_table_a2v_ca_video`, `scale_shift_table_a2v_ca_audio`) and gating parameters to control the influence each modality has on the other.

### Feed-Forward and Output Projection

Following attention, each stream passes through independent feed-forward networks (`self.ff` for video, `self.audio_ff` for audio). Final outputs are projected via `self.proj_out` (video) and `self.audio_proj_out` (audio), preserving the dimensional asymmetry throughout the entire forward pass.

## Modality Modes and Runtime Configuration

The `TransformerArgs` dataclass in [`packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py) (line 16) carries runtime tensors, masks, and timesteps through the blocks. The `LTXModel` supports three operational modes via the `LTXModelType` enum:

- **Audio-Video**: Both streams active with full cross-attention enabled
- **Video-Only**: Audio config omitted; blocks behave as standard video transformers
- **Audio-Only**: Video config omitted; blocks behave as standard audio transformers

This flexibility allows the same architecture to handle unimodal generation without requiring separate model implementations.

## Implementation Example

The following example demonstrates creating a dual-stream LTX-2 model with asymmetric dimensions and processing dummy video and audio tensors:

```python
import torch
from ltx_core.model.transformer.model import LTXModel, LTXModelType

# Dummy video tensor: (batch, frames, height, width, channels) → flattened to (B, N, C)

video_tensor = torch.randn(2, 16, 64, 64, 128)          # B=2, T=16, C=128

video_tensor = video_tensor.view(2, 16, -1)             # (B, T, C)

# Dummy audio tensor: (batch, time, channels)

audio_tensor = torch.randn(2, 200, 128)                # B=2, T=200, C=128

# Build a dual‑stream LTX model (default settings)

model = LTXModel(
    model_type=LTXModelType.AudioVideo,
    num_attention_heads=32,
    attention_head_dim=128,
    audio_num_attention_heads=32,
    audio_attention_head_dim=64,
    in_channels=128,
    audio_in_channels=128,
)

# Forward pass – returns (video_output, audio_output)

video_out, audio_out = model(video=video_tensor, audio=audio_tensor)

print("Video latent shape:", video_out.shape)   # e.g. (2, 16, 128)

print("Audio latent shape:", audio_out.shape)   # e.g. (2, 200, 128)

```

The example demonstrates the symmetric API (`model(video=…, audio=…)`) while the internal block layout remains asymmetric, with video and audio streams maintaining their distinct dimensional configurations throughout processing.

## Summary

- **LTX-2's asymmetric dual-stream transformer architecture** processes video and audio through separate transformer blocks with independent configurations, allowing each modality to maintain optimal hyperparameters for its specific characteristics.
- **Cross-attention mechanisms** operate bidirectionally with separate scale-shift tables and gating parameters, enabling controlled information flow between modalities without forcing dimensional alignment.
- **Three operational modes** (Audio-Video, Video-Only, Audio-Only) allow the same architecture to handle unimodal or multimodal inputs by selectively enabling stream-specific components.
- **Modality-specific AdaLN layers** apply timestep-dependent scaling and shifting through separate parameters, ensuring proper diffusion model conditioning for each stream independently.

## Frequently Asked Questions

### What makes LTX-2's dual-stream architecture "asymmetric"?

The architecture is asymmetric because the video and audio streams can have different hidden dimensions, attention head counts, and positional embedding configurations. For example, video might use 4096 dimensions (32 heads × 128 dimensions) while audio uses 2048 dimensions (32 heads × 64 dimensions), with independent cross-attention dimensions and gating parameters for each direction.

### How does cross-attention work between the video and audio streams?

Cross-attention operates through dedicated modules within `BasicAVTransformerBlock`: `audio_to_video_attn` allows audio tokens to query video representations, while `video_to_audio_attn` allows video tokens to query audio representations. Each direction uses separate scale-shift tables (`scale_shift_table_a2v_ca_video`, `scale_shift_table_a2v_ca_audio`) to control the gating and influence of the cross-modal interaction.

### Can LTX-2 operate with only video or only audio input?

Yes, the `LTXModel` supports three modes via the `LTXModelType` enum. When initialized as `Video-Only`, the audio configuration is omitted and blocks behave as standard video transformers. In `Audio-Only` mode, the video configuration is omitted. This flexibility allows the same architecture to handle unimodal generation without requiring separate model implementations.

### Where are the AdaLN parameters defined for each modality?

AdaLayerNorm parameters are defined in [`packages/ltx-core/src/ltx_core/model/transformer/adaln.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/adaln.py) and instantiated separately for each stream within `LTXModel`. Each modality maintains its own scale-shift tables (e.g., `self.scale_shift_table` for video, `audio_scale_shift_table` for audio) and additional tables for cross-attention gates, enabling timestep-dependent modulation that respects the asymmetric dimensional settings.