# How LTX-2's Dual-Stream Transformer Handles Audio-Video Synchronization

> Discover how LTX-2's dual-stream transformer achieves audio-video synchronization using parallel processing and cross-attention. Learn about frame-level alignment and AdaLN modulation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: internals
- Published: 2026-06-20

---

**LTX-2 synchronizes audio and video through a single transformer block containing two parallel streams that process each modality independently before fusing them via temporally-gated cross-attention, ensuring frame-level alignment through per-timestep AdaLN modulation.**

The LTX-2 video generation model by Lightricks employs a sophisticated dual-stream transformer architecture to maintain tight synchronization between audio and visual content. This design processes video and audio in parallel streams within a single transformer block, using bidirectional cross-attention mechanisms and adaptive layer normalization to align temporal features at the architecture level. Understanding this implementation reveals how the open-source model achieves precise audio-video coherence without sacrificing modality-specific representations.

## Architecture Overview of the Dual-Stream Transformer

At the core of LTX-2's synchronization capability is the `BasicAVTransformerBlock` class defined in [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py). This block maintains separate processing paths for video and audio while enabling controlled information exchange between modalities.

### Separate Modality Configurations

The block receives distinct `TransformerConfig` objects for video and audio (lines 86-94), defining dimension sizes, attention heads, and Ada-LN cross-attention settings. This configuration allows the architecture to handle unimodal inputs gracefully—if a modality is absent, its config is set to `None`, and the block processes only the available stream.

### Independent Self-Attention Processing

Each modality first undergoes dedicated self-attention to preserve temporal ordering. According to lines 70-88 in [`transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer.py), video features pass through `self.attn1` and `self.attn2`, while audio features use `self.audio_attn1` and `self.audio_attn2`. This separation ensures that frame sequences and audio samples retain their internal structure before any cross-modal mixing occurs.

## Cross-Modal Synchronization Mechanisms

Following independent processing, the dual-stream transformer aligns modalities through specialized cross-attention layers that operate bidirectionally.

### Bidirectional Cross-Attention

The block implements two distinct cross-attention mechanisms (lines 51-74):

- **Audio-to-Video** (`self.audio_to_video_attn`): Video queries attend to audio keys and values, allowing visual features to incorporate relevant sound context.
- **Video-to-Audio** (`self.video_to_audio_attn`): Audio queries attend to video keys and values, enabling sound generation to synchronize with visual events.

These layers are instantiated conditionally only when both modalities are present, using the same `Attention` implementation but swapping modality dimensions to ensure direct temporal correspondence.

### Temporally-Gated AdaLN Modulation

Before attention operations, hidden states undergo `AdaZero` normalization and modulation via learned scale-shift tables (`self.scale_shift_table` and `self.audio_scale_shift_table`). As implemented in lines 115-138, the method `self.get_av_ca_ada_values` extracts per-timestep gate values (`gate_out_a2v` and `gate_out_v2a`) from dedicated tables (`self.scale_shift_table_a2v_ca_*`).

This **time-aligned gating** mechanism allows the model to amplify or attenuate cross-attention strength based on the exact alignment between audio and video timesteps, creating precise synchronization points.

## Residual Integration and Output Processing

After cross-attention, outputs merge back into their respective streams through residual connections (lines 140-166). The cross-attention results are multiplied by gate tensors and optional perturbation masks before addition, ensuring that synchronized information integrates smoothly without overwriting modality-specific features.

Finally, each stream passes through dedicated feed-forward networks (`self.ff` for video, `self.audio_ff` for audio) with additional Ada-LN scaling (lines 168-190). This final processing step ensures that residual cross-modal information is properly normalized before the block returns updated `TransformerArgs` objects for both modalities.

## Implementation Example

The following code demonstrates how to instantiate and use the dual-stream transformer block for audio-video synchronization:

```python

# Example: Building a dual‑stream transformer block for a video‑to‑audio task

from ltx_core.model.transformer.transformer import (
    BasicAVTransformerBlock,
    TransformerConfig,
    TransformerOpsConfig,
)

# Video config: 768‑dim, 12 heads, cross‑attention enabled

video_cfg = TransformerConfig(
    dim=768,
    heads=12,
    d_head=64,
    context_dim=768,            # video‑to‑audio cross‑attention dim

    apply_gated_attention=False,
    cross_attention_adaln=True,
)

# Audio config: 512‑dim, 8 heads, cross‑attention enabled

audio_cfg = TransformerConfig(
    dim=512,
    heads=8,
    d_head=64,
    context_dim=512,
    apply_gated_attention=False,
    cross_attention_adaln=True,
)

# Optional custom ops (default PyTorch ops are used if omitted)

ops = TransformerOpsConfig()

block = BasicAVTransformerBlock(
    video=video_cfg,
    audio=audio_cfg,
    rope_type="SPLIT",   # LTXRopeType.SPLIT

    norm_eps=1e-6,
    ops=ops,
)

# Dummy inputs (B=2, T=16 for video; B=2, T=64 for audio)

import torch
video_tensor = torch.randn(2, 16, 768)   # [batch, frames, dim]

audio_tensor = torch.randn(2, 64, 512)   # [batch, samples, dim]

# Wrap tensors in TransformerArgs (simplified version)

from ltx_core.model.transformer.transformer_args import TransformerArgs

video_args = TransformerArgs(
    x=video_tensor,
    enabled=True,
    positional_embeddings=None,
    self_attention_mask=None,
    cross_positional_embeddings=None,
    cross_attn_skip_all=False,
    cross_scale_shift_timestep=torch.zeros(2, 1, 1),  # placeholder

    cross_gate_timestep=torch.zeros(2, 1, 1),
    cross_attn_perturbation_mask=1.0,
)

audio_args = TransformerArgs(
    x=audio_tensor,
    enabled=True,
    positional_embeddings=None,
    self_attention_mask=None,
    cross_positional_embeddings=None,
    cross_attn_skip_all=False,
    cross_scale_shift_timestep=torch.zeros(2, 1, 1),
    cross_gate_timestep=torch.zeros(2, 1, 1),
    cross_attn_perturbation_mask=1.0,
)

# Forward pass – returns synchronized video and audio tensors

synced_video, synced_audio = block(video_args, audio_args)
print(synced_video.x.shape, synced_audio.x.shape)

```

The key source files implementing this architecture include:
- [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py) – defines `BasicAVTransformerBlock` and the core dual-stream logic
- [`packages/ltx-core/src/ltx_core/model/transformer/model.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/model.py) – constructs the full model using `BasicAVTransformerBlock` instances
- [`packages/ltx-core/src/ltx_core/model/transformer/attention.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/attention.py) – implements the underlying `Attention` class
- [`packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer_args.py) – containers for per-layer inputs

## Summary

- **Dual-stream architecture** maintains separate processing paths for video and audio during self-attention, preserving modality-specific temporal structures.
- **Bidirectional cross-attention** enables direct information flow between streams, with video attending to audio and vice versa.
- **Temporally-gated AdaLN** provides per-timestep modulation through `get_av_ca_ada_values`, allowing precise alignment of audio chunks with video frames.
- **Conditional instantiation** allows the block to handle unimodal inputs by omitting cross-attention when one modality is absent.
- **Residual merging** with gating mechanisms ensures synchronized information integrates without overwriting original features.

## Frequently Asked Questions

### What is the primary advantage of a dual-stream transformer over single-stream architectures?

The dual-stream design preserves modality-specific representations through independent self-attention while enabling precise temporal alignment via cross-attention. This prevents early fusion from diluting distinct audio and video features, as implemented in the separate `TransformerConfig` processing paths of `BasicAVTransformerBlock`.

### How does LTX-2 handle inputs with only video or only audio?

The `BasicAVTransformerBlock` conditionally instantiates cross-attention layers only when both modalities are present (around line 51 in [`transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/transformer.py)). When processing unimodal input, the unused configuration is set to `None`, and the block skips the cross-attention stage, processing only the available stream through its dedicated self-attention and feed-forward layers.

### What role does AdaLN play in audio-video synchronization?

AdaLN (Adaptive Layer Normalization) provides per-timestep modulation through learned scale-shift tables (`self.scale_shift_table_a2v_ca_*`). The `get_av_ca_ada_values` method generates gate tensors (`gate_out_a2v`, `gate_out_v2a`) that scale cross-attention based on temporal alignment, allowing the model to emphasize or de-emphasize specific audio-video correspondences at specific timesteps.

### Where is the cross-attention mechanism implemented in the codebase?

The bidirectional cross-attention is implemented in [`packages/ltx-core/src/ltx_core/model/transformer/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer/transformer.py) within the `BasicAVTransformerBlock` class, specifically around lines 51-74. The audio-to-video attention uses `self.audio_to_video_attn`, while video-to-audio uses `self.video_to_audio_attn`, both leveraging the `Attention` class from [`attention.py`](https://github.com/Lightricks/LTX-2/blob/main/attention.py).