# How LTX-2's DiT-Based Architecture Works for Audio-Video Generation

> Discover how LTX-2's DiT-based architecture generates synchronized audio and video using a unified latent representation. Explore its innovative approach to media creation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: architecture
- Published: 2026-08-20

---

**LTX-2 uses a single DiT (Diffusion Image Transformer) backbone with lightweight connector modules and separate VAEs to generate synchronized audio and video from unified latent representations.**

The LTX-2 model from Lightricks represents a unified approach to multimodal generation, built entirely around a **DiT-based architecture** that extends beyond images to orchestrate both video frames and audio waveforms. This article breaks down how the transformer core, conditioning system, and modality-specific decoders work together based on the actual source code implementation.

## Unified DiT Transformer Core

At the heart of LTX-2 lies a **single transformer checkpoint** that contains all diffusion parameters. Rather than maintaining separate models for each modality, the architecture stores everything in one DiT while using additional *connector* modules to bridge the transformer to modality-specific components.

The command-line argument parser in [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) makes this explicit at line 549, describing the checkpoint format as *"transformer safetensors (DiT + connectors)"*. This design choice reduces parameter redundancy and enables shared representations across audio and video.

```python

# From the pipeline initialization - DiT loaded once, shared across modalities

pipeline = DistilledPipeline(
    transformer_path="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
    # ... VAEs and other components passed separately

)

```

The transformer itself operates on a **shared latent sequence** that encodes information for both modalities simultaneously.

## Multimodal Conditioning System

The DiT receives conditioning through specialized injection layers that handle three input types: **text prompts**, **video frames**, and **audio waveforms**. The trainer code implements this via `VideoConditionBy…` and `AudioConditionBy…` classes, which project latent representations into the diffusion process at the appropriate layers.

Key aspects of the conditioning architecture:

- **Text conditioning** uses a dedicated text encoder (Gemma-based in LTX-2.5)
- **Video conditioning** operates through the `VideoConditionBy…` projection classes
- **Audio conditioning** uses parallel `AudioConditionBy…` mechanisms
- **Timing masks** align audio frames with video timesteps for synchronization

This unified conditioning allows the DiT to learn cross-modal correlations—essential for lip-sync accuracy and sound-to-visual correspondence.

## Separate VAEs for Modality Decoding

While the DiT generates unified latents, the final pixel-level and waveform outputs require separate decoders. LTX-2 ships with:

| Component | Purpose | Checkpoint Example |
|-----------|---------|------------------|
| **Video VAE** | Decodes to pixel-level video frames | `ltx-2.5-video-vae-bf16.safetensors` |
| **Audio VAE** | Decodes to raw waveform samples | `ltx-2.5-audio-vae-bf16.safetensors` |

This separation allows each VAE to optimize for its modality's specific characteristics—spatial coherence for video, temporal frequency structure for audio—while the DiT handles the shared generative process.

## Optimized Linear Layers with CUDA Kernels

Performance-critical DiT operations are accelerated through custom kernels in the `ltx_kernels` package. The optimization documentation at [`packages/ltx-pipelines/docs/optimization.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/docs/optimization.md) line 15 specifically notes that **"DiT Linears"** are the target of NVFP4 casting, enabling:

- **FP4/FP8 precision** on Blackwell-class GPUs
- Significant memory reduction for the 22B parameter model
- Faster inference without quality degradation

The kernel implementation in [`packages/ltx-kernels/csrc/ops/rms_norm_split_rope.cpp`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-kernels/csrc/ops/rms_norm_split_rope.cpp) provides efficient fused operations including RMS normalization with split RoPE (Rotary Position Embedding) application—core building blocks of the DiT forward pass.

## Audio-Video Synchronization Mechanism

The DiT's shared latent representation enables inherent synchronization. The pipeline maintains temporal alignment through:

1. **Conditioning masks** that explicitly map audio frame indices to video timesteps
2. **Unified noise scheduling** applied to the combined latent sequence
3. **Cross-attention patterns** learned during training that correlate phoneme timings with mouth movements

This architecture eliminates the need for post-hoc synchronization or separate alignment models.

## Complete Inference Example

The following runnable example demonstrates the full architecture in action:

```python
import torch
from ltx_pipelines.distilled import DistilledPipeline

pipeline = DistilledPipeline(
    transformer_path="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
    text_encoder_path="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
    video_vae_path="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
    audio_vae_path="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
    spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
    device="cuda",
)

output = pipeline(
    prompt=(
        "A close‑up of a man speaking clearly while a subtle sniff "
        "sound is heard. The camera stays static, the background is a "
        "softly blurred beige wall."
    ),
    num_frames=121,
    seed=42,
)

# output[0]: video tensor, output[1]: audio tensor

video, audio = output

```

The `DistilledPipeline` class (in [`packages/ltx-pipelines/src/ltx_pipelines/distilled.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/distilled.py)) encapsulates the full DiT-based architecture: loading the shared transformer, routing through connectors, applying conditioning, and decoding through separate VAEs.

## Key Source Files

| File Path | Role in Architecture |
|-----------|---------------------|
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) | Defines checkpoint format as "DiT + connectors" (line 549) |
| [`packages/ltx-pipelines/docs/optimization.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/docs/optimization.md) | Documents NVFP4 optimization for DiT linear layers (line 15) |
| [`packages/ltx-pipelines/src/ltx_pipelines/distilled.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/distilled.py) | Main pipeline implementing DiT + connectors + VAEs |
| [`packages/ltx-kernels/csrc/ops/rms_norm_split_rope.cpp`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-kernels/csrc/ops/rms_norm_split_rope.cpp) | Custom CUDA kernels for efficient DiT operations |
| [`packages/ltx-core/src/ltx_core/utils.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/utils.py) | Core utilities for latent handling |

## Summary

- **Single DiT backbone** handles all generation parameters, shared across audio and video
- **Connector modules** link the transformer to modality-specific VAEs without duplicating parameters
- **Unified conditioning** through `VideoConditionBy…` and `AudioConditionBy…` classes enables cross-modal learning
- **Separate VAEs** decode shared latents into video frames and audio waveforms
- **Optimized kernels** in `ltx_kernels` accelerate DiT linear layers with FP4/FP8 support
- **Built-in synchronization** via conditioning masks eliminates need for post-processing alignment

## Frequently Asked Questions

### What does DiT mean in LTX-2's architecture?

**DiT stands for Diffusion Image Transformer**, originally developed for image generation. LTX-2 extends this architecture to video and audio by using a single transformer for latent diffusion, with lightweight adapters (connectors) for each modality. The core transformer checkpoint contains all ~22B parameters, making it more efficient than training separate diffusion models.

### How does LTX-2 keep audio and video synchronized?

The DiT generates a **shared latent sequence** where temporal positions are aligned through **conditioning masks** that explicitly map audio samples to video frames. During training, the model learns cross-modal attention patterns that correlate phonemes with mouth movements, eliminating the need for separate lip-sync models or post-hoc synchronization.

### Why does LTX-2 use separate VAEs if it has a unified DiT?

Separate VAEs allow **modality-specific optimization** while maintaining efficiency. The video VAE compresses spatial-temporal visual information, while the audio VAE handles frequency-domain structure. Both decode from the same DiT-generated latents, but their architectures differ—using shared parameters here would compromise reconstruction quality for either video or audio.

### Where can I find the DiT optimization code for faster inference?

The optimized linear layers are implemented in the `ltx_kernels` package, specifically `packages/ltx-kernels/csrc/ops/`. The [`rms_norm_split_rope.cpp`](https://github.com/Lightricks/LTX-2/blob/main/rms_norm_split_rope.cpp) file implements fused operations used throughout the DiT. For FP4/FP8 quantization, see [`packages/ltx-pipelines/docs/optimization.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/docs/optimization.md) at line 15, which documents the NVFP4 casting path for "DiT Linears" on compatible hardware.