How LTX-2's DiT-Based Architecture Works for Audio-Video Generation
LTX-2 uses a single DiT (Diffusion Image Transformer) backbone with lightweight connector modules and separate VAEs to generate synchronized audio and video from unified latent representations.
The LTX-2 model from Lightricks represents a unified approach to multimodal generation, built entirely around a DiT-based architecture that extends beyond images to orchestrate both video frames and audio waveforms. This article breaks down how the transformer core, conditioning system, and modality-specific decoders work together based on the actual source code implementation.
Unified DiT Transformer Core
At the heart of LTX-2 lies a single transformer checkpoint that contains all diffusion parameters. Rather than maintaining separate models for each modality, the architecture stores everything in one DiT while using additional connector modules to bridge the transformer to modality-specific components.
The command-line argument parser in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py makes this explicit at line 549, describing the checkpoint format as "transformer safetensors (DiT + connectors)". This design choice reduces parameter redundancy and enables shared representations across audio and video.
# From the pipeline initialization - DiT loaded once, shared across modalities
pipeline = DistilledPipeline(
transformer_path="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
# ... VAEs and other components passed separately
)
The transformer itself operates on a shared latent sequence that encodes information for both modalities simultaneously.
Multimodal Conditioning System
The DiT receives conditioning through specialized injection layers that handle three input types: text prompts, video frames, and audio waveforms. The trainer code implements this via VideoConditionBy… and AudioConditionBy… classes, which project latent representations into the diffusion process at the appropriate layers.
Key aspects of the conditioning architecture:
- Text conditioning uses a dedicated text encoder (Gemma-based in LTX-2.5)
- Video conditioning operates through the
VideoConditionBy…projection classes - Audio conditioning uses parallel
AudioConditionBy…mechanisms - Timing masks align audio frames with video timesteps for synchronization
This unified conditioning allows the DiT to learn cross-modal correlations—essential for lip-sync accuracy and sound-to-visual correspondence.
Separate VAEs for Modality Decoding
While the DiT generates unified latents, the final pixel-level and waveform outputs require separate decoders. LTX-2 ships with:
| Component | Purpose | Checkpoint Example |
|---|---|---|
| Video VAE | Decodes to pixel-level video frames | ltx-2.5-video-vae-bf16.safetensors |
| Audio VAE | Decodes to raw waveform samples | ltx-2.5-audio-vae-bf16.safetensors |
This separation allows each VAE to optimize for its modality's specific characteristics—spatial coherence for video, temporal frequency structure for audio—while the DiT handles the shared generative process.
Optimized Linear Layers with CUDA Kernels
Performance-critical DiT operations are accelerated through custom kernels in the ltx_kernels package. The optimization documentation at packages/ltx-pipelines/docs/optimization.md line 15 specifically notes that "DiT Linears" are the target of NVFP4 casting, enabling:
- FP4/FP8 precision on Blackwell-class GPUs
- Significant memory reduction for the 22B parameter model
- Faster inference without quality degradation
The kernel implementation in packages/ltx-kernels/csrc/ops/rms_norm_split_rope.cpp provides efficient fused operations including RMS normalization with split RoPE (Rotary Position Embedding) application—core building blocks of the DiT forward pass.
Audio-Video Synchronization Mechanism
The DiT's shared latent representation enables inherent synchronization. The pipeline maintains temporal alignment through:
- Conditioning masks that explicitly map audio frame indices to video timesteps
- Unified noise scheduling applied to the combined latent sequence
- Cross-attention patterns learned during training that correlate phoneme timings with mouth movements
This architecture eliminates the need for post-hoc synchronization or separate alignment models.
Complete Inference Example
The following runnable example demonstrates the full architecture in action:
import torch
from ltx_pipelines.distilled import DistilledPipeline
pipeline = DistilledPipeline(
transformer_path="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
text_encoder_path="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
video_vae_path="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
audio_vae_path="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
device="cuda",
)
output = pipeline(
prompt=(
"A close‑up of a man speaking clearly while a subtle sniff "
"sound is heard. The camera stays static, the background is a "
"softly blurred beige wall."
),
num_frames=121,
seed=42,
)
# output[0]: video tensor, output[1]: audio tensor
video, audio = output
The DistilledPipeline class (in packages/ltx-pipelines/src/ltx_pipelines/distilled.py) encapsulates the full DiT-based architecture: loading the shared transformer, routing through connectors, applying conditioning, and decoding through separate VAEs.
Key Source Files
| File Path | Role in Architecture |
|---|---|
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py |
Defines checkpoint format as "DiT + connectors" (line 549) |
packages/ltx-pipelines/docs/optimization.md |
Documents NVFP4 optimization for DiT linear layers (line 15) |
packages/ltx-pipelines/src/ltx_pipelines/distilled.py |
Main pipeline implementing DiT + connectors + VAEs |
packages/ltx-kernels/csrc/ops/rms_norm_split_rope.cpp |
Custom CUDA kernels for efficient DiT operations |
packages/ltx-core/src/ltx_core/utils.py |
Core utilities for latent handling |
Summary
- Single DiT backbone handles all generation parameters, shared across audio and video
- Connector modules link the transformer to modality-specific VAEs without duplicating parameters
- Unified conditioning through
VideoConditionBy…andAudioConditionBy…classes enables cross-modal learning - Separate VAEs decode shared latents into video frames and audio waveforms
- Optimized kernels in
ltx_kernelsaccelerate DiT linear layers with FP4/FP8 support - Built-in synchronization via conditioning masks eliminates need for post-processing alignment
Frequently Asked Questions
What does DiT mean in LTX-2's architecture?
DiT stands for Diffusion Image Transformer, originally developed for image generation. LTX-2 extends this architecture to video and audio by using a single transformer for latent diffusion, with lightweight adapters (connectors) for each modality. The core transformer checkpoint contains all ~22B parameters, making it more efficient than training separate diffusion models.
How does LTX-2 keep audio and video synchronized?
The DiT generates a shared latent sequence where temporal positions are aligned through conditioning masks that explicitly map audio samples to video frames. During training, the model learns cross-modal attention patterns that correlate phonemes with mouth movements, eliminating the need for separate lip-sync models or post-hoc synchronization.
Why does LTX-2 use separate VAEs if it has a unified DiT?
Separate VAEs allow modality-specific optimization while maintaining efficiency. The video VAE compresses spatial-temporal visual information, while the audio VAE handles frequency-domain structure. Both decode from the same DiT-generated latents, but their architectures differ—using shared parameters here would compromise reconstruction quality for either video or audio.
Where can I find the DiT optimization code for faster inference?
The optimized linear layers are implemented in the ltx_kernels package, specifically packages/ltx-kernels/csrc/ops/. The rms_norm_split_rope.cpp file implements fused operations used throughout the DiT. For FP4/FP8 quantization, see packages/ltx-pipelines/docs/optimization.md at line 15, which documents the NVFP4 casting path for "DiT Linears" on compatible hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →