Generating Video from Audio Using LTX-2 A2VidPipelineTwoStage: Complete Implementation Guide
LTX-2 provides a dedicated A2VidPipelineTwoStage class that converts audio files into high-resolution videos through a coarse-to-fine two-stage diffusion process.
The A2VidPipelineTwoStage in the Lightricks/LTX-2 repository implements professional-grade audio-to-video generation. This pipeline orchestrates transformer-based diffusion, spatial upsampling, and multi-modal conditioning to produce temporally synchronized video output. Below is a complete breakdown of the architecture, execution flow, and practical implementation patterns.
How A2VidPipelineTwoStage Works: Two-Stage Architecture
The A2VidPipelineTwoStage operates through two distinct diffusion stages that progressively refine video quality while maintaining audio conditioning throughout.
Stage 1: Coarse Video Generation
In the first stage, the transformer generates video at half the target resolution while keeping the audio latent frozen. Only the video modality undergoes denoising, using audio embeddings as conditioning signal.
Key implementation details from [a2vid_two_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L103-L125):
- The
DiffusionStageruns withGuidedDenoiserfor video-only diffusion - Audio latents are computed via
AudioConditionerand remain fixed - Image conditionings are encoded through
VideoEncoder
Stage 2: Upsampling and Refinement
The second stage upscales and refines the coarse output:
- Spatial upsampling: The
VideoUpsampler(defined in [blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L99-L110)) scales the low-resolution latent to target resolution - Distilled refinement: A second diffusion pass uses
SimpleDenoiserwith distilled LoRA weights for higher-fidelity output - Audio preservation: Audio latents remain frozen during video refinement
Core Components of the LTX-2 Audio-to-Video Pipeline
| Component | Purpose | Source Location |
|---|---|---|
A2VidPipelineTwoStage |
Orchestrates models, tiling, encoding, and two-stage diffusion | a2vid_two_stage.py#L53-L60 |
PromptEncoder |
Encodes text/negative prompts into video and audio contexts | a2vid_two_stage.py#L80-L89 |
AudioConditioner / VideoEncoder |
Encode raw audio to VAE latents and image conditionings | a2vid_two_stage.py#L96-L103 |
DiffusionStage |
Runs transformer with appropriate LoRA weights for each stage | a2vid_two_stage.py#L103-L125 |
VideoUpsampler |
Spatial upsampling between stages | blocks.py#L99-L110 |
VideoDecoder |
Decodes final latent to frames with optional tiling | a2vid_two_stage.py#L99-L105 |
MultiModalGuider |
Applies classifier-free guidance across video and audio modalities | a2vid_two_stage.py#L30-L37 |
Complete Execution Flow
The A2VidPipelineTwoStage follows this processing sequence as implemented in the source:
-
CLI arg parsing — handles
--audio-path,--audio-start-time,--audio-max-duration, resolution, and sampling parameters (a2vid_two_stage.py#L11-L19) -
Pipeline instantiation — loads model paths, LoRA lists, upsampler, and optional quantization settings (
a2vid_two_stage.py#L31-L38) -
Audio decoding — raw audio processed via
decode_audio_from_file(a2vid_two_stage.py#L95-L100) -
Stage 1 diffusion — image conditionings encoded,
GuidedDenoiserruns with frozen audio (a2vid_two_stage.py#L24-L58) -
Latent upsampling —
VideoUpsamplerscales to target resolution (a2vid_two_stage.py#L260-L262) -
Stage 2 refinement —
SimpleDenoiserwith distilled LoRA refines video (a2vid_two_stage.py#L73-L90) -
Final decoding —
VideoDecoderproduces frames, output muxed with original audio (a2vid_two_stage.py#L99-L105)
Running Audio-to-Video Generation: Code Examples
Command-Line Usage
python -m ltx_pipelines.a2vid_two_stage \
--prompt "A sunny day at the beach" \
--negative-prompt "" \
--seed 42 \
--height 720 \
--width 1280 \
--num-frames 120 \
--frame-rate 30 \
--num-inference-steps 50 \
--audio-path /path/to/music.wav \
--output-path result.mp4 \
--model-paths <model_dir> \
--spatial-upsampler-path <upsampler_dir>
Key parameters:
--audio-path— source audio file for conditioning--num-inference-steps— diffusion iterations (higher = better quality, slower)--spatial-upsampler-path— required for stage 2 resolution boost
Python API: Direct Pipeline Invocation
from ltx_pipelines.a2vid_two_stage import A2VidPipelineTwoStage
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.guider_params import MultiModalGuiderParams
from ltx_core.loader import LoraPathStrengthAndSDOps
import torch
# Configure model locations
paths = ModelPaths(root="<model_root>")
distilled_lora = [
LoraPathStrengthAndSDOps(path="<distilled_lora_path>", strength=1.0)
]
spatial_upsampler = "<upsampler_path>"
# Initialize pipeline
pipeline = A2VidPipelineTwoStage(
model_paths=paths,
distilled_lora=distilled_lora,
spatial_upsampler_path=spatial_upsampler,
loras=[], # optional additional LoRAs
device=torch.device("cuda"),
)
# Generate video from audio
video_iter, audio, tiling = pipeline(
prompt="A futuristic city at night",
negative_prompt="low quality",
seed=1234,
height=1080,
width=1920,
num_frames=150,
frame_rate=30.0,
num_inference_steps=60,
video_guider_params=MultiModalGuiderParams(
cfg_scale=7.5,
modality_scale=1.5
),
images=[], # optional image conditionings
audio_path="song.wav",
)
# Process output frames
for frame_tensor in video_iter:
# Convert to numpy and encode to video
pass
Production Integration: Cached Pipeline Instance
from functools import lru_cache
@lru_cache(maxsize=1)
def get_preinitialized_a2v_pipeline():
"""Return cached pipeline to avoid repeated model loading."""
paths = ModelPaths(root="/models/ltx2")
return A2VidPipelineTwoStage(
model_paths=paths,
distilled_lora=[...],
spatial_upsampler_path="/models/upsampler",
device=torch.device("cuda"),
)
def generate_a2v_video(prompt: str, audio_file: str):
"""Generate video with pre-loaded pipeline."""
pipeline = get_preinitialized_a2v_pipeline()
video_iter, audio, _ = pipeline(
prompt=prompt,
negative_prompt="",
seed=0,
height=720,
width=1280,
num_frames=60,
frame_rate=24,
num_inference_steps=40,
video_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
images=[],
audio_path=audio_file,
)
# Encode to final video format
return video_iter, audio
Caching the pipeline instance eliminates model loading overhead—critical for server-side inference scenarios.
Key Source Files for Deep Customization
Summary
A2VidPipelineTwoStageimplements coarse-to-fine audio-to-video generation in LTX-2- Stage 1 generates half-resolution video with frozen audio conditioning
VideoUpsamplerbridges stages by spatially scaling latents- Stage 2 applies distilled LoRA refinement at full resolution
- Multi-modal guidance operates through
MultiModalGuiderwith independent CFG scales - Both CLI and Python APIs support full parameter control including tiling, quantization, and custom LoRA loading
Frequently Asked Questions
What audio formats does A2VidPipelineTwoStage support?
The pipeline accepts standard audio formats through decode_audio_from_file in [media_io.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py). WAV and MP3 are explicitly supported, with automatic resampling to the model's expected sample rate. Use --audio-start-time and --audio-max-duration to extract specific segments.
How does the spatial upsampler improve video quality?
The VideoUpsampler (defined in [blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L99-L110)) implements learned spatial interpolation that preserves temporal consistency better than naive resizing. According to the LTX-2 source, this upsampler is specifically trained to reduce artifacts when doubling spatial resolution between diffusion stages.
Can I use custom LoRA weights with A2VidPipelineTwoStage?
Yes. Pass additional LoRAs via the loras parameter in the constructor or --lora-path via CLI. The pipeline composes these with the required distilled_lora for stage 2. Each LoRA is wrapped in LoraPathStrengthAndSDOps to specify blending strength and selective layer application.
Why are there two separate guider configurations for video and audio?
The MultiModalGuiderParams (defined in [a2vid_two_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L30-L37)) separates cfg_scale (classifier-free guidance strength) from modality_scale (cross-modal conditioning weight). This decoupling allows fine-tuning how strongly the audio influences video generation independently from the unconditional guidance scale—critical for balancing temporal synchronization with visual fidelity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →