How to Use A2VidPipeline for Audio-to-Video Generation in LTX-2
Instantiate A2VidPipelineTwoStage with model paths, audio file, and generation parameters to produce synchronized video from audio input using a two-stage diffusion process.
The A2VidPipeline in Lightricks' LTX-2 repository provides a production-ready audio-to-video generation system that transforms audio tracks into visually coherent videos through a sophisticated two-stage diffusion pipeline. This guide walks through the complete implementation, from architecture to practical usage, based on the actual source code in packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py.
Understanding the Two-Stage Architecture
The A2VidPipelineTwoStage class implements a deliberate coarse-to-fine generation strategy that separates initial structure creation from high-fidelity refinement.
Stage 1: Low-Resolution Structure Generation
Purpose: Generate base video at half target resolution while preserving audio fidelity through frozen conditioning.
The first stage uses a GuidedDenoiser that processes video and audio latents together. Critically, the audio latent is frozen with noise_scale=0.0, meaning it acts purely as conditioning guidance rather than being noised and denoised. From the source:
video_state, _ = self.stage_1(
denoiser=GuidedDenoiser(...),
sigmas=stage_1_sigmas,
video=ModalitySpec(context=v_context_p, conditionings=stage_1_conditionings),
audio=ModalitySpec(context=a_context_p, frozen=True, noise_scale=0.0,
initial_latent=encoded_audio_latent),
...
)
Source: /packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py, lines 124-158
The GuidedDenoiser handles multimodal guidance through classifier-free guidance (CFG) and modality-specific scaling, letting the audio strongly influence visual generation patterns.
Stage 2: Upscaled Refinement with Distilled LoRA
Purpose: Double spatial resolution and refine details using a distilled LoRA for faster, higher-quality inference.
After upsampling via VideoUpsampler, the second stage employs SimpleDenoiser with specialized stage_2_loras:
video_state, _ = self.stage_2(
denoiser=SimpleDenoiser(v_context_p, a_context_p),
sigmas=stage_2_sigmas,
video=ModalitySpec(..., noise_scale=stage_2_sigmas[0].item(),
initial_latent=upscaled_video_latent),
audio=ModalitySpec(..., frozen=True, initial_latent=encoded_audio_latent),
...
)
Source: /packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py, lines 187-224
The distilled LoRA enables fewer inference steps while maintaining quality—a key efficiency optimization in the LTX-2 pipeline.
Core Pipeline Components
Initialization Flow
The __init__ method in A2VidPipelineTwoStage (lines 61-145) establishes all conditioning and generation modules:
self.prompt_encoder = PromptEncoder(...)
self.image_conditioner = ImageConditioner(...)
self.audio_conditioner = AudioConditioner(...)
self.stage_1 = DiffusionStage.from_checkpoint(...)
self.stage_2 = DiffusionStage.from_checkpoint(...)
self.upsampler = VideoUpsampler(...)
self.video_decoder = VideoDecoder(...)
All components share:
- dtype:
bfloat16for memory efficiency - device: Unified CUDA placement
- scheduler:
LTX2Schedulerfor consistent noise scheduling
Audio Conditioning Pipeline
Audio processing follows a clear encoding path in the __call__ method (lines 95-103):
decoded_audio = decode_audio_from_file(audio_path, self.device, ...)
encoded_audio_latent = self.audio_conditioner(
lambda enc: vae_encode_audio(decoded_audio, enc, None)
)
The AudioConditioner (defined in /packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) wraps the audio VAE and provides consistent latent representations of shape AudioLatentShape that remain frozen across both generation stages.
Command-Line Usage for A2VidPipeline
The fastest way to experiment with audio-to-video generation is through the built-in CLI:
python -m ltx_pipelines.a2vid_two_stage \
--model-paths /path/to/models \
--distilled-lora /path/to/distilled_lora.pt \
--spatial-upsampler-path /path/to/upsampler.pt \
--prompt "A sunrise over a futuristic city" \
--negative-prompt "" \
--seed 42 \
--height 720 \
--width 1280 \
--num-frames 120 \
--frame-rate 30 \
--num-inference-steps 50 \
--audio-path /data/audio/my_track.wav \
--output-path result.mp4
Essential Audio Parameters
| Parameter | Function | Default |
|---|---|---|
--audio-path |
Source audio file path | Required |
--audio-start-time |
Seconds to skip at audio start | 0.0 |
--audio-max-duration |
Maximum audio length in seconds | Match video duration |
Additional generation flags include --enhance-prompt for LLM-based prompt refinement and --video-cfg-guidance-scale for adjusting classifier-free guidance strength.
Programmatic A2VidPipeline Integration
For production systems requiring custom workflows, instantiate the pipeline directly:
from ltx_pipelines.a2vid_two_stage import A2VidPipelineTwoStage
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_core.loader.registry import Registry
import torch
# Prepare model configuration
model_paths = ModelPaths(root="/models")
distilled_lora = [...] # List[PathStrengthAndSDOps]
spatial_upsampler = "/models/upsampler.pt"
loras = [] # Optional additional LoRAs
# Initialize pipeline
pipeline = A2VidPipelineTwoStage(
model_paths=model_paths,
distilled_lora=distilled_lora,
spatial_upsampler_path=spatial_upsampler,
loras=loras,
device=torch.device("cuda"),
offload_mode=OffloadMode.NONE,
)
# Execute audio-to-video generation
video, audio, tiling_cfg = pipeline(
prompt="A rainy night in neon Tokyo",
negative_prompt="low quality",
seed=1234,
height=720,
width=1280,
num_frames=150,
frame_rate=30.0,
num_inference_steps=60,
video_guider_params=MultiModalGuiderParams(
cfg_scale=7.5,
modality_scale=1.0,
),
images=[], # Empty list = no image conditioning
audio_path="/data/audio/song.wav",
audio_start_time=0.0,
audio_max_duration=None,
)
Saving Generated Output
The pipeline returns raw tensors and audio objects that require encoding:
from ltx_pipelines.utils.media_io.encode import encode_video
from ltx_pipelines.utils.media_io.decode import get_video_chunks_number
tiling_cfg = tiling_cfg or AUTO_TILING
chunks = get_video_chunks_number(num_frames=150, tiling_cfg=tiling_cfg)
encode_video(
video=video,
fps=30.0,
audio=audio, # Original waveform, not VAE-decoded
output_path="generated.mp4",
video_chunks_number=chunks,
color_space=HDRColorSpace.SDR,
)
Memory Optimization with Tiling
For GPUs unable to hold full-resolution video latents, the A2VidPipeline supports automatic tiling:
- Set
tiling_config=AUTO_TILINGin the pipeline call - The
VideoDecoder(from/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) handles spatial tiling and seamless stitching - Tiling configuration is returned as
tiling_cfgfor proper encoding alignment
Key Source Files Reference
| Component | Implementation Location |
|---|---|
A2VidPipelineTwoStage |
/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py |
PromptEncoder, ImageConditioner, AudioConditioner, DiffusionStage, VideoUpsampler, VideoDecoder |
/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py |
GuidedDenoiser, SimpleDenoiser |
/packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py |
| Audio/video I/O utilities | /packages/ltx-pipelines/src/ltx_pipelines/utils/media_io/ |
LTX2Scheduler |
/packages/ltx-core/src/ltx_core/components/schedulers.py |
Summary
- A2VidPipeline implements two-stage audio-to-video generation: Stage 1 creates half-resolution structure with frozen audio guidance, Stage 2 upscales and refines with distilled LoRA
- Audio is encoded once via
AudioConditionerand kept frozen (noise_scale=0.0) across both stages to preserve synchronization fidelity - The pipeline supports both CLI execution (
python -m ltx_pipelines.a2vid_two_stage) and Python instantiation - Memory-constrained deployments benefit from automatic tiling in
VideoDecoder - The original audio waveform (not VAE-decoded) is returned and should be used for final encoding to maintain quality
Frequently Asked Questions
What audio formats does A2VidPipeline support?
The pipeline uses decode_audio_from_file from the media I/O utilities, which handles standard formats including WAV, MP3, and FLAC through underlying torchaudio backends. The audio is resampled to the model's expected sampling rate during encoding.
Why is the audio latent frozen instead of being generated?
Freezing the audio latent with noise_scale=0.0 ensures perfect audio fidelity preservation. Unlike video which becomes coherent through denoising, audio would degrade if subjected to the diffusion process. The frozen latent serves as deterministic conditioning that guides visual generation to match audio rhythm and structure.
How do image conditionings work with audio input?
The images parameter accepts a list of conditioning images that are processed by ImageConditioner at each stage's native resolution (half resolution for Stage 1, full resolution for Stage 2). These combine with audio conditioning through the multimodal GuidedDenoiser, allowing synchronized visual-audio generation with specific visual references.
What's the difference between regular LoRAs and distilled LoRA?
Regular LoRAs load into stage_1 for initial generation, while distilled LoRA (passed via distilled_lora parameter) specializes stage_2 for efficient high-quality refinement. The distilled variant enables faster convergence with fewer inference steps, making the second stage practical for production deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →