How to Use A2VidPipeline for Audio-to-Video Generation in LTX-2

Instantiate A2VidPipelineTwoStage with model paths, audio file, and generation parameters to produce synchronized video from audio input using a two-stage diffusion process.

The A2VidPipeline in Lightricks' LTX-2 repository provides a production-ready audio-to-video generation system that transforms audio tracks into visually coherent videos through a sophisticated two-stage diffusion pipeline. This guide walks through the complete implementation, from architecture to practical usage, based on the actual source code in packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py.

Understanding the Two-Stage Architecture

The A2VidPipelineTwoStage class implements a deliberate coarse-to-fine generation strategy that separates initial structure creation from high-fidelity refinement.

Stage 1: Low-Resolution Structure Generation

Purpose: Generate base video at half target resolution while preserving audio fidelity through frozen conditioning.

The first stage uses a GuidedDenoiser that processes video and audio latents together. Critically, the audio latent is frozen with noise_scale=0.0, meaning it acts purely as conditioning guidance rather than being noised and denoised. From the source:

video_state, _ = self.stage_1(
    denoiser=GuidedDenoiser(...),
    sigmas=stage_1_sigmas,
    video=ModalitySpec(context=v_context_p, conditionings=stage_1_conditionings),
    audio=ModalitySpec(context=a_context_p, frozen=True, noise_scale=0.0,
                      initial_latent=encoded_audio_latent),
    ...
)

Source: /packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py, lines 124-158

The GuidedDenoiser handles multimodal guidance through classifier-free guidance (CFG) and modality-specific scaling, letting the audio strongly influence visual generation patterns.

Stage 2: Upscaled Refinement with Distilled LoRA

Purpose: Double spatial resolution and refine details using a distilled LoRA for faster, higher-quality inference.

After upsampling via VideoUpsampler, the second stage employs SimpleDenoiser with specialized stage_2_loras:

video_state, _ = self.stage_2(
    denoiser=SimpleDenoiser(v_context_p, a_context_p),
    sigmas=stage_2_sigmas,
    video=ModalitySpec(..., noise_scale=stage_2_sigmas[0].item(),
                      initial_latent=upscaled_video_latent),
    audio=ModalitySpec(..., frozen=True, initial_latent=encoded_audio_latent),
    ...
)

Source: /packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py, lines 187-224

The distilled LoRA enables fewer inference steps while maintaining quality—a key efficiency optimization in the LTX-2 pipeline.

Core Pipeline Components

Initialization Flow

The __init__ method in A2VidPipelineTwoStage (lines 61-145) establishes all conditioning and generation modules:

self.prompt_encoder = PromptEncoder(...)
self.image_conditioner = ImageConditioner(...)
self.audio_conditioner = AudioConditioner(...)
self.stage_1 = DiffusionStage.from_checkpoint(...)
self.stage_2 = DiffusionStage.from_checkpoint(...)
self.upsampler = VideoUpsampler(...)
self.video_decoder = VideoDecoder(...)

All components share:

  • dtype: bfloat16 for memory efficiency
  • device: Unified CUDA placement
  • scheduler: LTX2Scheduler for consistent noise scheduling

Audio Conditioning Pipeline

Audio processing follows a clear encoding path in the __call__ method (lines 95-103):

decoded_audio = decode_audio_from_file(audio_path, self.device, ...)
encoded_audio_latent = self.audio_conditioner(
    lambda enc: vae_encode_audio(decoded_audio, enc, None)
)

The AudioConditioner (defined in /packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) wraps the audio VAE and provides consistent latent representations of shape AudioLatentShape that remain frozen across both generation stages.

Command-Line Usage for A2VidPipeline

The fastest way to experiment with audio-to-video generation is through the built-in CLI:

python -m ltx_pipelines.a2vid_two_stage \
    --model-paths /path/to/models \
    --distilled-lora /path/to/distilled_lora.pt \
    --spatial-upsampler-path /path/to/upsampler.pt \
    --prompt "A sunrise over a futuristic city" \
    --negative-prompt "" \
    --seed 42 \
    --height 720 \
    --width 1280 \
    --num-frames 120 \
    --frame-rate 30 \
    --num-inference-steps 50 \
    --audio-path /data/audio/my_track.wav \
    --output-path result.mp4

Essential Audio Parameters

Parameter Function Default
--audio-path Source audio file path Required
--audio-start-time Seconds to skip at audio start 0.0
--audio-max-duration Maximum audio length in seconds Match video duration

Additional generation flags include --enhance-prompt for LLM-based prompt refinement and --video-cfg-guidance-scale for adjusting classifier-free guidance strength.

Programmatic A2VidPipeline Integration

For production systems requiring custom workflows, instantiate the pipeline directly:

from ltx_pipelines.a2vid_two_stage import A2VidPipelineTwoStage
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_core.loader.registry import Registry
import torch

# Prepare model configuration

model_paths = ModelPaths(root="/models")
distilled_lora = [...]  # List[PathStrengthAndSDOps]

spatial_upsampler = "/models/upsampler.pt"
loras = []  # Optional additional LoRAs

# Initialize pipeline

pipeline = A2VidPipelineTwoStage(
    model_paths=model_paths,
    distilled_lora=distilled_lora,
    spatial_upsampler_path=spatial_upsampler,
    loras=loras,
    device=torch.device("cuda"),
    offload_mode=OffloadMode.NONE,
)

# Execute audio-to-video generation

video, audio, tiling_cfg = pipeline(
    prompt="A rainy night in neon Tokyo",
    negative_prompt="low quality",
    seed=1234,
    height=720,
    width=1280,
    num_frames=150,
    frame_rate=30.0,
    num_inference_steps=60,
    video_guider_params=MultiModalGuiderParams(
        cfg_scale=7.5,
        modality_scale=1.0,
    ),
    images=[],  # Empty list = no image conditioning

    audio_path="/data/audio/song.wav",
    audio_start_time=0.0,
    audio_max_duration=None,
)

Saving Generated Output

The pipeline returns raw tensors and audio objects that require encoding:

from ltx_pipelines.utils.media_io.encode import encode_video
from ltx_pipelines.utils.media_io.decode import get_video_chunks_number

tiling_cfg = tiling_cfg or AUTO_TILING
chunks = get_video_chunks_number(num_frames=150, tiling_cfg=tiling_cfg)

encode_video(
    video=video,
    fps=30.0,
    audio=audio,  # Original waveform, not VAE-decoded

    output_path="generated.mp4",
    video_chunks_number=chunks,
    color_space=HDRColorSpace.SDR,
)

Memory Optimization with Tiling

For GPUs unable to hold full-resolution video latents, the A2VidPipeline supports automatic tiling:

Key Source Files Reference

Component Implementation Location
A2VidPipelineTwoStage /packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py
PromptEncoder, ImageConditioner, AudioConditioner, DiffusionStage, VideoUpsampler, VideoDecoder /packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py
GuidedDenoiser, SimpleDenoiser /packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py
Audio/video I/O utilities /packages/ltx-pipelines/src/ltx_pipelines/utils/media_io/
LTX2Scheduler /packages/ltx-core/src/ltx_core/components/schedulers.py

Summary

  • A2VidPipeline implements two-stage audio-to-video generation: Stage 1 creates half-resolution structure with frozen audio guidance, Stage 2 upscales and refines with distilled LoRA
  • Audio is encoded once via AudioConditioner and kept frozen (noise_scale=0.0) across both stages to preserve synchronization fidelity
  • The pipeline supports both CLI execution (python -m ltx_pipelines.a2vid_two_stage) and Python instantiation
  • Memory-constrained deployments benefit from automatic tiling in VideoDecoder
  • The original audio waveform (not VAE-decoded) is returned and should be used for final encoding to maintain quality

Frequently Asked Questions

What audio formats does A2VidPipeline support?

The pipeline uses decode_audio_from_file from the media I/O utilities, which handles standard formats including WAV, MP3, and FLAC through underlying torchaudio backends. The audio is resampled to the model's expected sampling rate during encoding.

Why is the audio latent frozen instead of being generated?

Freezing the audio latent with noise_scale=0.0 ensures perfect audio fidelity preservation. Unlike video which becomes coherent through denoising, audio would degrade if subjected to the diffusion process. The frozen latent serves as deterministic conditioning that guides visual generation to match audio rhythm and structure.

How do image conditionings work with audio input?

The images parameter accepts a list of conditioning images that are processed by ImageConditioner at each stage's native resolution (half resolution for Stage 1, full resolution for Stage 2). These combine with audio conditioning through the multimodal GuidedDenoiser, allowing synchronized visual-audio generation with specific visual references.

What's the difference between regular LoRAs and distilled LoRA?

Regular LoRAs load into stage_1 for initial generation, while distilled LoRA (passed via distilled_lora parameter) specializes stage_2 for efficient high-quality refinement. The distilled variant enables faster convergence with fewer inference steps, making the second stage practical for production deployment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →