Generating Video from Audio Using LTX-2 A2VidPipelineTwoStage: Complete Implementation Guide

LTX-2 provides a dedicated A2VidPipelineTwoStage class that converts audio files into high-resolution videos through a coarse-to-fine two-stage diffusion process.

The A2VidPipelineTwoStage in the Lightricks/LTX-2 repository implements professional-grade audio-to-video generation. This pipeline orchestrates transformer-based diffusion, spatial upsampling, and multi-modal conditioning to produce temporally synchronized video output. Below is a complete breakdown of the architecture, execution flow, and practical implementation patterns.

How A2VidPipelineTwoStage Works: Two-Stage Architecture

The A2VidPipelineTwoStage operates through two distinct diffusion stages that progressively refine video quality while maintaining audio conditioning throughout.

Stage 1: Coarse Video Generation

In the first stage, the transformer generates video at half the target resolution while keeping the audio latent frozen. Only the video modality undergoes denoising, using audio embeddings as conditioning signal.

Key implementation details from [a2vid_two_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L103-L125):

  • The DiffusionStage runs with GuidedDenoiser for video-only diffusion
  • Audio latents are computed via AudioConditioner and remain fixed
  • Image conditionings are encoded through VideoEncoder

Stage 2: Upsampling and Refinement

The second stage upscales and refines the coarse output:

  1. Spatial upsampling: The VideoUpsampler (defined in [blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L99-L110)) scales the low-resolution latent to target resolution
  2. Distilled refinement: A second diffusion pass uses SimpleDenoiser with distilled LoRA weights for higher-fidelity output
  3. Audio preservation: Audio latents remain frozen during video refinement

Core Components of the LTX-2 Audio-to-Video Pipeline

Component Purpose Source Location
A2VidPipelineTwoStage Orchestrates models, tiling, encoding, and two-stage diffusion a2vid_two_stage.py#L53-L60
PromptEncoder Encodes text/negative prompts into video and audio contexts a2vid_two_stage.py#L80-L89
AudioConditioner / VideoEncoder Encode raw audio to VAE latents and image conditionings a2vid_two_stage.py#L96-L103
DiffusionStage Runs transformer with appropriate LoRA weights for each stage a2vid_two_stage.py#L103-L125
VideoUpsampler Spatial upsampling between stages blocks.py#L99-L110
VideoDecoder Decodes final latent to frames with optional tiling a2vid_two_stage.py#L99-L105
MultiModalGuider Applies classifier-free guidance across video and audio modalities a2vid_two_stage.py#L30-L37

Complete Execution Flow

The A2VidPipelineTwoStage follows this processing sequence as implemented in the source:

  1. CLI arg parsing — handles --audio-path, --audio-start-time, --audio-max-duration, resolution, and sampling parameters (a2vid_two_stage.py#L11-L19)

  2. Pipeline instantiation — loads model paths, LoRA lists, upsampler, and optional quantization settings (a2vid_two_stage.py#L31-L38)

  3. Audio decoding — raw audio processed via decode_audio_from_file (a2vid_two_stage.py#L95-L100)

  4. Stage 1 diffusion — image conditionings encoded, GuidedDenoiser runs with frozen audio (a2vid_two_stage.py#L24-L58)

  5. Latent upsampling — VideoUpsampler scales to target resolution (a2vid_two_stage.py#L260-L262)

  6. Stage 2 refinement — SimpleDenoiser with distilled LoRA refines video (a2vid_two_stage.py#L73-L90)

  7. Final decoding — VideoDecoder produces frames, output muxed with original audio (a2vid_two_stage.py#L99-L105)

Running Audio-to-Video Generation: Code Examples

Command-Line Usage

python -m ltx_pipelines.a2vid_two_stage \
    --prompt "A sunny day at the beach" \
    --negative-prompt "" \
    --seed 42 \
    --height 720 \
    --width 1280 \
    --num-frames 120 \
    --frame-rate 30 \
    --num-inference-steps 50 \
    --audio-path /path/to/music.wav \
    --output-path result.mp4 \
    --model-paths <model_dir> \
    --spatial-upsampler-path <upsampler_dir>

Key parameters:

  • --audio-path — source audio file for conditioning
  • --num-inference-steps — diffusion iterations (higher = better quality, slower)
  • --spatial-upsampler-path — required for stage 2 resolution boost

Python API: Direct Pipeline Invocation

from ltx_pipelines.a2vid_two_stage import A2VidPipelineTwoStage
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.guider_params import MultiModalGuiderParams
from ltx_core.loader import LoraPathStrengthAndSDOps
import torch

# Configure model locations

paths = ModelPaths(root="<model_root>")
distilled_lora = [
    LoraPathStrengthAndSDOps(path="<distilled_lora_path>", strength=1.0)
]
spatial_upsampler = "<upsampler_path>"

# Initialize pipeline

pipeline = A2VidPipelineTwoStage(
    model_paths=paths,
    distilled_lora=distilled_lora,
    spatial_upsampler_path=spatial_upsampler,
    loras=[],  # optional additional LoRAs

    device=torch.device("cuda"),
)

# Generate video from audio

video_iter, audio, tiling = pipeline(
    prompt="A futuristic city at night",
    negative_prompt="low quality",
    seed=1234,
    height=1080,
    width=1920,
    num_frames=150,
    frame_rate=30.0,
    num_inference_steps=60,
    video_guider_params=MultiModalGuiderParams(
        cfg_scale=7.5,
        modality_scale=1.5
    ),
    images=[],  # optional image conditionings

    audio_path="song.wav",
)

# Process output frames

for frame_tensor in video_iter:
    # Convert to numpy and encode to video

    pass

Production Integration: Cached Pipeline Instance

from functools import lru_cache

@lru_cache(maxsize=1)
def get_preinitialized_a2v_pipeline():
    """Return cached pipeline to avoid repeated model loading."""
    paths = ModelPaths(root="/models/ltx2")
    return A2VidPipelineTwoStage(
        model_paths=paths,
        distilled_lora=[...],
        spatial_upsampler_path="/models/upsampler",
        device=torch.device("cuda"),
    )

def generate_a2v_video(prompt: str, audio_file: str):
    """Generate video with pre-loaded pipeline."""
    pipeline = get_preinitialized_a2v_pipeline()
    
    video_iter, audio, _ = pipeline(
        prompt=prompt,
        negative_prompt="",
        seed=0,
        height=720,
        width=1280,
        num_frames=60,
        frame_rate=24,
        num_inference_steps=40,
        video_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
        images=[],
        audio_path=audio_file,
    )
    
    # Encode to final video format

    return video_iter, audio

Caching the pipeline instance eliminates model loading overhead—critical for server-side inference scenarios.

Key Source Files for Deep Customization

File Description
[packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py) Core two-stage pipeline and CLI
[packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) VideoUpsampler, VideoDecoder, AudioConditioner
[packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) Audio decoding, video encoding, HDR utilities
[packages/ltx-core/src/ltx_core/model/audio_vae.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/audio_vae.py) Audio VAE encoder (vae_encode_audio)
[packages/ltx-core/src/ltx_core/model/video_vae.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae.py) Video VAE with tiling support
[packages/ltx-core/src/ltx_core/components/guiders.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/guiders.py) MultiModalGuider for CFG
[packages/ltx-core/src/ltx_core/components/noisers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/noisers.py) GaussianNoiser for diffusion

Summary

  • A2VidPipelineTwoStage implements coarse-to-fine audio-to-video generation in LTX-2
  • Stage 1 generates half-resolution video with frozen audio conditioning
  • VideoUpsampler bridges stages by spatially scaling latents
  • Stage 2 applies distilled LoRA refinement at full resolution
  • Multi-modal guidance operates through MultiModalGuider with independent CFG scales
  • Both CLI and Python APIs support full parameter control including tiling, quantization, and custom LoRA loading

Frequently Asked Questions

What audio formats does A2VidPipelineTwoStage support?

The pipeline accepts standard audio formats through decode_audio_from_file in [media_io.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py). WAV and MP3 are explicitly supported, with automatic resampling to the model's expected sample rate. Use --audio-start-time and --audio-max-duration to extract specific segments.

How does the spatial upsampler improve video quality?

The VideoUpsampler (defined in [blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L99-L110)) implements learned spatial interpolation that preserves temporal consistency better than naive resizing. According to the LTX-2 source, this upsampler is specifically trained to reduce artifacts when doubling spatial resolution between diffusion stages.

Can I use custom LoRA weights with A2VidPipelineTwoStage?

Yes. Pass additional LoRAs via the loras parameter in the constructor or --lora-path via CLI. The pipeline composes these with the required distilled_lora for stage 2. Each LoRA is wrapped in LoraPathStrengthAndSDOps to specify blending strength and selective layer application.

Why are there two separate guider configurations for video and audio?

The MultiModalGuiderParams (defined in [a2vid_two_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L30-L37)) separates cfg_scale (classifier-free guidance strength) from modality_scale (cross-modal conditioning weight). This decoupling allows fine-tuning how strongly the audio influences video generation independently from the unconditional guidance scale—critical for balancing temporal synchronization with visual fidelity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →