# Generating Video from Audio Using LTX-2 A2VidPipelineTwoStage: Complete Implementation Guide

> Learn to generate high-resolution videos from audio with LTX-2 A2VidPipelineTwoStage. This guide details the complete implementation of a two-stage diffusion process.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-15

---

**LTX-2 provides a dedicated `A2VidPipelineTwoStage` class that converts audio files into high-resolution videos through a coarse-to-fine two-stage diffusion process.**

The `A2VidPipelineTwoStage` in the [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2) repository implements professional-grade audio-to-video generation. This pipeline orchestrates transformer-based diffusion, spatial upsampling, and multi-modal conditioning to produce temporally synchronized video output. Below is a complete breakdown of the architecture, execution flow, and practical implementation patterns.

## How A2VidPipelineTwoStage Works: Two-Stage Architecture

The `A2VidPipelineTwoStage` operates through **two distinct diffusion stages** that progressively refine video quality while maintaining audio conditioning throughout.

### Stage 1: Coarse Video Generation

In the first stage, the transformer generates video at **half the target resolution** while keeping the audio latent frozen. Only the video modality undergoes denoising, using audio embeddings as conditioning signal.

Key implementation details from [[`a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/a2vid_two_stage.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L103-L125):

- The `DiffusionStage` runs with `GuidedDenoiser` for video-only diffusion
- Audio latents are computed via `AudioConditioner` and remain fixed
- Image conditionings are encoded through `VideoEncoder`

### Stage 2: Upsampling and Refinement

The second stage upscales and refines the coarse output:

1. **Spatial upsampling**: The `VideoUpsampler` (defined in [[`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L99-L110)) scales the low-resolution latent to target resolution
2. **Distilled refinement**: A second diffusion pass uses `SimpleDenoiser` with **distilled LoRA weights** for higher-fidelity output
3. **Audio preservation**: Audio latents remain frozen during video refinement

## Core Components of the LTX-2 Audio-to-Video Pipeline

| Component | Purpose | Source Location |
|-----------|---------|---------------|
| `A2VidPipelineTwoStage` | Orchestrates models, tiling, encoding, and two-stage diffusion | [`a2vid_two_stage.py#L53-L60`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L53-L60) |
| `PromptEncoder` | Encodes text/negative prompts into video and audio contexts | [`a2vid_two_stage.py#L80-L89`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L80-L89) |
| `AudioConditioner` / `VideoEncoder` | Encode raw audio to VAE latents and image conditionings | [`a2vid_two_stage.py#L96-L103`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L96-L103) |
| `DiffusionStage` | Runs transformer with appropriate LoRA weights for each stage | [`a2vid_two_stage.py#L103-L125`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L103-L125) |
| `VideoUpsampler` | Spatial upsampling between stages | [`blocks.py#L99-L110`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L99-L110) |
| `VideoDecoder` | Decodes final latent to frames with optional tiling | [`a2vid_two_stage.py#L99-L105`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L99-L105) |
| `MultiModalGuider` | Applies classifier-free guidance across video and audio modalities | [`a2vid_two_stage.py#L30-L37`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L30-L37) |

## Complete Execution Flow

The `A2VidPipelineTwoStage` follows this processing sequence as implemented in the source:

1. **CLI arg parsing** — handles `--audio-path`, `--audio-start-time`, `--audio-max-duration`, resolution, and sampling parameters ([`a2vid_two_stage.py#L11-L19`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L11-L19))

2. **Pipeline instantiation** — loads model paths, LoRA lists, upsampler, and optional quantization settings ([`a2vid_two_stage.py#L31-L38`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L31-L38))

3. **Audio decoding** — raw audio processed via `decode_audio_from_file` ([`a2vid_two_stage.py#L95-L100`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L95-L100))

4. **Stage 1 diffusion** — image conditionings encoded, `GuidedDenoiser` runs with frozen audio ([`a2vid_two_stage.py#L24-L58`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L24-L58))

5. **Latent upsampling** — `VideoUpsampler` scales to target resolution ([`a2vid_two_stage.py#L260-L262`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L260-L262))

6. **Stage 2 refinement** — `SimpleDenoiser` with distilled LoRA refines video ([`a2vid_two_stage.py#L73-L90`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L73-L90))

7. **Final decoding** — `VideoDecoder` produces frames, output muxed with original audio ([`a2vid_two_stage.py#L99-L105`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L99-L105))

## Running Audio-to-Video Generation: Code Examples

### Command-Line Usage

```bash
python -m ltx_pipelines.a2vid_two_stage \
    --prompt "A sunny day at the beach" \
    --negative-prompt "" \
    --seed 42 \
    --height 720 \
    --width 1280 \
    --num-frames 120 \
    --frame-rate 30 \
    --num-inference-steps 50 \
    --audio-path /path/to/music.wav \
    --output-path result.mp4 \
    --model-paths <model_dir> \
    --spatial-upsampler-path <upsampler_dir>

```

**Key parameters:**
- `--audio-path` — source audio file for conditioning
- `--num-inference-steps` — diffusion iterations (higher = better quality, slower)
- `--spatial-upsampler-path` — required for stage 2 resolution boost

### Python API: Direct Pipeline Invocation

```python
from ltx_pipelines.a2vid_two_stage import A2VidPipelineTwoStage
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.guider_params import MultiModalGuiderParams
from ltx_core.loader import LoraPathStrengthAndSDOps
import torch

# Configure model locations

paths = ModelPaths(root="<model_root>")
distilled_lora = [
    LoraPathStrengthAndSDOps(path="<distilled_lora_path>", strength=1.0)
]
spatial_upsampler = "<upsampler_path>"

# Initialize pipeline

pipeline = A2VidPipelineTwoStage(
    model_paths=paths,
    distilled_lora=distilled_lora,
    spatial_upsampler_path=spatial_upsampler,
    loras=[],  # optional additional LoRAs

    device=torch.device("cuda"),
)

# Generate video from audio

video_iter, audio, tiling = pipeline(
    prompt="A futuristic city at night",
    negative_prompt="low quality",
    seed=1234,
    height=1080,
    width=1920,
    num_frames=150,
    frame_rate=30.0,
    num_inference_steps=60,
    video_guider_params=MultiModalGuiderParams(
        cfg_scale=7.5,
        modality_scale=1.5
    ),
    images=[],  # optional image conditionings

    audio_path="song.wav",
)

# Process output frames

for frame_tensor in video_iter:
    # Convert to numpy and encode to video

    pass

```

### Production Integration: Cached Pipeline Instance

```python
from functools import lru_cache

@lru_cache(maxsize=1)
def get_preinitialized_a2v_pipeline():
    """Return cached pipeline to avoid repeated model loading."""
    paths = ModelPaths(root="/models/ltx2")
    return A2VidPipelineTwoStage(
        model_paths=paths,
        distilled_lora=[...],
        spatial_upsampler_path="/models/upsampler",
        device=torch.device("cuda"),
    )

def generate_a2v_video(prompt: str, audio_file: str):
    """Generate video with pre-loaded pipeline."""
    pipeline = get_preinitialized_a2v_pipeline()
    
    video_iter, audio, _ = pipeline(
        prompt=prompt,
        negative_prompt="",
        seed=0,
        height=720,
        width=1280,
        num_frames=60,
        frame_rate=24,
        num_inference_steps=40,
        video_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
        images=[],
        audio_path=audio_file,
    )
    
    # Encode to final video format

    return video_iter, audio

```

Caching the pipeline instance eliminates model loading overhead—critical for server-side inference scenarios.

## Key Source Files for Deep Customization

| File | Description |
|------|-------------|
| [[`packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py) | Core two-stage pipeline and CLI |
| [[`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) | `VideoUpsampler`, `VideoDecoder`, `AudioConditioner` |
| [[`packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py) | Audio decoding, video encoding, HDR utilities |
| [[`packages/ltx-core/src/ltx_core/model/audio_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/audio_vae.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/audio_vae.py) | Audio VAE encoder (`vae_encode_audio`) |
| [[`packages/ltx-core/src/ltx_core/model/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae.py) | Video VAE with tiling support |
| [[`packages/ltx-core/src/ltx_core/components/guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/guiders.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/guiders.py) | `MultiModalGuider` for CFG |
| [[`packages/ltx-core/src/ltx_core/components/noisers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/noisers.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/noisers.py) | `GaussianNoiser` for diffusion |

## Summary

- **`A2VidPipelineTwoStage`** implements coarse-to-fine audio-to-video generation in LTX-2
- **Stage 1** generates half-resolution video with frozen audio conditioning
- **`VideoUpsampler`** bridges stages by spatially scaling latents
- **Stage 2** applies distilled LoRA refinement at full resolution
- **Multi-modal guidance** operates through `MultiModalGuider` with independent CFG scales
- Both CLI and Python APIs support full parameter control including tiling, quantization, and custom LoRA loading

## Frequently Asked Questions

### What audio formats does A2VidPipelineTwoStage support?

The pipeline accepts standard audio formats through `decode_audio_from_file` in [[`media_io.py`](https://github.com/Lightricks/LTX-2/blob/main/media_io.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io.py). WAV and MP3 are explicitly supported, with automatic resampling to the model's expected sample rate. Use `--audio-start-time` and `--audio-max-duration` to extract specific segments.

### How does the spatial upsampler improve video quality?

The `VideoUpsampler` (defined in [[`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py#L99-L110)) implements learned spatial interpolation that preserves temporal consistency better than naive resizing. According to the LTX-2 source, this upsampler is specifically trained to reduce artifacts when doubling spatial resolution between diffusion stages.

### Can I use custom LoRA weights with A2VidPipelineTwoStage?

Yes. Pass additional LoRAs via the `loras` parameter in the constructor or `--lora-path` via CLI. The pipeline composes these with the required `distilled_lora` for stage 2. Each LoRA is wrapped in `LoraPathStrengthAndSDOps` to specify blending strength and selective layer application.

### Why are there two separate guider configurations for video and audio?

The `MultiModalGuiderParams` (defined in [[`a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/a2vid_two_stage.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py#L30-L37)) separates `cfg_scale` (classifier-free guidance strength) from `modality_scale` (cross-modal conditioning weight). This decoupling allows fine-tuning how strongly the audio influences video generation independently from the unconditional guidance scale—critical for balancing temporal synchronization with visual fidelity.