# How to Use A2VidPipeline for Audio-to-Video Generation in LTX-2

> Generate synchronized video from audio using A2VidPipeline in LTX-2. Learn to instantiate A2VidPipelineTwoStage and leverage its two-stage diffusion process for audio-to-video creation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-14

---

**Instantiate `A2VidPipelineTwoStage` with model paths, audio file, and generation parameters to produce synchronized video from audio input using a two-stage diffusion process.**

The **A2VidPipeline** in Lightricks' LTX-2 repository provides a production-ready **audio-to-video generation** system that transforms audio tracks into visually coherent videos through a sophisticated two-stage diffusion pipeline. This guide walks through the complete implementation, from architecture to practical usage, based on the actual source code in [`packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py).

## Understanding the Two-Stage Architecture

The `A2VidPipelineTwoStage` class implements a deliberate **coarse-to-fine generation strategy** that separates initial structure creation from high-fidelity refinement.

### Stage 1: Low-Resolution Structure Generation

**Purpose**: Generate base video at half target resolution while preserving audio fidelity through frozen conditioning.

The first stage uses a `GuidedDenoiser` that processes video and audio latents together. Critically, the audio latent is **frozen** with `noise_scale=0.0`, meaning it acts purely as conditioning guidance rather than being noised and denoised. From the source:

```python
video_state, _ = self.stage_1(
    denoiser=GuidedDenoiser(...),
    sigmas=stage_1_sigmas,
    video=ModalitySpec(context=v_context_p, conditionings=stage_1_conditionings),
    audio=ModalitySpec(context=a_context_p, frozen=True, noise_scale=0.0,
                      initial_latent=encoded_audio_latent),
    ...
)

```

*Source: [`/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main//packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py), lines 124-158*

The `GuidedDenoiser` handles multimodal guidance through classifier-free guidance (CFG) and modality-specific scaling, letting the audio strongly influence visual generation patterns.

### Stage 2: Upscaled Refinement with Distilled LoRA

**Purpose**: Double spatial resolution and refine details using a **distilled LoRA** for faster, higher-quality inference.

After upsampling via `VideoUpsampler`, the second stage employs `SimpleDenoiser` with specialized `stage_2_loras`:

```python
video_state, _ = self.stage_2(
    denoiser=SimpleDenoiser(v_context_p, a_context_p),
    sigmas=stage_2_sigmas,
    video=ModalitySpec(..., noise_scale=stage_2_sigmas[0].item(),
                      initial_latent=upscaled_video_latent),
    audio=ModalitySpec(..., frozen=True, initial_latent=encoded_audio_latent),
    ...
)

```

*Source: [`/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main//packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py), lines 187-224*

The **distilled LoRA** enables fewer inference steps while maintaining quality—a key efficiency optimization in the LTX-2 pipeline.

## Core Pipeline Components

### Initialization Flow

The `__init__` method in `A2VidPipelineTwoStage` (lines 61-145) establishes all conditioning and generation modules:

```python
self.prompt_encoder = PromptEncoder(...)
self.image_conditioner = ImageConditioner(...)
self.audio_conditioner = AudioConditioner(...)
self.stage_1 = DiffusionStage.from_checkpoint(...)
self.stage_2 = DiffusionStage.from_checkpoint(...)
self.upsampler = VideoUpsampler(...)
self.video_decoder = VideoDecoder(...)

```

All components share:
- **dtype**: `bfloat16` for memory efficiency
- **device**: Unified CUDA placement
- **scheduler**: `LTX2Scheduler` for consistent noise scheduling

### Audio Conditioning Pipeline

Audio processing follows a clear encoding path in the `__call__` method (lines 95-103):

```python
decoded_audio = decode_audio_from_file(audio_path, self.device, ...)
encoded_audio_latent = self.audio_conditioner(
    lambda enc: vae_encode_audio(decoded_audio, enc, None)
)

```

The `AudioConditioner` (defined in [`/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main//packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py)) wraps the audio VAE and provides consistent latent representations of shape `AudioLatentShape` that remain frozen across both generation stages.

## Command-Line Usage for A2VidPipeline

The fastest way to experiment with **audio-to-video generation** is through the built-in CLI:

```bash
python -m ltx_pipelines.a2vid_two_stage \
    --model-paths /path/to/models \
    --distilled-lora /path/to/distilled_lora.pt \
    --spatial-upsampler-path /path/to/upsampler.pt \
    --prompt "A sunrise over a futuristic city" \
    --negative-prompt "" \
    --seed 42 \
    --height 720 \
    --width 1280 \
    --num-frames 120 \
    --frame-rate 30 \
    --num-inference-steps 50 \
    --audio-path /data/audio/my_track.wav \
    --output-path result.mp4

```

### Essential Audio Parameters

| Parameter | Function | Default |
|-----------|----------|---------|
| `--audio-path` | Source audio file path | **Required** |
| `--audio-start-time` | Seconds to skip at audio start | `0.0` |
| `--audio-max-duration` | Maximum audio length in seconds | Match video duration |

Additional generation flags include `--enhance-prompt` for LLM-based prompt refinement and `--video-cfg-guidance-scale` for adjusting classifier-free guidance strength.

## Programmatic A2VidPipeline Integration

For production systems requiring custom workflows, instantiate the pipeline directly:

```python
from ltx_pipelines.a2vid_two_stage import A2VidPipelineTwoStage
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_core.loader.registry import Registry
import torch

# Prepare model configuration

model_paths = ModelPaths(root="/models")
distilled_lora = [...]  # List[PathStrengthAndSDOps]

spatial_upsampler = "/models/upsampler.pt"
loras = []  # Optional additional LoRAs

# Initialize pipeline

pipeline = A2VidPipelineTwoStage(
    model_paths=model_paths,
    distilled_lora=distilled_lora,
    spatial_upsampler_path=spatial_upsampler,
    loras=loras,
    device=torch.device("cuda"),
    offload_mode=OffloadMode.NONE,
)

# Execute audio-to-video generation

video, audio, tiling_cfg = pipeline(
    prompt="A rainy night in neon Tokyo",
    negative_prompt="low quality",
    seed=1234,
    height=720,
    width=1280,
    num_frames=150,
    frame_rate=30.0,
    num_inference_steps=60,
    video_guider_params=MultiModalGuiderParams(
        cfg_scale=7.5,
        modality_scale=1.0,
    ),
    images=[],  # Empty list = no image conditioning

    audio_path="/data/audio/song.wav",
    audio_start_time=0.0,
    audio_max_duration=None,
)

```

### Saving Generated Output

The pipeline returns raw tensors and audio objects that require encoding:

```python
from ltx_pipelines.utils.media_io.encode import encode_video
from ltx_pipelines.utils.media_io.decode import get_video_chunks_number

tiling_cfg = tiling_cfg or AUTO_TILING
chunks = get_video_chunks_number(num_frames=150, tiling_cfg=tiling_cfg)

encode_video(
    video=video,
    fps=30.0,
    audio=audio,  # Original waveform, not VAE-decoded

    output_path="generated.mp4",
    video_chunks_number=chunks,
    color_space=HDRColorSpace.SDR,
)

```

## Memory Optimization with Tiling

For GPUs unable to hold full-resolution video latents, the **A2VidPipeline** supports automatic tiling:

- Set `tiling_config=AUTO_TILING` in the pipeline call
- The `VideoDecoder` (from [`/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main//packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py)) handles spatial tiling and seamless stitching
- Tiling configuration is returned as `tiling_cfg` for proper encoding alignment

## Key Source Files Reference

| Component | Implementation Location |
|-----------|------------------------|
| `A2VidPipelineTwoStage` | [`/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main//packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py) |
| `PromptEncoder`, `ImageConditioner`, `AudioConditioner`, `DiffusionStage`, `VideoUpsampler`, `VideoDecoder` | [`/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main//packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) |
| `GuidedDenoiser`, `SimpleDenoiser` | [`/packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py`](https://github.com/Lightricks/LTX-2/blob/main//packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py) |
| Audio/video I/O utilities | `/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io/` |
| `LTX2Scheduler` | [`/packages/ltx-core/src/ltx_core/components/schedulers.py`](https://github.com/Lightricks/LTX-2/blob/main//packages/ltx-core/src/ltx_core/components/schedulers.py) |

## Summary

- **A2VidPipeline** implements two-stage audio-to-video generation: Stage 1 creates half-resolution structure with frozen audio guidance, Stage 2 upscales and refines with distilled LoRA
- Audio is encoded once via `AudioConditioner` and kept frozen (`noise_scale=0.0`) across both stages to preserve synchronization fidelity
- The pipeline supports both CLI execution (`python -m ltx_pipelines.a2vid_two_stage`) and Python instantiation
- Memory-constrained deployments benefit from automatic tiling in `VideoDecoder`
- The original audio waveform (not VAE-decoded) is returned and should be used for final encoding to maintain quality

## Frequently Asked Questions

### What audio formats does A2VidPipeline support?

The pipeline uses `decode_audio_from_file` from the media I/O utilities, which handles standard formats including WAV, MP3, and FLAC through underlying `torchaudio` backends. The audio is resampled to the model's expected sampling rate during encoding.

### Why is the audio latent frozen instead of being generated?

Freezing the audio latent with `noise_scale=0.0` ensures **perfect audio fidelity preservation**. Unlike video which becomes coherent through denoising, audio would degrade if subjected to the diffusion process. The frozen latent serves as deterministic conditioning that guides visual generation to match audio rhythm and structure.

### How do image conditionings work with audio input?

The `images` parameter accepts a list of conditioning images that are processed by `ImageConditioner` at each stage's native resolution (half resolution for Stage 1, full resolution for Stage 2). These combine with audio conditioning through the multimodal `GuidedDenoiser`, allowing synchronized visual-audio generation with specific visual references.

### What's the difference between regular LoRAs and distilled LoRA?

Regular LoRAs load into `stage_1` for initial generation, while **distilled LoRA** (passed via `distilled_lora` parameter) specializes `stage_2` for efficient high-quality refinement. The distilled variant enables faster convergence with fewer inference steps, making the second stage practical for production deployment.