How the A2Vid Pipeline Leverages Audio Conditioning for Video Generation in LTX-2

The A2VidPipelineTwoStage converts input audio into a frozen latent representation that guides both stages of diffusion without being denoised, enabling cross-modal video generation synchronized to acoustic cues.

Audio conditioning in video generation requires precise alignment between temporal sound patterns and visual motion. The A2VidPipelineTwoStage class in Lightricks' LTX-2 repository implements this through a specialized two-stage diffusion process where audio serves as an immutable context rather than a generative target. This article examines the complete audio conditioning pipeline from waveform decoding through final video synthesis.

Audio Encoding: From Waveform to Latent Representation

The pipeline begins by transforming raw audio into a compressed latent space suitable for transformer attention. This happens in packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py through a coordinated sequence of decoding and encoding operations.

Decoding and VAE Encoding

The decode_audio_from_file function loads the audio waveform, which is then processed by vae_encode_audio imported from packages/ltx-core/model/audio_vae.py:


# From packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py lines 2000-2003

audio_latent = self.vae_encode_audio(decoded_audio)

# audio_latent is then trimmed to match expected frames

audio_latent = audio_latent[:, :, :audio_shape.frames, :, :]

The resulting AudioLatentShape tensor captures temporal structure at a compressed frame rate, enabling efficient attention during diffusion. The latent is explicitly truncated to audio_shape.frames to ensure dimensional alignment with the video generation timeline.

AudioConditioner: Encapsulating Audio State

The pipeline wraps audio processing in an AudioConditioner context manager, defined in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py. This block handles device placement, dtype conversion, and conditioning metadata registration:


# From packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py lines 96-100

self.audio_conditioner = AudioConditioner(
    audio_encoder_path=self.model_paths.audio_vae_path,
    device=self.device,
    dtype=self.dtype,
)

When invoked as a context manager, the AudioConditioner executes the encoding pipeline and yields the prepared latent for downstream consumption.

Frozen Modality: Preserving Audio Integrity During Diffusion

The core innovation of audio conditioning in A2Vid is treating audio as a frozen modality—a context stream that participates in cross-attention but never receives noise or gradient updates.

ModalitySpec Configuration

Both diffusion stages receive identical audio specifications with frozen=True and noise_scale=0.0:

Stage 1 audio specification (lines 251-256):

ModalitySpec(
    latent=audio_latent,
    frozen=True,           # Never denoised

    noise_scale=0.0,       # Zero noise injection

    latent_mean=audio_stats["latent_mean"],
    latent_std=audio_stats["latent_std"],
)

Stage 2 audio specification (lines 91-96):

ModalitySpec(
    latent=audio_latent_upsampled,
    frozen=True,
    noise_scale=0.0,
    latent_mean=audio_stats["latent_mean"],
    latent_std=audio_stats["latent_std"],
)

These specifications ensure the audio latent remains bitwise identical from the initial VAE encoding through both diffusion stages. The transformer can attend to audio tokens via cross-modal attention, but the audio representation itself never drifts.

Cross-Modal Guidance: Aligning Video to Audio Context

The frozen audio latent propagates through the denoising infrastructure to influence video generation dynamics. Two denoiser implementations handle this differently across stages.

Stage 1: GuidedDenoiser with MultiModalGuider

The first stage uses GuidedDenoiser (from packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py) which constructs a MultiModalGuider. This guider receives the audio context (a_context_p) alongside visual conditioning:


# Conceptual flow from GuidedDenoiser implementation

multi_modal_guider = MultiModalGuider(
    video_context=v_context_p,
    audio_context=a_context_p,    # Frozen audio conditioning

    cfg_scale=params.cfg_scale,
    modality_scale=params.modality_scale,
)

The modality_scale parameter (set to 3.0 in typical configurations) controls the strength of audio-video cross-attention, allowing users to tune how aggressively motion follows rhythmic or timbral audio features.

Stage 2: SimpleDenoiser with Preserved Context

The second stage employs SimpleDenoiser without classifier-free guidance, but maintains the same frozen audio latent. This preserves audio-video alignment during spatial upsampling without introducing guidance artifacts:


# SimpleDenoiser receives audio context directly

denoised = self.simple_denoiser(
    noisy_latent,
    timestep,
    context={
        "video": v_context_p,
        "audio": a_context_p,    # Unchanged from Stage 1

    },
)

Complete Inference Example

The following runnable example demonstrates audio-conditioned video generation with explicit modality scaling:

from ltx_pipelines import A2VidPipelineTwoStage
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines.utils.media_io.encode import encode_video
from ltx_pipelines.utils.enums import HDRColorSpace

# 1. Load model paths from monolithic checkpoint

model_paths = ModelPaths.from_monolith()

# 2. Instantiate two-stage pipeline

pipeline = A2VidPipelineTwoStage(
    model_paths=model_paths,
    distilled_lora=[],                     # (path, strength, ops) tuples

    spatial_upsampler_path="spatial_upsampler.pt",
    loras=[],
)

# 3. Generate video with audio conditioning

video, audio, tiling_cfg = pipeline(
    prompt="A dancer moving to the rhythm",
    negative_prompt="static, blurry, low quality",
    seed=42,
    height=512,
    width=768,
    num_frames=120,
    frame_rate=30.0,
    num_inference_steps=30,
    video_guider_params=MultiModalGuiderParams(
        cfg_scale=3.0,
        modality_scale=3.0,               # Audio-video cross-attention strength

    ),
    images=[],                             # Optional visual conditioning

    audio_path="input_music.wav",          # Audio conditioning source

)

# 4. Encode with original audio preserved

encode_video(
    video=video,
    fps=30.0,
    audio=audio,                           # Original decoded waveform

    output_path="audio_synced_output.mp4",
    video_chunks_number=1,
    color_space=HDRColorSpace.SDR,
)

Audio Fidelity Preservation

A critical design decision in A2VidPipelineTwoStage is returning the original decoded audio rather than the VAE-reconstructed version. From lines 3001-3004:


# Return original audio for maximum fidelity

return Video(video_frames), Audio(original_audio), tiling_config

This prevents generative artifacts from the audio VAE decoder from degrading output quality. The audio latent serves only as an internal conditioning signal; users receive the pristine input waveform synchronized to their generated video.

Key Architectural Files

File Responsibility
packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py Core pipeline implementation with audio encoding, two-stage diffusion, and output handling
packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py AudioConditioner and other conditioning block definitions
packages/ltx-core/model/audio_vae.py vae_encode_audio and audio VAE architecture
packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py GuidedDenoiser, SimpleDenoiser, and MultiModalGuider implementations
packages/ltx-pipelines/docs/conditioning.md Documentation on multimodal conditioning patterns

Summary

  • Audio encoding: Raw waveforms pass through decode_audio_from_file → vae_encode_audio to produce shaped latents in packages/ltx-core/model/audio_vae.py
  • Frozen modality: ModalitySpec(frozen=True, noise_scale=0.0) prevents any noise injection or latent drift during diffusion
  • Cross-modal guidance: MultiModalGuider in GuidedDenoiser applies audio context to video generation via modality_scale parameter
  • Fidelity preservation: Original decoded audio is returned, bypassing VAE reconstruction artifacts
  • Two-stage consistency: Both stages receive identical frozen audio conditioning for coherent temporal alignment

Frequently Asked Questions

How does A2VidPipelineTwoStage prevent audio quality degradation during generation?

The pipeline preserves audio fidelity by returning the original decoded waveform rather than the VAE-reconstructed version. While the audio VAE latent guides video generation internally, users receive the pristine input audio file unchanged, avoiding any compression artifacts from the generative audio decoder.

What does frozen=True accomplish in the audio ModalitySpec?

Setting frozen=True with noise_scale=0.0 instructs the diffusion loop to skip all noise injection and denoising operations for the audio stream. The audio latent remains exactly as output from the VAE encoder, serving as stable cross-modal context that the video transformer can attend to without perturbation.

Can I control how strongly audio influences video motion?

Yes. The MultiModalGuiderParams.modality_scale parameter (typically set to 3.0) directly controls audio-video cross-attention strength. Higher values amplify rhythmic and timbral audio cues in visual dynamics; lower values produce more independent video motion. This is passed to GuidedDenoiser in Stage 1, while Stage 2 preserves the established alignment without additional scaling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →