How to Use the LipDub Pipeline for Audio-Driven Video Generation in LTX-2

The LipDub pipeline in LTX-2 is a two-stage diffusion system that generates lip-synchronized videos by conditioning on reference audio and visual identity through IC-LoRA, implemented in the LipDubPipeline class at packages/ltx-pipelines/src/ltx_pipelines/lipdub.py.

The LTX-2 LipDub pipeline enables developers to create audio-driven video generation where a subject's lip movements match an input audio track while preserving their visual identity from a reference video. This specialized implementation combines audio conditioning with IC-LoRA (Identity-Conserving LoRA) video reference conditioning to produce temporally coherent, lip-synchronized results. Whether you are building a command-line tool or integrating into a Python application, understanding the pipeline's architecture and API is essential for effective deployment.

Architecture of the LipDub Pipeline

The LipDubPipeline class encapsulates a complete multi-modal generation workflow that marries three conditioning modalities: text prompts, reference video (IC-LoRA), and reference audio.

According to the source code at LipDubPipeline.__init__, the pipeline initializes six core components:

  • Prompt encoder – Encodes text prompts and optional prompt-images into conditioning contexts
  • Image conditioner – Prepares image-conditioned latent patches (maintained for API consistency, though unused in LipDub mode)
  • Audio conditioner – Encodes reference audio into VAE latents using the audio VAE
  • Diffusion stage – Executes the first-stage low-resolution diffusion with combined IC-LoRA video and audio conditioning
  • Spatial upsampler – Doubles the resolution of video latents between diffusion stages
  • Video and audio decoders – Convert final latents back into MP4 video frames and waveform audio

Reference Video and Audio Processing

Before diffusion begins, the pipeline processes the input reference video to extract both visual identity cues and the target audio track.

Frame Alignment and Audio Extraction

The pipeline first snaps the video's frame count to the nearest 8k+1 format (the model's required temporal length) using the _snap_frames_to_8k1 helper function at lines 45-48. Simultaneously, it extracts the audio stream via decode_audio_from_file and encodes it through the audio VAE using vae_encode_audio as implemented in _encode_reference_audio_vae_latent.

Audio Patchification and Position Encoding

The encoded audio latent undergoes patchification through patchify_lipdub_audio_reference_latent (lines 69-86), which uses the AudioPatchifier class to split the latent into one-frame patches. When negative_positions=True, the function shifts position tensors negatively, placing reference audio tokens before the generation window. This positioning allows the model to treat the reference audio as a fixed identity cue rather than part of the generation target.

Generation Flow in LipDubPipeline.__call__

The generation process follows a strict two-stage workflow implemented in the __call__ method (lines 45-62 and 200-207):

  1. Metadata extraction – Determines frame count and frame-rate from the reference video
  2. Prompt encoding – Generates video_context and audio_context tensors via the prompt encoder
  3. Stage-1 conditioning – Combines image latents, IC-LoRA video reference, and audio reference latents
  4. Stage-1 diffusion – Runs the low-resolution diffusion using SimpleDenoiser with the combined conditioning
  5. Spatial upsampling – Increases resolution via self.upsampler
  6. Stage-2 diffusion – Refines the up-sampled video while keeping the audio latent frozen
  7. Decoding – Renders final output via self.video_decoder and self.audio_decoder

Command-Line Usage

The lipdub_arg_parser function at [packages/ltx-pipelines/src/ltx_pipelines/utils/args.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py#L69-L78) builds a CLI that requires exactly one IC-LoRA checkpoint for identity injection.

python -m ltx_pipelines.lipdub \
  --distilled-checkpoint-path /path/to/distilled.safetensors \
  --spatial-upsampler-path   /path/to/spatial_upsampler.safetensors \
  --gemma-root              /path/to/gemma_root \
  --lora                    /path/to/ic_lora.safetensors 1.0 \
  --prompt                  "A talking parrot in a jungle" \
  --seed                    42 \
  --height                  512 \
  --width                   512 \
  --reference-video         /path/to/source_video.mp4 \
  --output-path             /tmp/parrot_lipdub.mp4

Key parameters:

  • --distilled-checkpoint-path – Base LTX-2 distilled model weights
  • --spatial-upsampler-path – Model checkpoint for resolution doubling
  • --lora – IC-LoRA checkpoint with strength (exactly one required for identity preservation)
  • --reference-video – Source video providing both audio track and visual identity

The CLI automatically derives frame count and frame-rate from the reference video, snaps the count to the nearest 8k+1, and outputs an MP4 at the specified resolution.

Programmatic Usage

For integration into larger systems, instantiate LipDubPipeline directly and bypass the CLI argument parsing.

from ltx_pipelines.lipdub import LipDubPipeline
from ltx_pipelines.utils.media_io import encode_video

# Initialize pipeline with required checkpoints

pipeline = LipDubPipeline(
    distilled_checkpoint_path="distilled.safetensors",
    spatial_upsampler_path="spatial_upsampler.safetensors",
    gemma_root="gemma_root",
    ic_lora=("ic_lora.safetensors", 1.0),  # Tuple of (path, strength)

    quantization="int8",                    # Optional: int8, int4, or None

    compilation_config=None,                # Optional: torch.compile config

    offload_mode="model",                   # Optional: cpu offloading strategy

)

# Generate video and audio tensors

video_tensor, audio_tensor = pipeline(
    prompt="A singing robot",
    seed=123,
    height=512,
    width=512,
    images=[],                              # Empty for pure audio-driven mode

    reference_video_path="source.mp4",
    reference_strength=1.0,
)

# Encode to MP4

encode_video(
    video=video_tensor,
    fps=30,                                 # Match reference video fps

    audio=audio_tensor,
    output_path="output_lipdub.mp4",
    video_chunks_number=1,
)

Implementation notes:

  • Pass images=[] when using only audio and video reference conditioning
  • The reference_strength parameter controls how strongly the IC-LoRA identity influences the output
  • Frame count inference happens automatically inside __call__ based on the reference video duration

Summary

  • The LipDub pipeline requires exactly one IC-LoRA checkpoint to preserve the reference video's visual identity during generation
  • Input videos are automatically processed to 8k+1 frame lengths to satisfy model architecture constraints
  • Audio conditioning uses negative position encoding to separate reference identity audio from the generation target
  • The two-stage diffusion process first generates low-resolution video aligned to audio, then upsamples and refines while freezing audio latents
  • Both CLI and programmatic APIs are available in packages/ltx-pipelines/src/ltx_pipelines/lipdub.py

Frequently Asked Questions

What is the 8k+1 frame requirement in the LipDub pipeline?

The LTX-2 model architecture requires input sequences where the frame count follows the formula 8k+1 (where k is an integer). The _snap_frames_to_8k1 function automatically rounds your reference video's frame count to the nearest valid value, ensuring compatibility with the model's temporal attention mechanisms without manual intervention.

Why does the LipDub pipeline require an IC-LoRA checkpoint?

The IC-LoRA (Identity-Conserving LoRA) is required because the pipeline must preserve the visual identity of the subject from the reference video while changing only the lip movements to match new audio. According to the main function implementation at lines 91-105, the CLI validates that exactly one LoRA is supplied, as this provides the identity injection mechanism that keeps the subject's facial features consistent across the generated frames.

Can I use the LipDub pipeline with custom audio instead of extracting from the reference video?

Currently, the LipDubPipeline class in lipdub.py extracts audio directly from the reference_video_path parameter using decode_audio_from_file. To use custom audio, you would need to either create a video file containing your target audio track, or modify the _encode_reference_audio_vae_latent method to accept an independent audio file path instead of extracting from the video container.

How does the pipeline handle memory constraints during generation?

The pipeline supports multiple offloading strategies via the offload_mode parameter (including "model" and "sequential") and quantization options via quantization (such as "int8" or "int4"). These features, configured during LipDubPipeline initialization, allow the pipeline to run on GPUs with limited VRAM by offloading unused layers to CPU or compressing model weights, respectively.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →