# How to Use the LipDub Pipeline for Audio-Driven Video Generation in LTX-2

> Generate lip-synchronized videos with the LTX-2 LipDub pipeline. Learn how to use this audio-driven video generation system for stunning visual results.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-06-20

---

**The LipDub pipeline in LTX-2 is a two-stage diffusion system that generates lip-synchronized videos by conditioning on reference audio and visual identity through IC-LoRA, implemented in the `LipDubPipeline` class at [`packages/ltx-pipelines/src/ltx_pipelines/lipdub.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/lipdub.py).**

The LTX-2 LipDub pipeline enables developers to create audio-driven video generation where a subject's lip movements match an input audio track while preserving their visual identity from a reference video. This specialized implementation combines audio conditioning with IC-LoRA (Identity-Conserving LoRA) video reference conditioning to produce temporally coherent, lip-synchronized results. Whether you are building a command-line tool or integrating into a Python application, understanding the pipeline's architecture and API is essential for effective deployment.

## Architecture of the LipDub Pipeline

The `LipDubPipeline` class encapsulates a complete multi-modal generation workflow that marries three conditioning modalities: **text prompts**, **reference video (IC-LoRA)**, and **reference audio**.

According to the source code at [`LipDubPipeline.__init__`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/lipdub.py#L51-L66), the pipeline initializes six core components:

- **Prompt encoder** – Encodes text prompts and optional prompt-images into conditioning contexts
- **Image conditioner** – Prepares image-conditioned latent patches (maintained for API consistency, though unused in LipDub mode)
- **Audio conditioner** – Encodes reference audio into VAE latents using the audio VAE
- **Diffusion stage** – Executes the first-stage low-resolution diffusion with combined IC-LoRA video and audio conditioning
- **Spatial upsampler** – Doubles the resolution of video latents between diffusion stages
- **Video and audio decoders** – Convert final latents back into MP4 video frames and waveform audio

## Reference Video and Audio Processing

Before diffusion begins, the pipeline processes the input reference video to extract both visual identity cues and the target audio track.

### Frame Alignment and Audio Extraction

The pipeline first snaps the video's frame count to the nearest **8k+1** format (the model's required temporal length) using the `_snap_frames_to_8k1` helper function at lines 45-48. Simultaneously, it extracts the audio stream via `decode_audio_from_file` and encodes it through the audio VAE using `vae_encode_audio` as implemented in [`_encode_reference_audio_vae_latent`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/lipdub.py#L38-L44).

### Audio Patchification and Position Encoding

The encoded audio latent undergoes **patchification** through `patchify_lipdub_audio_reference_latent` (lines 69-86), which uses the `AudioPatchifier` class to split the latent into one-frame patches. When `negative_positions=True`, the function shifts position tensors negatively, placing reference audio tokens *before* the generation window. This positioning allows the model to treat the reference audio as a fixed identity cue rather than part of the generation target.

## Generation Flow in `LipDubPipeline.__call__`

The generation process follows a strict two-stage workflow implemented in the `__call__` method (lines 45-62 and 200-207):

1. **Metadata extraction** – Determines frame count and frame-rate from the reference video
2. **Prompt encoding** – Generates `video_context` and `audio_context` tensors via the prompt encoder
3. **Stage-1 conditioning** – Combines image latents, IC-LoRA video reference, and audio reference latents
4. **Stage-1 diffusion** – Runs the low-resolution diffusion using `SimpleDenoiser` with the combined conditioning
5. **Spatial upsampling** – Increases resolution via `self.upsampler`
6. **Stage-2 diffusion** – Refines the up-sampled video while keeping the audio latent frozen
7. **Decoding** – Renders final output via `self.video_decoder` and `self.audio_decoder`

## Command-Line Usage

The `lipdub_arg_parser` function at [[`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py#L69-L78) builds a CLI that requires exactly one IC-LoRA checkpoint for identity injection.

```bash
python -m ltx_pipelines.lipdub \
  --distilled-checkpoint-path /path/to/distilled.safetensors \
  --spatial-upsampler-path   /path/to/spatial_upsampler.safetensors \
  --gemma-root              /path/to/gemma_root \
  --lora                    /path/to/ic_lora.safetensors 1.0 \
  --prompt                  "A talking parrot in a jungle" \
  --seed                    42 \
  --height                  512 \
  --width                   512 \
  --reference-video         /path/to/source_video.mp4 \
  --output-path             /tmp/parrot_lipdub.mp4

```

**Key parameters:**
- `--distilled-checkpoint-path` – Base LTX-2 distilled model weights
- `--spatial-upsampler-path` – Model checkpoint for resolution doubling
- `--lora` – **IC-LoRA** checkpoint with strength (exactly one required for identity preservation)
- `--reference-video` – Source video providing both audio track and visual identity

The CLI automatically derives frame count and frame-rate from the reference video, snaps the count to the nearest 8k+1, and outputs an MP4 at the specified resolution.

## Programmatic Usage

For integration into larger systems, instantiate `LipDubPipeline` directly and bypass the CLI argument parsing.

```python
from ltx_pipelines.lipdub import LipDubPipeline
from ltx_pipelines.utils.media_io import encode_video

# Initialize pipeline with required checkpoints

pipeline = LipDubPipeline(
    distilled_checkpoint_path="distilled.safetensors",
    spatial_upsampler_path="spatial_upsampler.safetensors",
    gemma_root="gemma_root",
    ic_lora=("ic_lora.safetensors", 1.0),  # Tuple of (path, strength)

    quantization="int8",                    # Optional: int8, int4, or None

    compilation_config=None,                # Optional: torch.compile config

    offload_mode="model",                   # Optional: cpu offloading strategy

)

# Generate video and audio tensors

video_tensor, audio_tensor = pipeline(
    prompt="A singing robot",
    seed=123,
    height=512,
    width=512,
    images=[],                              # Empty for pure audio-driven mode

    reference_video_path="source.mp4",
    reference_strength=1.0,
)

# Encode to MP4

encode_video(
    video=video_tensor,
    fps=30,                                 # Match reference video fps

    audio=audio_tensor,
    output_path="output_lipdub.mp4",
    video_chunks_number=1,
)

```

**Implementation notes:**
- Pass `images=[]` when using only audio and video reference conditioning
- The `reference_strength` parameter controls how strongly the IC-LoRA identity influences the output
- Frame count inference happens automatically inside `__call__` based on the reference video duration

## Summary

- The LipDub pipeline requires **exactly one IC-LoRA** checkpoint to preserve the reference video's visual identity during generation
- Input videos are automatically processed to **8k+1 frame lengths** to satisfy model architecture constraints
- Audio conditioning uses **negative position encoding** to separate reference identity audio from the generation target
- The two-stage diffusion process first generates low-resolution video aligned to audio, then upsamples and refines while freezing audio latents
- Both CLI and programmatic APIs are available in [`packages/ltx-pipelines/src/ltx_pipelines/lipdub.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/lipdub.py)

## Frequently Asked Questions

### What is the 8k+1 frame requirement in the LipDub pipeline?

The LTX-2 model architecture requires input sequences where the frame count follows the formula `8k+1` (where k is an integer). The `_snap_frames_to_8k1` function automatically rounds your reference video's frame count to the nearest valid value, ensuring compatibility with the model's temporal attention mechanisms without manual intervention.

### Why does the LipDub pipeline require an IC-LoRA checkpoint?

The IC-LoRA (Identity-Conserving LoRA) is required because the pipeline must preserve the visual identity of the subject from the reference video while changing only the lip movements to match new audio. According to the `main` function implementation at lines 91-105, the CLI validates that exactly one LoRA is supplied, as this provides the identity injection mechanism that keeps the subject's facial features consistent across the generated frames.

### Can I use the LipDub pipeline with custom audio instead of extracting from the reference video?

Currently, the `LipDubPipeline` class in [`lipdub.py`](https://github.com/Lightricks/LTX-2/blob/main/lipdub.py) extracts audio directly from the `reference_video_path` parameter using `decode_audio_from_file`. To use custom audio, you would need to either create a video file containing your target audio track, or modify the `_encode_reference_audio_vae_latent` method to accept an independent audio file path instead of extracting from the video container.

### How does the pipeline handle memory constraints during generation?

The pipeline supports multiple offloading strategies via the `offload_mode` parameter (including `"model"` and `"sequential"`) and quantization options via `quantization` (such as `"int8"` or `"int4"`). These features, configured during `LipDubPipeline` initialization, allow the pipeline to run on GPUs with limited VRAM by offloading unused layers to CPU or compressing model weights, respectively.