# How to Use the Dub-It Pipeline for Audio Dubbing with Lip Sync in LTX-2

> Master audio dubbing with lip sync in LTX-2 using the Dub-It pipeline. Learn how this innovative process ensures perfect synchronization for your videos.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-14

---

**The Dub-It pipeline in LTX-2 replaces a video's audio track while maintaining perfect lip synchronization through a two-stage diffusion process that freezes the original audio latent.**

The **Dub-It pipeline** is a specialized inference pipeline in Lightricks' LTX-2 that enables high-quality audio dubbing with automatic lip-sync. Unlike standard video generation, it preserves the exact timing and content of a reference audio track while regenerating visual frames to match. This guide explains how to run Dub-It from both command-line and Python interfaces according to the official LTX-2 source code.

## Architecture of the Dub-It Pipeline

The pipeline is implemented in [`packages/ltx-pipelines/src/ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/dubit.py) and consists of five core components working in sequence:

### Core Components

- **`DubItPipeline` class** — Orchestrates the entire dubbing process. It loads the prompt encoder, video/audio conditioners, diffusion stage, and spatial up-sampler. Located at lines 67-84 in [`dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/dubit.py).

- **Audio conditioning module** — Extracts reference audio from the input video, encodes it via the audio VAE, and constructs a RoPE-based reference token using `AudioConditionByReferenceLatent`. See lines 86-92.

- **Video conditioning module** — Builds image conditionings and injects IC-LoRA video reference tokens for visual fidelity. Implemented at lines 149-180.

- **Two-stage diffusion** — Stage 1 runs joint diffusion on both video and audio latents. The audio latent is marked `frozen=True` for Stage 2, ensuring timing preservation. Found at lines 279-326.

- **Spatial up-sampler** — Doubles video resolution in latent space between stages (lines 300-312), while audio remains untouched.

### Lip Sync Guarantee Mechanism

The pipeline guarantees lip sync through four specific technical steps:

1. **Audio latent extraction** — `self._encode_reference_audio_vae_latent` reads the reference video's audio track via `decode_audio_from_file` and encodes it with `vae_encode_audio`.

2. **RoPE position alignment** — `patchify_dubit_audio_reference_latent` (lines 335-355) creates patchified latents with frame-aligned position maps.

3. **Frozen audio refinement** — During Stage 2 diffusion, the audio latent receives zero additional noise—only the video latent is denoised at higher resolution.

4. **Direct audio decoding** — The final audio is decoded directly from the frozen latent without modification, then muxed with the regenerated video.

## Command-Line Usage

The simplest way to run Dub-It is through the dedicated CLI. Argument parsing is handled by `dubit_arg_parser` in [`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py) (lines 80-88).

### Required Arguments

| Flag | Purpose |
|------|---------|
| `--prompt` | Text description of desired visual content |
| `--reference-video` | Source video providing **both audio track and frame count** |
| `--spatial-upsampler-path` | Path to the up-sampler checkpoint (Stage 2) |
| `--lora` | Path to **Dub-It IC-LoRA** model (exactly one required) |
| `--output-path` | Destination MP4 file path |

### Optional Arguments

| Flag | Default | Description |
|------|---------|-------------|
| `--height` / `--width` | — | Target resolution (must be divisible by 64) |
| `--reference-strength` | 1.0 | Strength of video reference conditioning |
| `--seed` | random | Reproducibility seed |
| `--enhance-prompt` | False | LLM-based prompt enhancement |
| `--compile` | False | PyTorch compilation for speed |

### CLI Example

```bash
python -m ltx_pipelines.dubit \
  --prompt "A chef presenting a new dish" \
  --reference-video /data/original_clip.mp4 \
  --spatial-upsampler-path /models/spatial_upsampler.safetensors \
  --lora /models/dubit_ic_lora.safetensors \
  --output-path /results/dubbed_output.mp4 \
  --height 512 \
  --width 512 \
  --seed 42

```

No frame-rate or frame-count arguments are needed—these are inferred automatically from the reference video container via `_verify_media_path_args`.

## Programmatic Usage

For integration into larger workflows, instantiate `DubItPipeline` directly:

### Python Example

```python
from ltx_pipelines.dubit import DubItPipeline, patchify_dubit_audio_reference_latent
from ltx_pipelines.utils.model_paths import ModelPaths

# Step 1: Configure model paths

model_paths = ModelPaths(
    transformer="path/to/transformer.safetensors",
    video_vae="path/to/video_vae.safetensors",
    audio_vae="path/to/audio_vae.safetensors",
)

# Step 2: Initialize pipeline with spatial upsampler and IC-LoRA

pipeline = DubItPipeline(
    model_paths=model_paths,
    spatial_upsampler_path="path/to/spatial_upsampler.safetensors",
    ic_lora=ModelPaths.lora_from_path("path/to/dubit_ic_lora.safetensors"),
    device="cuda",
)

# Step 3: Run inference

video, audio, _ = pipeline(
    prompt="A robot delivering a speech",
    seed=123,
    height=512,
    width=512,
    images=[],                         # No additional image conditioning

    reference_video_path="original.mp4",
    reference_strength=1.0,
)

# Step 4: Save output using media I/O helpers

from ltx_pipelines.utils.media_io import (
    encode_video,
    get_videostream_metadata,
    get_video_chunks_number,
)

meta = get_videostream_metadata("original.mp4")
encode_video(
    video=video,
    fps=int(meta.fps),
    audio=audio,
    output_path="dubbed_result.mp4",
    video_chunks_number=get_video_chunks_number(meta.frames, None),
)

```

**Critical implementation detail:** The `audio` tensor returned by `pipeline.__call__` is decoded directly from the frozen reference latent—guaranteeing bit-identical timing to the original source.

## Key Implementation Files

| File | Contents | Line References |
|------|----------|-----------------|
| [`ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/dubit.py) | Full pipeline: audio extraction, joint diffusion, up-sampling | 67-84 (init), 86-92 (audio cond), 149-180 (video cond), 279-326 (diffusion), 300-312 (upsampler), 335-355 (patchify) |
| [`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py) | CLI argument parser with Dub-It-specific validation | 80-88 |
| [`ltx_pipelines/utils/media_io/__init__.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/media_io/__init__.py) | Audio decoding (`decode_audio_from_file`) and video encoding (`encode_video`) | — |
| [`ltx_core/model/audio_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/audio_vae.py) | Audio VAE encoder/decoder for latent extraction | — |
| [`ltx_core/conditioning.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/conditioning.py) | `AudioConditionByReferenceLatent` class definition | — |

## Summary

- **Dub-It** is a two-stage diffusion pipeline specifically designed for audio-dubbed video generation with guaranteed lip sync.
- The **reference video** provides both the audio track to preserve and the frame count/timing to match—no manual alignment needed.
- **Audio latents are frozen** during Stage 2, ensuring decoded audio matches the original speech waveform exactly.
- **IC-LoRA conditioning** maintains visual identity from the reference while allowing prompt-driven visual changes.
- Both **CLI and Python APIs** are fully supported in `ltx_pipelines.dubit`.

## Frequently Asked Questions

### What makes Dub-It different from standard video generation pipelines?

Standard pipelines generate both audio and video from scratch or condition on separate audio inputs. Dub-It specifically **extracts and freezes the reference audio latent** from an existing video, ensuring the generated lip movements match the original speech timing with frame-level precision. This is implemented through `AudioConditionByReferenceLatent` and the `frozen=True` flag in Stage 2 diffusion.

### Why does Dub-It require a spatial up-sampler?

The pipeline runs diffusion at a lower resolution in Stage 1 for efficiency, then uses the spatial up-sampler to double the video resolution in Stage 2. The audio latent bypasses this up-sampling entirely—only video latents are processed—maintaining the original audio quality while improving visual fidelity.

### Can I use multiple LoRAs with Dub-It?

No. The `dubit_arg_parser` explicitly enforces exactly one LoRA path via `--lora`. This LoRA must be a **Dub-It IC-LoRA** trained for video reference conditioning. The constraint is validated in [`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py) to ensure proper audio-visual alignment.

### How does the pipeline handle frame rate mismatches?

It doesn't need to. Frame rate and frame count are **inferred directly from the reference video container** via `_verify_media_path_args`. The RoPE position embeddings in `patchify_dubit_audio_reference_latent` automatically align audio tokens to video frames based on these extracted metadata, eliminating manual configuration.