# How the Dub-It Pipeline Works for Audio Rephrasing with Lip Sync in LTX-2

> Discover the Dub-It pipeline in LTX-2 a two stage diffusion system for audio rephrasing with lip sync. Learn how it preserves original audio and maintains perfect sync.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-20

---

**The Dub-It pipeline in LTX-2 is a two-stage diffusion system that replaces a video's visual content while preserving the original audio track and maintaining perfect lip synchronization through frozen audio latents and RoPE-based conditioning.**

This guide explains the complete technical architecture of Lightricks' **Dub-It audio rephrasing pipeline**, from reference audio extraction to joint video-audio diffusion. The pipeline is implemented as `DubItPipeline` in the LTX-2 inference framework, enabling high-quality dubbing without traditional frame-rate constraints.

## Dub-It Pipeline Architecture

The `DubItPipeline` class in [`packages/ltx-pipelines/src/ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/dubit.py) orchestrates a sophisticated two-stage process that keeps audio and video tightly coupled.

### Core Components

| Component | Role | Location |
|-----------|------|----------|
| **`DubItPipeline`** | Main orchestrator loading prompt encoder, conditioners, diffusion, and upsampler | [`ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/dubit.py) lines 67-84 |
| **Audio VAE encoder** | Extracts and encodes reference audio from source video | [`ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/dubit.py) lines 86-92 |
| **IC-LoRA conditioner** | Adds video reference tokens for identity preservation | [`ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/dubit.py) lines 149-180 |
| **Two-stage diffusion** | Joint video-audio generation with frozen audio latents | [`ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/dubit.py) lines 279-326 |
| **Spatial upsampler** | Doubles video resolution in latent space | [`ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/dubit.py) lines 300-312 |

The pipeline operates **without explicit frame-rate or frame-count arguments**—these are inferred directly from the reference video container via `_verify_media_path_args`.

## Audio Conditioning and Lip-Sync Mechanism

### Step 1: Reference Audio Extraction

The pipeline extracts audio from the source video using `decode_audio_from_file`, then encodes it through the audio VAE:

```python

# Internal pipeline flow (simplified)

reference_audio = decode_audio_from_file(reference_video_path)
audio_latent = vae_encode_audio(reference_audio)

```

This creates `AudioConditionByReferenceLatent` (defined in [`ltx_core/conditioning.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/conditioning.py)), a conditioning token that maintains temporal alignment with video frames.

### Step 2: RoPE-Based Temporal Alignment

The `patchify_dubit_audio_reference_latent` function (lines 335-355) performs critical patchification:

- Splits audio latent into **time-aligned tokens**
- Builds a **RoPE position map** linking each audio token to its corresponding video frame
- Preserves phase relationships for lip-sync accuracy

### Step 3: Frozen Audio in Second Stage

The key innovation ensuring lip-sync: the audio latent is marked `frozen=True` for stage 2 diffusion. While the video latent undergoes spatial upsampling, the audio receives **no additional noise**, guaranteeing decoded audio matches the original waveform precisely.

## Command-Line Usage

The `dubit_arg_parser` in [`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py) (lines 80-88) provides a minimal CLI requiring only:

| Flag | Purpose |
|------|---------|
| `--prompt` | Visual description for the generated content |
| `--reference-video` | Source video providing **audio track and frame count** |
| `--spatial-upsampler-path` | Checkpoint for resolution doubling |
| `--lora` | **Dub-It IC-LoRA** model (exactly one required) |
| `--height` / `--width` | Target resolution (multiple of 64) |
| `--reference-strength` | Video conditioning intensity (default 1.0) |

### Basic CLI Example

```bash
python -m ltx_pipelines.dubit \
  --prompt "A chef presenting a new dish" \
  --reference-video /data/original_clip.mp4 \
  --spatial-upsampler-path /models/spatial_upsampler.safetensors \
  --lora /models/dubit_ic_lora.safetensors \
  --output-path /results/dubbed_output.mp4 \
  --height 512 \
  --width 512 \
  --seed 42

```

Execution flow:
1. Parse arguments and validate media paths
2. Load `DubItPipeline` with IC-LoRA weights
3. Extract and encode reference audio
4. Run joint diffusion (stage 1)
5. Upsample video latent, run stage 2 with **frozen audio**
6. Mux final video with original audio into output MP4

## Programmatic Implementation

For custom integrations, instantiate `DubItPipeline` directly:

```python
from ltx_pipelines.dubit import DubItPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Configure model paths

model_paths = ModelPaths(
    transformer="path/to/transformer.safetensors",
    video_vae="path/to/video_vae.safetensors",
    audio_vae="path/to/audio_vae.safetensors",
)

# Initialize pipeline with Dub-It IC-LoRA

pipeline = DubItPipeline(
    model_paths=model_paths,
    spatial_upsampler_path="path/to/spatial_upsampler.safetensors",
    ic_lora=ModelPaths.lora_from_path("path/to/dubit_ic_lora.safetensors"),
    device="cuda",
)

# Generate dubbed video

video, audio, _ = pipeline(
    prompt="A robot delivering a speech",
    seed=123,
    height=512,
    width=512,
    images=[],  # optional image conditioning

    reference_video_path="original.mp4",
    reference_strength=1.0,
)

# Export with preserved audio sync

from ltx_pipelines.utils.media_io import (
    encode_video, get_videostream_metadata, get_video_chunks_number
)

meta = get_videostream_metadata("original.mp4")
encode_video(
    video=video,
    fps=int(meta.fps),
    audio=audio,  # Original reference audio, perfectly synced

    output_path="dubbed_result.mp4",
    video_chunks_number=get_video_chunks_number(meta.frames, None),
)

```

The returned `audio` tensor is **identical** to the decoded reference audio—no generation artifacts or timing drift.

## Key Implementation Files

| File | Content | Link |
|------|---------|------|
| [`ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/dubit.py) | Complete `DubItPipeline` implementation | [source](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/dubit.py) |
| [`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py) | `dubit_arg_parser` CLI definition | [source](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py) |
| [`ltx_pipelines/utils/media_io/__init__.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/media_io/__init__.py) | `decode_audio_from_file`, `encode_video` helpers | [source](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/media_io/__init__.py) |
| [`ltx_core/model/audio_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/audio_vae.py) | Audio VAE encoder/decoder | [source](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/audio_vae.py) |
| [`ltx_core/conditioning.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/conditioning.py) | `AudioConditionByReferenceLatent` definition | [source](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/conditioning.py) |

## Summary

- **Dub-It is a two-stage diffusion pipeline** that regenerates video visuals while freezing audio latents to preserve exact timing
- **RoPE-based conditioning** aligns audio tokens with video frames for natural lip movement
- **No manual frame-rate specification**—timing is inferred from the reference video container
- **IC-LoRA video reference** maintains subject identity across visual changes
- **Spatial upsampling** operates only on video, leaving audio untouched for perfect sync

## Frequently Asked Questions

### How does Dub-It maintain lip synchronization without explicit phoneme alignment?

The pipeline uses **frozen audio latents** and **RoPE position encoding** rather than traditional phoneme detection. By encoding the reference audio through the audio VAE and preventing any diffusion noise in stage 2, the decoded audio matches the original waveform exactly. The `patchify_dubit_audio_reference_latent` function creates temporally-aligned tokens that the video diffusion attends to naturally.

### What makes the Dub-It LoRA different from standard IC-LoRA models?

Dub-It requires **exactly one IC-LoRA checkpoint** specifically trained for audio-guided video generation. This LoRA encodes video reference information through the standard IC-LoRA token mechanism (lines 149-180 in [`dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/dubit.py)), but is optimized for cases where audio timing must be preserved while visual content changes.

### Can I use Dub-It with custom audio instead of extracted reference audio?

According to the source implementation, the pipeline is designed around **reference video conditioning**—the audio is extracted automatically via `decode_audio_from_file`. For custom audio, you would need to modify the `_encode_reference_audio_vae_latent` call or create a video container with your target audio track.

### Why does the pipeline require a spatial upsampler as a separate argument?

The **two-stage architecture** separates generation from resolution enhancement. Stage 1 produces base-resolution latents for both modalities; only the video latent proceeds through `spatial_upsampler_path` (lines 300-312). This design keeps the audio pathway lightweight while allowing flexible video resolution scaling without re-running audio encoding.