# How to Use the LTX-2 DubItPipeline for Re-Voicing and Lip-Sync Synchronization

> Master LTX-2 DubItPipeline for seamless re-voicing and lip-sync synchronization. Preserve speaker identity and facial expressions with this powerful two-stage diffusion process.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-15

---

**The LTX-2 DubItPipeline rewrites spoken content while preserving the original speaker's identity, lip movements, and facial expressions using a two-stage diffusion process with a single IC-LoRA applied to both stages.**

The **DubItPipeline** in Lightricks/LTX-2 enables developers to generate synchronized video and audio from a reference clip and new text prompt. This pipeline leverages the distilled LTX-2.5 model and specialized IC-LoRA conditioning to achieve state-of-the-art re-voicing results. Below is a comprehensive guide to its architecture, API, and practical usage.

## How DubItPipeline Works

The pipeline operates as a **two-stage generation engine** that processes reference video frames and audio simultaneously. Both visual and audio cues from the reference drive the generation, with the audio latent appended as frozen reference tokens to ensure lip-sync alignment.

### Core Architecture Components

The pipeline initializes six primary components in [`packages/ltx-pipelines/src/ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/dubit.py):

- **PromptEncoder** – encodes the input text prompt
- **ImageConditioner** – handles visual conditioning signals
- **AudioConditioner** – processes audio reference latents
- **DiffusionStage** – runs the noise-to-latent diffusion process
- **VideoUpsampler** – spatially upsamples low-resolution video latents
- **VideoDecoder / AudioDecoder** – VAE-based decoders for final output

All components operate in **bfloat16** precision on the same device. The IC-LoRA is loaded via `LoraPathStrengthAndSDOps`, with its down-scale factor read from LoRA metadata.

### Two-Stage Conditioning Process

**Stage 1: Initial Generation**

1. Image conditioning is built with `combined_image_conditionings`
2. Reference video frames are injected via `append_ic_lora_reference_video_conditionings` — this adds IC-LoRA tokens aligned to the reference
3. Audio conditioning extracts the reference audio stream, encodes it with `vae_encode_audio`, and patchifies via `patchify_dubit_audio_reference_latent` (adding negative RoPE positions for reference)
4. Diffusion runs with `DISTILLED_SIGMAS` to produce low-resolution video and audio latents

**Stage 2: Upscaling and Refinement**

1. The video latent is upsampled by the spatial upsampler
2. Video conditionings are recomputed for target resolution
3. Audio conditionings reuse the frozen reference latent from stage 1
4. Final diffusion produces the high-resolution output

The pipeline returns `(decoded_video, decoded_audio, tiling_config)`.

## Running DubItPipeline from the Command Line

The fastest way to use the pipeline is via the `ltx_pipelines.dubit` CLI entry point. No frame count or frame rate parameters are required — these are inferred automatically from the reference video.

```bash
uv run python -m ltx_pipelines.dubit \
    --transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
    --text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
    --video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
    --audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
    --spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
    --lora Lightricks/LTX-2.3-22b-IC-LoRA-DubIt \
    --reference-video path/to/original_clip.mp4 \
    --prompt "Hello, I'm excited to announce the new feature!" \
    --seed 12345 \
    --height 720 \
    --width 1280 \
    --output-path dubit_result.mp4

```

**Key CLI behavior:** The `--reference-video` flag provides both visual frames and audio VAE latents. The pipeline automatically derives frame count and FPS from the source file.

## Using the DubItPipeline Python API

For custom workflows, instantiate `DubItPipeline` directly and call it with your parameters.

```python
from ltx_pipelines.dubit import DubItPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Build paths to required model files

model_paths = ModelPaths(
    transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
    text_encoder="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
    video_vae="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
    audio_vae="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
    spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)

pipeline = DubItPipeline(
    model_paths=model_paths,
    spatial_upsampler_path=model_paths.spatial_upsampler,
    ic_lora=ModelPaths.lora_path("Lightricks/LTX-2.3-22b-IC-LoRA-DubIt"),
)

video, audio, _ = pipeline(
    prompt="Welcome to the next generation of video synthesis!",
    seed=42,
    height=720,
    width=1280,
    images=[],
    reference_video_path="reference.mp4",
)

# Save the result

from ltx_pipelines.utils.media_io import encode_video
encode_video(video, fps=30, audio=audio, output_path="dubbed.mp4")

```

The `pipeline()` call returns decoded video frames, decoded audio samples, and the tiling configuration used during VAE decoding.

## Customizing Conditioning Parameters

To modify reference strength or inject additional image conditionings, use the internal `_create_stage_conditionings` method or replicate its logic in your code.

```python

# Example: increase reference strength to 1.5

custom_conditionings = pipeline._create_stage_conditionings(
    images=[],
    reference_video_path="reference.mp4",
    reference_strength=1.5,
    height=360,
    width=640,
    num_frames=121,
    video_encoder=your_video_encoder,
    encode_tiling=your_tiling_cfg,
)

```

This returns a list of conditioning objects that can be passed directly to `pipeline.stage` for fine-grained control over the diffusion process.

## Key Source Files and Functions

Understanding these files helps when debugging or extending the pipeline:

| File | Purpose |
|------|---------|
| [`packages/ltx-pipelines/src/ltx_pipelines/dubit.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/dubit.py) | Main `DubItPipeline` implementation |
| [`packages/ltx-core/src/ltx_core/conditioning.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/conditioning.py) | `AudioConditionByReferenceLatent` and conditioning helpers |
| [`packages/ltx-core/src/ltx_core/model/audio_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/audio_vae.py) | `vae_encode_audio` and audio latent operations |
| [`packages/ltx-core/src/ltx_core/model/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae.py) | Video VAE encoder, decoder, and tiling utilities |
| [`packages/ltx-pipelines/utils/media_io.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/utils/media_io.py) | `encode_video` and media I/O functions |

Critical functions to know:

- `append_ic_lora_reference_video_conditionings` — injects IC-LoRA video tokens from reference
- `patchify_dubit_audio_reference_latent` — converts audio VAE latents to patches with RoPE positions
- `AudioConditionByReferenceLatent` — wrapper for audio conditioning objects

## Model Requirements for DubItPipeline

The pipeline requires these specific checkpoint files:

| Component | Required File |
|-----------|---------------|
| Transformer | `ltx-2.5-22b-distilled-transformer-bf16.safetensors` |
| Text Encoder | `gemma4-12b-with-proj-ltx-2.5-bf16.safetensors` |
| Video VAE | `ltx-2.5-video-vae-bf16.safetensors` |
| Audio VAE | `ltx-2.5-audio-vae-bf16.safetensors` |
| Spatial Upsampler | `ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors` |
| IC-LoRA | `Lightricks/LTX-2.3-22b-IC-LoRA-DubIt` |

All components use **bfloat16** precision. The distilled transformer enables faster inference with fewer sampling steps compared to the full model.

## Summary

- **DubItPipeline** performs re-voicing with lip-sync by conditioning on both video frames and audio latents from a reference clip
- A **single IC-LoRA** (`Lightricks/LTX-2.3-22b-IC-LoRA-DubIt`) is applied in both diffusion stages for consistent identity preservation
- The **two-stage architecture** generates low-res latents first, then upsamples and refines for final output
- **CLI usage** requires no frame-rate or frame-count parameters — these are inferred from the reference video
- **Python API** provides full control via `DubItPipeline` class and `_create_stage_conditionings` for custom conditioning

## Frequently Asked Questions

### What makes DubItPipeline different from standard text-to-video generation?

**DubItPipeline preserves the original speaker's facial identity and lip synchronization** by using the reference video's audio latent as frozen conditioning tokens. Standard text-to-video pipelines generate entirely new content without this cross-modal alignment mechanism.

### Why does the pipeline use the same IC-LoRA in both stages?

**Applying `Lightricks/LTX-2.3-22b-IC-LoRA-DubIt` consistently across stages maintains visual-audio coherence.** The LoRA's image-conditioning tokens carry speaker identity information that must persist from the initial generation through the upsampling refinement. The down-scale factor from LoRA metadata adjusts token scaling automatically.

### How does the audio reference latent work for lip-sync?

**The audio VAE latent is patchified with negative RoPE positions via `patchify_dubit_audio_reference_latent` and appended as frozen tokens.** This positional encoding scheme aligns the reference audio with the target generation timeline, allowing the model to match mouth movements to the new speech content while preserving the original speaker's vocal characteristics.