How to Use the Dub-It Pipeline for Audio Dubbing with Lip Sync in LTX-2

The Dub-It pipeline in LTX-2 replaces a video's audio track while maintaining perfect lip synchronization through a two-stage diffusion process that freezes the original audio latent.

The Dub-It pipeline is a specialized inference pipeline in Lightricks' LTX-2 that enables high-quality audio dubbing with automatic lip-sync. Unlike standard video generation, it preserves the exact timing and content of a reference audio track while regenerating visual frames to match. This guide explains how to run Dub-It from both command-line and Python interfaces according to the official LTX-2 source code.

Architecture of the Dub-It Pipeline

The pipeline is implemented in packages/ltx-pipelines/src/ltx_pipelines/dubit.py and consists of five core components working in sequence:

Core Components

  • DubItPipeline class — Orchestrates the entire dubbing process. It loads the prompt encoder, video/audio conditioners, diffusion stage, and spatial up-sampler. Located at lines 67-84 in dubit.py.

  • Audio conditioning module — Extracts reference audio from the input video, encodes it via the audio VAE, and constructs a RoPE-based reference token using AudioConditionByReferenceLatent. See lines 86-92.

  • Video conditioning module — Builds image conditionings and injects IC-LoRA video reference tokens for visual fidelity. Implemented at lines 149-180.

  • Two-stage diffusion — Stage 1 runs joint diffusion on both video and audio latents. The audio latent is marked frozen=True for Stage 2, ensuring timing preservation. Found at lines 279-326.

  • Spatial up-sampler — Doubles video resolution in latent space between stages (lines 300-312), while audio remains untouched.

Lip Sync Guarantee Mechanism

The pipeline guarantees lip sync through four specific technical steps:

  1. Audio latent extraction — self._encode_reference_audio_vae_latent reads the reference video's audio track via decode_audio_from_file and encodes it with vae_encode_audio.

  2. RoPE position alignment — patchify_dubit_audio_reference_latent (lines 335-355) creates patchified latents with frame-aligned position maps.

  3. Frozen audio refinement — During Stage 2 diffusion, the audio latent receives zero additional noise—only the video latent is denoised at higher resolution.

  4. Direct audio decoding — The final audio is decoded directly from the frozen latent without modification, then muxed with the regenerated video.

Command-Line Usage

The simplest way to run Dub-It is through the dedicated CLI. Argument parsing is handled by dubit_arg_parser in ltx_pipelines/utils/args.py (lines 80-88).

Required Arguments

Flag Purpose
--prompt Text description of desired visual content
--reference-video Source video providing both audio track and frame count
--spatial-upsampler-path Path to the up-sampler checkpoint (Stage 2)
--lora Path to Dub-It IC-LoRA model (exactly one required)
--output-path Destination MP4 file path

Optional Arguments

Flag Default Description
--height / --width — Target resolution (must be divisible by 64)
--reference-strength 1.0 Strength of video reference conditioning
--seed random Reproducibility seed
--enhance-prompt False LLM-based prompt enhancement
--compile False PyTorch compilation for speed

CLI Example

python -m ltx_pipelines.dubit \
  --prompt "A chef presenting a new dish" \
  --reference-video /data/original_clip.mp4 \
  --spatial-upsampler-path /models/spatial_upsampler.safetensors \
  --lora /models/dubit_ic_lora.safetensors \
  --output-path /results/dubbed_output.mp4 \
  --height 512 \
  --width 512 \
  --seed 42

No frame-rate or frame-count arguments are needed—these are inferred automatically from the reference video container via _verify_media_path_args.

Programmatic Usage

For integration into larger workflows, instantiate DubItPipeline directly:

Python Example

from ltx_pipelines.dubit import DubItPipeline, patchify_dubit_audio_reference_latent
from ltx_pipelines.utils.model_paths import ModelPaths

# Step 1: Configure model paths

model_paths = ModelPaths(
    transformer="path/to/transformer.safetensors",
    video_vae="path/to/video_vae.safetensors",
    audio_vae="path/to/audio_vae.safetensors",
)

# Step 2: Initialize pipeline with spatial upsampler and IC-LoRA

pipeline = DubItPipeline(
    model_paths=model_paths,
    spatial_upsampler_path="path/to/spatial_upsampler.safetensors",
    ic_lora=ModelPaths.lora_from_path("path/to/dubit_ic_lora.safetensors"),
    device="cuda",
)

# Step 3: Run inference

video, audio, _ = pipeline(
    prompt="A robot delivering a speech",
    seed=123,
    height=512,
    width=512,
    images=[],                         # No additional image conditioning

    reference_video_path="original.mp4",
    reference_strength=1.0,
)

# Step 4: Save output using media I/O helpers

from ltx_pipelines.utils.media_io import (
    encode_video,
    get_videostream_metadata,
    get_video_chunks_number,
)

meta = get_videostream_metadata("original.mp4")
encode_video(
    video=video,
    fps=int(meta.fps),
    audio=audio,
    output_path="dubbed_result.mp4",
    video_chunks_number=get_video_chunks_number(meta.frames, None),
)

Critical implementation detail: The audio tensor returned by pipeline.__call__ is decoded directly from the frozen reference latent—guaranteeing bit-identical timing to the original source.

Key Implementation Files

File Contents Line References
ltx_pipelines/dubit.py Full pipeline: audio extraction, joint diffusion, up-sampling 67-84 (init), 86-92 (audio cond), 149-180 (video cond), 279-326 (diffusion), 300-312 (upsampler), 335-355 (patchify)
ltx_pipelines/utils/args.py CLI argument parser with Dub-It-specific validation 80-88
ltx_pipelines/utils/media_io/__init__.py Audio decoding (decode_audio_from_file) and video encoding (encode_video) —
ltx_core/model/audio_vae.py Audio VAE encoder/decoder for latent extraction —
ltx_core/conditioning.py AudioConditionByReferenceLatent class definition —

Summary

  • Dub-It is a two-stage diffusion pipeline specifically designed for audio-dubbed video generation with guaranteed lip sync.
  • The reference video provides both the audio track to preserve and the frame count/timing to match—no manual alignment needed.
  • Audio latents are frozen during Stage 2, ensuring decoded audio matches the original speech waveform exactly.
  • IC-LoRA conditioning maintains visual identity from the reference while allowing prompt-driven visual changes.
  • Both CLI and Python APIs are fully supported in ltx_pipelines.dubit.

Frequently Asked Questions

What makes Dub-It different from standard video generation pipelines?

Standard pipelines generate both audio and video from scratch or condition on separate audio inputs. Dub-It specifically extracts and freezes the reference audio latent from an existing video, ensuring the generated lip movements match the original speech timing with frame-level precision. This is implemented through AudioConditionByReferenceLatent and the frozen=True flag in Stage 2 diffusion.

Why does Dub-It require a spatial up-sampler?

The pipeline runs diffusion at a lower resolution in Stage 1 for efficiency, then uses the spatial up-sampler to double the video resolution in Stage 2. The audio latent bypasses this up-sampling entirely—only video latents are processed—maintaining the original audio quality while improving visual fidelity.

Can I use multiple LoRAs with Dub-It?

No. The dubit_arg_parser explicitly enforces exactly one LoRA path via --lora. This LoRA must be a Dub-It IC-LoRA trained for video reference conditioning. The constraint is validated in ltx_pipelines/utils/args.py to ensure proper audio-visual alignment.

How does the pipeline handle frame rate mismatches?

It doesn't need to. Frame rate and frame count are inferred directly from the reference video container via _verify_media_path_args. The RoPE position embeddings in patchify_dubit_audio_reference_latent automatically align audio tokens to video frames based on these extracted metadata, eliminating manual configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →