How the Dub-It Pipeline Works for Audio Rephrasing with Lip Sync in LTX-2

The Dub-It pipeline in LTX-2 is a two-stage diffusion system that replaces a video's visual content while preserving the original audio track and maintaining perfect lip synchronization through frozen audio latents and RoPE-based conditioning.

This guide explains the complete technical architecture of Lightricks' Dub-It audio rephrasing pipeline, from reference audio extraction to joint video-audio diffusion. The pipeline is implemented as DubItPipeline in the LTX-2 inference framework, enabling high-quality dubbing without traditional frame-rate constraints.

Dub-It Pipeline Architecture

The DubItPipeline class in packages/ltx-pipelines/src/ltx_pipelines/dubit.py orchestrates a sophisticated two-stage process that keeps audio and video tightly coupled.

Core Components

Component Role Location
DubItPipeline Main orchestrator loading prompt encoder, conditioners, diffusion, and upsampler ltx_pipelines/dubit.py lines 67-84
Audio VAE encoder Extracts and encodes reference audio from source video ltx_pipelines/dubit.py lines 86-92
IC-LoRA conditioner Adds video reference tokens for identity preservation ltx_pipelines/dubit.py lines 149-180
Two-stage diffusion Joint video-audio generation with frozen audio latents ltx_pipelines/dubit.py lines 279-326
Spatial upsampler Doubles video resolution in latent space ltx_pipelines/dubit.py lines 300-312

The pipeline operates without explicit frame-rate or frame-count arguments—these are inferred directly from the reference video container via _verify_media_path_args.

Audio Conditioning and Lip-Sync Mechanism

Step 1: Reference Audio Extraction

The pipeline extracts audio from the source video using decode_audio_from_file, then encodes it through the audio VAE:


# Internal pipeline flow (simplified)

reference_audio = decode_audio_from_file(reference_video_path)
audio_latent = vae_encode_audio(reference_audio)

This creates AudioConditionByReferenceLatent (defined in ltx_core/conditioning.py), a conditioning token that maintains temporal alignment with video frames.

Step 2: RoPE-Based Temporal Alignment

The patchify_dubit_audio_reference_latent function (lines 335-355) performs critical patchification:

  • Splits audio latent into time-aligned tokens
  • Builds a RoPE position map linking each audio token to its corresponding video frame
  • Preserves phase relationships for lip-sync accuracy

Step 3: Frozen Audio in Second Stage

The key innovation ensuring lip-sync: the audio latent is marked frozen=True for stage 2 diffusion. While the video latent undergoes spatial upsampling, the audio receives no additional noise, guaranteeing decoded audio matches the original waveform precisely.

Command-Line Usage

The dubit_arg_parser in ltx_pipelines/utils/args.py (lines 80-88) provides a minimal CLI requiring only:

Flag Purpose
--prompt Visual description for the generated content
--reference-video Source video providing audio track and frame count
--spatial-upsampler-path Checkpoint for resolution doubling
--lora Dub-It IC-LoRA model (exactly one required)
--height / --width Target resolution (multiple of 64)
--reference-strength Video conditioning intensity (default 1.0)

Basic CLI Example

python -m ltx_pipelines.dubit \
  --prompt "A chef presenting a new dish" \
  --reference-video /data/original_clip.mp4 \
  --spatial-upsampler-path /models/spatial_upsampler.safetensors \
  --lora /models/dubit_ic_lora.safetensors \
  --output-path /results/dubbed_output.mp4 \
  --height 512 \
  --width 512 \
  --seed 42

Execution flow:

  1. Parse arguments and validate media paths
  2. Load DubItPipeline with IC-LoRA weights
  3. Extract and encode reference audio
  4. Run joint diffusion (stage 1)
  5. Upsample video latent, run stage 2 with frozen audio
  6. Mux final video with original audio into output MP4

Programmatic Implementation

For custom integrations, instantiate DubItPipeline directly:

from ltx_pipelines.dubit import DubItPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Configure model paths

model_paths = ModelPaths(
    transformer="path/to/transformer.safetensors",
    video_vae="path/to/video_vae.safetensors",
    audio_vae="path/to/audio_vae.safetensors",
)

# Initialize pipeline with Dub-It IC-LoRA

pipeline = DubItPipeline(
    model_paths=model_paths,
    spatial_upsampler_path="path/to/spatial_upsampler.safetensors",
    ic_lora=ModelPaths.lora_from_path("path/to/dubit_ic_lora.safetensors"),
    device="cuda",
)

# Generate dubbed video

video, audio, _ = pipeline(
    prompt="A robot delivering a speech",
    seed=123,
    height=512,
    width=512,
    images=[],  # optional image conditioning

    reference_video_path="original.mp4",
    reference_strength=1.0,
)

# Export with preserved audio sync

from ltx_pipelines.utils.media_io import (
    encode_video, get_videostream_metadata, get_video_chunks_number
)

meta = get_videostream_metadata("original.mp4")
encode_video(
    video=video,
    fps=int(meta.fps),
    audio=audio,  # Original reference audio, perfectly synced

    output_path="dubbed_result.mp4",
    video_chunks_number=get_video_chunks_number(meta.frames, None),
)

The returned audio tensor is identical to the decoded reference audio—no generation artifacts or timing drift.

Key Implementation Files

File Content Link
ltx_pipelines/dubit.py Complete DubItPipeline implementation source
ltx_pipelines/utils/args.py dubit_arg_parser CLI definition source
ltx_pipelines/utils/media_io/__init__.py decode_audio_from_file, encode_video helpers source
ltx_core/model/audio_vae.py Audio VAE encoder/decoder source
ltx_core/conditioning.py AudioConditionByReferenceLatent definition source

Summary

  • Dub-It is a two-stage diffusion pipeline that regenerates video visuals while freezing audio latents to preserve exact timing
  • RoPE-based conditioning aligns audio tokens with video frames for natural lip movement
  • No manual frame-rate specification—timing is inferred from the reference video container
  • IC-LoRA video reference maintains subject identity across visual changes
  • Spatial upsampling operates only on video, leaving audio untouched for perfect sync

Frequently Asked Questions

How does Dub-It maintain lip synchronization without explicit phoneme alignment?

The pipeline uses frozen audio latents and RoPE position encoding rather than traditional phoneme detection. By encoding the reference audio through the audio VAE and preventing any diffusion noise in stage 2, the decoded audio matches the original waveform exactly. The patchify_dubit_audio_reference_latent function creates temporally-aligned tokens that the video diffusion attends to naturally.

What makes the Dub-It LoRA different from standard IC-LoRA models?

Dub-It requires exactly one IC-LoRA checkpoint specifically trained for audio-guided video generation. This LoRA encodes video reference information through the standard IC-LoRA token mechanism (lines 149-180 in dubit.py), but is optimized for cases where audio timing must be preserved while visual content changes.

Can I use Dub-It with custom audio instead of extracted reference audio?

According to the source implementation, the pipeline is designed around reference video conditioning—the audio is extracted automatically via decode_audio_from_file. For custom audio, you would need to modify the _encode_reference_audio_vae_latent call or create a video container with your target audio track.

Why does the pipeline require a spatial upsampler as a separate argument?

The two-stage architecture separates generation from resolution enhancement. Stage 1 produces base-resolution latents for both modalities; only the video latent proceeds through spatial_upsampler_path (lines 300-312). This design keeps the audio pathway lightweight while allowing flexible video resolution scaling without re-running audio encoding.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →