How the Dub-It Pipeline Works for Audio Rephrasing with Lip Sync in LTX-2
The Dub-It pipeline in LTX-2 is a two-stage diffusion system that replaces a video's visual content while preserving the original audio track and maintaining perfect lip synchronization through frozen audio latents and RoPE-based conditioning.
This guide explains the complete technical architecture of Lightricks' Dub-It audio rephrasing pipeline, from reference audio extraction to joint video-audio diffusion. The pipeline is implemented as DubItPipeline in the LTX-2 inference framework, enabling high-quality dubbing without traditional frame-rate constraints.
Dub-It Pipeline Architecture
The DubItPipeline class in packages/ltx-pipelines/src/ltx_pipelines/dubit.py orchestrates a sophisticated two-stage process that keeps audio and video tightly coupled.
Core Components
| Component | Role | Location |
|---|---|---|
DubItPipeline |
Main orchestrator loading prompt encoder, conditioners, diffusion, and upsampler | ltx_pipelines/dubit.py lines 67-84 |
| Audio VAE encoder | Extracts and encodes reference audio from source video | ltx_pipelines/dubit.py lines 86-92 |
| IC-LoRA conditioner | Adds video reference tokens for identity preservation | ltx_pipelines/dubit.py lines 149-180 |
| Two-stage diffusion | Joint video-audio generation with frozen audio latents | ltx_pipelines/dubit.py lines 279-326 |
| Spatial upsampler | Doubles video resolution in latent space | ltx_pipelines/dubit.py lines 300-312 |
The pipeline operates without explicit frame-rate or frame-count arguments—these are inferred directly from the reference video container via _verify_media_path_args.
Audio Conditioning and Lip-Sync Mechanism
Step 1: Reference Audio Extraction
The pipeline extracts audio from the source video using decode_audio_from_file, then encodes it through the audio VAE:
# Internal pipeline flow (simplified)
reference_audio = decode_audio_from_file(reference_video_path)
audio_latent = vae_encode_audio(reference_audio)
This creates AudioConditionByReferenceLatent (defined in ltx_core/conditioning.py), a conditioning token that maintains temporal alignment with video frames.
Step 2: RoPE-Based Temporal Alignment
The patchify_dubit_audio_reference_latent function (lines 335-355) performs critical patchification:
- Splits audio latent into time-aligned tokens
- Builds a RoPE position map linking each audio token to its corresponding video frame
- Preserves phase relationships for lip-sync accuracy
Step 3: Frozen Audio in Second Stage
The key innovation ensuring lip-sync: the audio latent is marked frozen=True for stage 2 diffusion. While the video latent undergoes spatial upsampling, the audio receives no additional noise, guaranteeing decoded audio matches the original waveform precisely.
Command-Line Usage
The dubit_arg_parser in ltx_pipelines/utils/args.py (lines 80-88) provides a minimal CLI requiring only:
| Flag | Purpose |
|---|---|
--prompt |
Visual description for the generated content |
--reference-video |
Source video providing audio track and frame count |
--spatial-upsampler-path |
Checkpoint for resolution doubling |
--lora |
Dub-It IC-LoRA model (exactly one required) |
--height / --width |
Target resolution (multiple of 64) |
--reference-strength |
Video conditioning intensity (default 1.0) |
Basic CLI Example
python -m ltx_pipelines.dubit \
--prompt "A chef presenting a new dish" \
--reference-video /data/original_clip.mp4 \
--spatial-upsampler-path /models/spatial_upsampler.safetensors \
--lora /models/dubit_ic_lora.safetensors \
--output-path /results/dubbed_output.mp4 \
--height 512 \
--width 512 \
--seed 42
Execution flow:
- Parse arguments and validate media paths
- Load
DubItPipelinewith IC-LoRA weights - Extract and encode reference audio
- Run joint diffusion (stage 1)
- Upsample video latent, run stage 2 with frozen audio
- Mux final video with original audio into output MP4
Programmatic Implementation
For custom integrations, instantiate DubItPipeline directly:
from ltx_pipelines.dubit import DubItPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
# Configure model paths
model_paths = ModelPaths(
transformer="path/to/transformer.safetensors",
video_vae="path/to/video_vae.safetensors",
audio_vae="path/to/audio_vae.safetensors",
)
# Initialize pipeline with Dub-It IC-LoRA
pipeline = DubItPipeline(
model_paths=model_paths,
spatial_upsampler_path="path/to/spatial_upsampler.safetensors",
ic_lora=ModelPaths.lora_from_path("path/to/dubit_ic_lora.safetensors"),
device="cuda",
)
# Generate dubbed video
video, audio, _ = pipeline(
prompt="A robot delivering a speech",
seed=123,
height=512,
width=512,
images=[], # optional image conditioning
reference_video_path="original.mp4",
reference_strength=1.0,
)
# Export with preserved audio sync
from ltx_pipelines.utils.media_io import (
encode_video, get_videostream_metadata, get_video_chunks_number
)
meta = get_videostream_metadata("original.mp4")
encode_video(
video=video,
fps=int(meta.fps),
audio=audio, # Original reference audio, perfectly synced
output_path="dubbed_result.mp4",
video_chunks_number=get_video_chunks_number(meta.frames, None),
)
The returned audio tensor is identical to the decoded reference audio—no generation artifacts or timing drift.
Key Implementation Files
| File | Content | Link |
|---|---|---|
ltx_pipelines/dubit.py |
Complete DubItPipeline implementation |
source |
ltx_pipelines/utils/args.py |
dubit_arg_parser CLI definition |
source |
ltx_pipelines/utils/media_io/__init__.py |
decode_audio_from_file, encode_video helpers |
source |
ltx_core/model/audio_vae.py |
Audio VAE encoder/decoder | source |
ltx_core/conditioning.py |
AudioConditionByReferenceLatent definition |
source |
Summary
- Dub-It is a two-stage diffusion pipeline that regenerates video visuals while freezing audio latents to preserve exact timing
- RoPE-based conditioning aligns audio tokens with video frames for natural lip movement
- No manual frame-rate specification—timing is inferred from the reference video container
- IC-LoRA video reference maintains subject identity across visual changes
- Spatial upsampling operates only on video, leaving audio untouched for perfect sync
Frequently Asked Questions
How does Dub-It maintain lip synchronization without explicit phoneme alignment?
The pipeline uses frozen audio latents and RoPE position encoding rather than traditional phoneme detection. By encoding the reference audio through the audio VAE and preventing any diffusion noise in stage 2, the decoded audio matches the original waveform exactly. The patchify_dubit_audio_reference_latent function creates temporally-aligned tokens that the video diffusion attends to naturally.
What makes the Dub-It LoRA different from standard IC-LoRA models?
Dub-It requires exactly one IC-LoRA checkpoint specifically trained for audio-guided video generation. This LoRA encodes video reference information through the standard IC-LoRA token mechanism (lines 149-180 in dubit.py), but is optimized for cases where audio timing must be preserved while visual content changes.
Can I use Dub-It with custom audio instead of extracted reference audio?
According to the source implementation, the pipeline is designed around reference video conditioning—the audio is extracted automatically via decode_audio_from_file. For custom audio, you would need to modify the _encode_reference_audio_vae_latent call or create a video container with your target audio track.
Why does the pipeline require a spatial upsampler as a separate argument?
The two-stage architecture separates generation from resolution enhancement. Stage 1 produces base-resolution latents for both modalities; only the video latent proceeds through spatial_upsampler_path (lines 300-312). This design keeps the audio pathway lightweight while allowing flexible video resolution scaling without re-running audio encoding.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →