How to Use the LTX-2 DubItPipeline for Re-Voicing and Lip-Sync Synchronization
The LTX-2 DubItPipeline rewrites spoken content while preserving the original speaker's identity, lip movements, and facial expressions using a two-stage diffusion process with a single IC-LoRA applied to both stages.
The DubItPipeline in Lightricks/LTX-2 enables developers to generate synchronized video and audio from a reference clip and new text prompt. This pipeline leverages the distilled LTX-2.5 model and specialized IC-LoRA conditioning to achieve state-of-the-art re-voicing results. Below is a comprehensive guide to its architecture, API, and practical usage.
How DubItPipeline Works
The pipeline operates as a two-stage generation engine that processes reference video frames and audio simultaneously. Both visual and audio cues from the reference drive the generation, with the audio latent appended as frozen reference tokens to ensure lip-sync alignment.
Core Architecture Components
The pipeline initializes six primary components in packages/ltx-pipelines/src/ltx_pipelines/dubit.py:
- PromptEncoder – encodes the input text prompt
- ImageConditioner – handles visual conditioning signals
- AudioConditioner – processes audio reference latents
- DiffusionStage – runs the noise-to-latent diffusion process
- VideoUpsampler – spatially upsamples low-resolution video latents
- VideoDecoder / AudioDecoder – VAE-based decoders for final output
All components operate in bfloat16 precision on the same device. The IC-LoRA is loaded via LoraPathStrengthAndSDOps, with its down-scale factor read from LoRA metadata.
Two-Stage Conditioning Process
Stage 1: Initial Generation
- Image conditioning is built with
combined_image_conditionings - Reference video frames are injected via
append_ic_lora_reference_video_conditionings— this adds IC-LoRA tokens aligned to the reference - Audio conditioning extracts the reference audio stream, encodes it with
vae_encode_audio, and patchifies viapatchify_dubit_audio_reference_latent(adding negative RoPE positions for reference) - Diffusion runs with
DISTILLED_SIGMASto produce low-resolution video and audio latents
Stage 2: Upscaling and Refinement
- The video latent is upsampled by the spatial upsampler
- Video conditionings are recomputed for target resolution
- Audio conditionings reuse the frozen reference latent from stage 1
- Final diffusion produces the high-resolution output
The pipeline returns (decoded_video, decoded_audio, tiling_config).
Running DubItPipeline from the Command Line
The fastest way to use the pipeline is via the ltx_pipelines.dubit CLI entry point. No frame count or frame rate parameters are required — these are inferred automatically from the reference video.
uv run python -m ltx_pipelines.dubit \
--transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
--text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
--video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
--audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
--spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
--lora Lightricks/LTX-2.3-22b-IC-LoRA-DubIt \
--reference-video path/to/original_clip.mp4 \
--prompt "Hello, I'm excited to announce the new feature!" \
--seed 12345 \
--height 720 \
--width 1280 \
--output-path dubit_result.mp4
Key CLI behavior: The --reference-video flag provides both visual frames and audio VAE latents. The pipeline automatically derives frame count and FPS from the source file.
Using the DubItPipeline Python API
For custom workflows, instantiate DubItPipeline directly and call it with your parameters.
from ltx_pipelines.dubit import DubItPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
# Build paths to required model files
model_paths = ModelPaths(
transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
text_encoder="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
video_vae="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
audio_vae="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)
pipeline = DubItPipeline(
model_paths=model_paths,
spatial_upsampler_path=model_paths.spatial_upsampler,
ic_lora=ModelPaths.lora_path("Lightricks/LTX-2.3-22b-IC-LoRA-DubIt"),
)
video, audio, _ = pipeline(
prompt="Welcome to the next generation of video synthesis!",
seed=42,
height=720,
width=1280,
images=[],
reference_video_path="reference.mp4",
)
# Save the result
from ltx_pipelines.utils.media_io import encode_video
encode_video(video, fps=30, audio=audio, output_path="dubbed.mp4")
The pipeline() call returns decoded video frames, decoded audio samples, and the tiling configuration used during VAE decoding.
Customizing Conditioning Parameters
To modify reference strength or inject additional image conditionings, use the internal _create_stage_conditionings method or replicate its logic in your code.
# Example: increase reference strength to 1.5
custom_conditionings = pipeline._create_stage_conditionings(
images=[],
reference_video_path="reference.mp4",
reference_strength=1.5,
height=360,
width=640,
num_frames=121,
video_encoder=your_video_encoder,
encode_tiling=your_tiling_cfg,
)
This returns a list of conditioning objects that can be passed directly to pipeline.stage for fine-grained control over the diffusion process.
Key Source Files and Functions
Understanding these files helps when debugging or extending the pipeline:
| File | Purpose |
|---|---|
packages/ltx-pipelines/src/ltx_pipelines/dubit.py |
Main DubItPipeline implementation |
packages/ltx-core/src/ltx_core/conditioning.py |
AudioConditionByReferenceLatent and conditioning helpers |
packages/ltx-core/src/ltx_core/model/audio_vae.py |
vae_encode_audio and audio latent operations |
packages/ltx-core/src/ltx_core/model/video_vae.py |
Video VAE encoder, decoder, and tiling utilities |
packages/ltx-pipelines/utils/media_io.py |
encode_video and media I/O functions |
Critical functions to know:
append_ic_lora_reference_video_conditionings— injects IC-LoRA video tokens from referencepatchify_dubit_audio_reference_latent— converts audio VAE latents to patches with RoPE positionsAudioConditionByReferenceLatent— wrapper for audio conditioning objects
Model Requirements for DubItPipeline
The pipeline requires these specific checkpoint files:
| Component | Required File |
|---|---|
| Transformer | ltx-2.5-22b-distilled-transformer-bf16.safetensors |
| Text Encoder | gemma4-12b-with-proj-ltx-2.5-bf16.safetensors |
| Video VAE | ltx-2.5-video-vae-bf16.safetensors |
| Audio VAE | ltx-2.5-audio-vae-bf16.safetensors |
| Spatial Upsampler | ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors |
| IC-LoRA | Lightricks/LTX-2.3-22b-IC-LoRA-DubIt |
All components use bfloat16 precision. The distilled transformer enables faster inference with fewer sampling steps compared to the full model.
Summary
- DubItPipeline performs re-voicing with lip-sync by conditioning on both video frames and audio latents from a reference clip
- A single IC-LoRA (
Lightricks/LTX-2.3-22b-IC-LoRA-DubIt) is applied in both diffusion stages for consistent identity preservation - The two-stage architecture generates low-res latents first, then upsamples and refines for final output
- CLI usage requires no frame-rate or frame-count parameters — these are inferred from the reference video
- Python API provides full control via
DubItPipelineclass and_create_stage_conditioningsfor custom conditioning
Frequently Asked Questions
What makes DubItPipeline different from standard text-to-video generation?
DubItPipeline preserves the original speaker's facial identity and lip synchronization by using the reference video's audio latent as frozen conditioning tokens. Standard text-to-video pipelines generate entirely new content without this cross-modal alignment mechanism.
Why does the pipeline use the same IC-LoRA in both stages?
Applying Lightricks/LTX-2.3-22b-IC-LoRA-DubIt consistently across stages maintains visual-audio coherence. The LoRA's image-conditioning tokens carry speaker identity information that must persist from the initial generation through the upsampling refinement. The down-scale factor from LoRA metadata adjusts token scaling automatically.
How does the audio reference latent work for lip-sync?
The audio VAE latent is patchified with negative RoPE positions via patchify_dubit_audio_reference_latent and appended as frozen tokens. This positional encoding scheme aligns the reference audio with the target generation timeline, allowing the model to match mouth movements to the new speech content while preserving the original speaker's vocal characteristics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →