How to Use the LTX-2 DubItPipeline for Re-Voicing and Lip-Sync Synchronization

The LTX-2 DubItPipeline rewrites spoken content while preserving the original speaker's identity, lip movements, and facial expressions using a two-stage diffusion process with a single IC-LoRA applied to both stages.

The DubItPipeline in Lightricks/LTX-2 enables developers to generate synchronized video and audio from a reference clip and new text prompt. This pipeline leverages the distilled LTX-2.5 model and specialized IC-LoRA conditioning to achieve state-of-the-art re-voicing results. Below is a comprehensive guide to its architecture, API, and practical usage.

How DubItPipeline Works

The pipeline operates as a two-stage generation engine that processes reference video frames and audio simultaneously. Both visual and audio cues from the reference drive the generation, with the audio latent appended as frozen reference tokens to ensure lip-sync alignment.

Core Architecture Components

The pipeline initializes six primary components in packages/ltx-pipelines/src/ltx_pipelines/dubit.py:

  • PromptEncoder – encodes the input text prompt
  • ImageConditioner – handles visual conditioning signals
  • AudioConditioner – processes audio reference latents
  • DiffusionStage – runs the noise-to-latent diffusion process
  • VideoUpsampler – spatially upsamples low-resolution video latents
  • VideoDecoder / AudioDecoder – VAE-based decoders for final output

All components operate in bfloat16 precision on the same device. The IC-LoRA is loaded via LoraPathStrengthAndSDOps, with its down-scale factor read from LoRA metadata.

Two-Stage Conditioning Process

Stage 1: Initial Generation

  1. Image conditioning is built with combined_image_conditionings
  2. Reference video frames are injected via append_ic_lora_reference_video_conditionings — this adds IC-LoRA tokens aligned to the reference
  3. Audio conditioning extracts the reference audio stream, encodes it with vae_encode_audio, and patchifies via patchify_dubit_audio_reference_latent (adding negative RoPE positions for reference)
  4. Diffusion runs with DISTILLED_SIGMAS to produce low-resolution video and audio latents

Stage 2: Upscaling and Refinement

  1. The video latent is upsampled by the spatial upsampler
  2. Video conditionings are recomputed for target resolution
  3. Audio conditionings reuse the frozen reference latent from stage 1
  4. Final diffusion produces the high-resolution output

The pipeline returns (decoded_video, decoded_audio, tiling_config).

Running DubItPipeline from the Command Line

The fastest way to use the pipeline is via the ltx_pipelines.dubit CLI entry point. No frame count or frame rate parameters are required — these are inferred automatically from the reference video.

uv run python -m ltx_pipelines.dubit \
    --transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
    --text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
    --video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
    --audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
    --spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
    --lora Lightricks/LTX-2.3-22b-IC-LoRA-DubIt \
    --reference-video path/to/original_clip.mp4 \
    --prompt "Hello, I'm excited to announce the new feature!" \
    --seed 12345 \
    --height 720 \
    --width 1280 \
    --output-path dubit_result.mp4

Key CLI behavior: The --reference-video flag provides both visual frames and audio VAE latents. The pipeline automatically derives frame count and FPS from the source file.

Using the DubItPipeline Python API

For custom workflows, instantiate DubItPipeline directly and call it with your parameters.

from ltx_pipelines.dubit import DubItPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

# Build paths to required model files

model_paths = ModelPaths(
    transformer="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
    text_encoder="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
    video_vae="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
    audio_vae="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
    spatial_upsampler="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
)

pipeline = DubItPipeline(
    model_paths=model_paths,
    spatial_upsampler_path=model_paths.spatial_upsampler,
    ic_lora=ModelPaths.lora_path("Lightricks/LTX-2.3-22b-IC-LoRA-DubIt"),
)

video, audio, _ = pipeline(
    prompt="Welcome to the next generation of video synthesis!",
    seed=42,
    height=720,
    width=1280,
    images=[],
    reference_video_path="reference.mp4",
)

# Save the result

from ltx_pipelines.utils.media_io import encode_video
encode_video(video, fps=30, audio=audio, output_path="dubbed.mp4")

The pipeline() call returns decoded video frames, decoded audio samples, and the tiling configuration used during VAE decoding.

Customizing Conditioning Parameters

To modify reference strength or inject additional image conditionings, use the internal _create_stage_conditionings method or replicate its logic in your code.


# Example: increase reference strength to 1.5

custom_conditionings = pipeline._create_stage_conditionings(
    images=[],
    reference_video_path="reference.mp4",
    reference_strength=1.5,
    height=360,
    width=640,
    num_frames=121,
    video_encoder=your_video_encoder,
    encode_tiling=your_tiling_cfg,
)

This returns a list of conditioning objects that can be passed directly to pipeline.stage for fine-grained control over the diffusion process.

Key Source Files and Functions

Understanding these files helps when debugging or extending the pipeline:

File Purpose
packages/ltx-pipelines/src/ltx_pipelines/dubit.py Main DubItPipeline implementation
packages/ltx-core/src/ltx_core/conditioning.py AudioConditionByReferenceLatent and conditioning helpers
packages/ltx-core/src/ltx_core/model/audio_vae.py vae_encode_audio and audio latent operations
packages/ltx-core/src/ltx_core/model/video_vae.py Video VAE encoder, decoder, and tiling utilities
packages/ltx-pipelines/utils/media_io.py encode_video and media I/O functions

Critical functions to know:

  • append_ic_lora_reference_video_conditionings — injects IC-LoRA video tokens from reference
  • patchify_dubit_audio_reference_latent — converts audio VAE latents to patches with RoPE positions
  • AudioConditionByReferenceLatent — wrapper for audio conditioning objects

Model Requirements for DubItPipeline

The pipeline requires these specific checkpoint files:

Component Required File
Transformer ltx-2.5-22b-distilled-transformer-bf16.safetensors
Text Encoder gemma4-12b-with-proj-ltx-2.5-bf16.safetensors
Video VAE ltx-2.5-video-vae-bf16.safetensors
Audio VAE ltx-2.5-audio-vae-bf16.safetensors
Spatial Upsampler ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors
IC-LoRA Lightricks/LTX-2.3-22b-IC-LoRA-DubIt

All components use bfloat16 precision. The distilled transformer enables faster inference with fewer sampling steps compared to the full model.

Summary

  • DubItPipeline performs re-voicing with lip-sync by conditioning on both video frames and audio latents from a reference clip
  • A single IC-LoRA (Lightricks/LTX-2.3-22b-IC-LoRA-DubIt) is applied in both diffusion stages for consistent identity preservation
  • The two-stage architecture generates low-res latents first, then upsamples and refines for final output
  • CLI usage requires no frame-rate or frame-count parameters — these are inferred from the reference video
  • Python API provides full control via DubItPipeline class and _create_stage_conditionings for custom conditioning

Frequently Asked Questions

What makes DubItPipeline different from standard text-to-video generation?

DubItPipeline preserves the original speaker's facial identity and lip synchronization by using the reference video's audio latent as frozen conditioning tokens. Standard text-to-video pipelines generate entirely new content without this cross-modal alignment mechanism.

Why does the pipeline use the same IC-LoRA in both stages?

Applying Lightricks/LTX-2.3-22b-IC-LoRA-DubIt consistently across stages maintains visual-audio coherence. The LoRA's image-conditioning tokens carry speaker identity information that must persist from the initial generation through the upsampling refinement. The down-scale factor from LoRA metadata adjusts token scaling automatically.

How does the audio reference latent work for lip-sync?

The audio VAE latent is patchified with negative RoPE positions via patchify_dubit_audio_reference_latent and appended as frozen tokens. This positional encoding scheme aligns the reference audio with the target generation timeline, allowing the model to match mouth movements to the new speech content while preserving the original speaker's vocal characteristics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →