LTX-2 Pipeline for Lip Dubbing and Audio Rephrasing: Implementation Guide
The LTX-2 framework provides two specialized pipelines for these tasks: LipDubPipeline for lip-synced video generation with audio reference conditioning, and T2AOneStagePipeline for text-to-audio rephrasing using only audio weights.
The Lightricks LTX-2 repository offers production-ready inference pipelines for advanced video and audio manipulation. For developers building lip-sync applications or audio replacement tools, understanding which LTX-2 pipeline handles lip dubbing versus audio rephrasing is critical for optimal performance and resource usage.
LTX-2 Lip Dubbing Pipeline (LipDubPipeline)
For lip dubbing tasks that require synchronizing lip movements to a reference audio track while maintaining video consistency, use the LipDubPipeline class.
Core Architecture
Located in packages/ltx-pipelines/src/ltx_pipelines/lipdub.py, this pipeline implements a two-stage diffusion process:
- Reference Processing: Loads a distilled checkpoint and spatial upsampler, then extracts the audio track from the reference video using IC-LoRA conditioning
- Audio Encoding: Encodes the reference audio through the audio VAE (
vae_encode_audio), patches it viaAudioPatchifier, and injects it as anAudioConditionByReferenceLatentconditioning token - Two-Stage Diffusion:
- First pass generates low-resolution video latent
- Second pass upsamples while keeping the audio latent frozen (
audio_frozen=True)
CLI Usage Example
Run the lip dubbing pipeline from the command line:
python -m ltx_pipelines.lipdub \
--distilled-checkpoint-path <distilled.ckpt> \
--spatial-upsampler-path <spatial_up.ckpt> \
--gemma-root <gemma_root> \
--lora <ic_lora_path>:<strength> \
--reference-video <path/to/source_video.mp4> \
--prompt "A smiling person says hello" \
--seed 42 \
--height 720 \
--width 1280 \
--output-path ./output_lipdub.mp4
Internally, LipDubPipeline.__call__ builds a GaussianNoiser, encodes prompts via PromptEncoder, and manages the VideoUpsampler between diffusion stages. Final output is multiplexed using encode_video from the media I/O utilities.
LTX-2 Audio Rephrasing Pipeline (T2AOneStagePipeline)
For audio rephrasing (text-to-audio generation), the T2AOneStagePipeline provides a lightweight, audio-only solution that avoids loading video weights entirely.
Core Architecture
Located in packages/ltx-pipelines/src/ltx_pipelines/t2a_one_stage.py, this pipeline is configured specifically for audio generation:
- Model Configuration: Uses
LTXAudioOnlyModelConfiguratorto load only audio weights, reducing memory footprint - Single-Stage Diffusion: Performs one diffusion pass on audio latents using
DiffusionStagewith multimodal guidance - Guidance Mechanism: Supports CFG (Classifier-Free Guidance), STG (Self-Guidance), and other modalities via
FactoryGuidedDenoiserandMultiModalGuiderFactory
Python API Implementation
Integrate the audio rephrasing pipeline programmatically:
from ltx_pipelines.t2a_one_stage import T2AOneStagePipeline
from ltx_pipelines.utils.args import detect_checkpoint_path, default_1_stage_t2a_arg_parser
from ltx_pipelines.utils.constants import detect_params
from ltx_pipelines.utils.media_io import encode_audio
# Detect checkpoint and parameters
checkpoint = detect_checkpoint_path()
params = detect_params(checkpoint)
# Initialize pipeline with audio-only configuration
pipeline = T2AOneStagePipeline(
checkpoint_path=checkpoint,
gemma_root="/opt/gemma",
loras=(),
)
# Generate rephrased audio
audio = pipeline(
prompt="A calm voice says hello world",
negative_prompt="",
seed=1234,
num_frames=500,
frame_rate=24.0,
num_inference_steps=50,
audio_guider_params=MultiModalGuiderParams(
cfg_scale=7.0,
stg_scale=1.0,
rescale_scale=0.0,
modality_scale=1.0,
skip_step=0,
stg_blocks=0,
),
)
# Export to file
encode_audio(audio=audio, output_path="rephrased.wav")
The pipeline outputs raw audio through AudioDecoder (defined in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py), ready for encoding via encode_audio.
Shared Components and Audio VAE
Both pipelines rely on common building blocks from packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py:
- PromptEncoder: Handles text conditioning for both pipelines
- AudioDecoder: Decodes audio latents into waveforms
- DiffusionStage: Core diffusion logic shared across modalities
The audio VAE implementation in packages/ltx-core/src/ltx_core/model/audio_vae/audio_vae.py provides the vae_encode_audio function used by LipDubPipeline to process reference audio, ensuring consistent latent representations across both pipelines.
Summary
- LipDubPipeline (
packages/ltx-pipelines/src/ltx_pipelines/lipdub.py) handles video+audio conditioning for lip-sync tasks using two-stage diffusion and frozen audio latents - T2AOneStagePipeline (
packages/ltx-pipelines/src/ltx_pipelines/t2a_one_stage.py) provides lightweight text-to-audio generation usingLTXAudioOnlyModelConfiguratorand single-stage diffusion - Both pipelines use
AudioConditionByReferenceLatentfor conditioning and shareAudioDecoderfrom the utilities module - The audio VAE in
ltx_corestandardizes audio latent encoding across the framework
Frequently Asked Questions
What is the difference between LipDubPipeline and T2AOneStagePipeline in LTX-2?
LipDubPipeline is a two-stage video generation pipeline that accepts video and audio references to produce lip-synced output, while T2AOneStagePipeline is an audio-only pipeline for pure text-to-audio generation without video components. The former uses LTXVideoModelConfigurator and spatial upsamplers, whereas the latter uses LTXAudioOnlyModelConfigurator to minimize memory usage.
How does LTX-2 handle audio conditioning in the lip dubbing pipeline?
The pipeline extracts audio from the reference video using _encode_reference_audio_vae_latent, encodes it through the audio VAE, and patches it into AudioConditionByReferenceLatent tokens. These tokens are injected into the diffusion process alongside video conditionings, with the audio latent remaining frozen during the upsampling stage to preserve speech content.
Can I use T2AOneStagePipeline for video generation?
No. T2AOneStagePipeline explicitly configures the model using LTXAudioOnlyModelConfigurator, which excludes video weights entirely. For video generation, use LipDubPipeline or other video-specific pipelines in the ltx_pipelines package that load full video model weights.
What checkpoint configuration is required for audio rephrasing in LTX-2?
Audio rephrasing requires a checkpoint compatible with LTXAudioOnlyModelConfigurator. The pipeline uses detect_checkpoint_path() and detect_params() to automatically configure the correct audio VAE and model architecture, loading only the audio components necessary for the DiffusionStage to perform single-pass generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →