# LTX-2 Pipeline for Lip Dubbing and Audio Rephrasing: Implementation Guide

> Master lip dubbing and audio rephrasing with the LTX-2 framework. Explore the LipDubPipeline and T2AOneStagePipeline for advanced audio generation and manipulation.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-06-21

---

**The LTX-2 framework provides two specialized pipelines for these tasks: `LipDubPipeline` for lip-synced video generation with audio reference conditioning, and `T2AOneStagePipeline` for text-to-audio rephrasing using only audio weights.**

The Lightricks LTX-2 repository offers production-ready inference pipelines for advanced video and audio manipulation. For developers building lip-sync applications or audio replacement tools, understanding which LTX-2 pipeline handles lip dubbing versus audio rephrasing is critical for optimal performance and resource usage.

## LTX-2 Lip Dubbing Pipeline (LipDubPipeline)

For **lip dubbing** tasks that require synchronizing lip movements to a reference audio track while maintaining video consistency, use the `LipDubPipeline` class.

### Core Architecture

Located in [`packages/ltx-pipelines/src/ltx_pipelines/lipdub.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/lipdub.py), this pipeline implements a two-stage diffusion process:

1. **Reference Processing**: Loads a distilled checkpoint and spatial upsampler, then extracts the audio track from the reference video using IC-LoRA conditioning
2. **Audio Encoding**: Encodes the reference audio through the audio VAE (`vae_encode_audio`), patches it via `AudioPatchifier`, and injects it as an `AudioConditionByReferenceLatent` conditioning token
3. **Two-Stage Diffusion**: 
   - First pass generates low-resolution video latent
   - Second pass upsamples while keeping the audio latent frozen (`audio_frozen=True`)

### CLI Usage Example

Run the lip dubbing pipeline from the command line:

```bash
python -m ltx_pipelines.lipdub \
    --distilled-checkpoint-path <distilled.ckpt> \
    --spatial-upsampler-path <spatial_up.ckpt> \
    --gemma-root <gemma_root> \
    --lora <ic_lora_path>:<strength> \
    --reference-video <path/to/source_video.mp4> \
    --prompt "A smiling person says hello" \
    --seed 42 \
    --height 720 \
    --width 1280 \
    --output-path ./output_lipdub.mp4

```

Internally, `LipDubPipeline.__call__` builds a `GaussianNoiser`, encodes prompts via `PromptEncoder`, and manages the `VideoUpsampler` between diffusion stages. Final output is multiplexed using `encode_video` from the media I/O utilities.

## LTX-2 Audio Rephrasing Pipeline (T2AOneStagePipeline)

For **audio rephrasing** (text-to-audio generation), the `T2AOneStagePipeline` provides a lightweight, audio-only solution that avoids loading video weights entirely.

### Core Architecture

Located in [`packages/ltx-pipelines/src/ltx_pipelines/t2a_one_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/t2a_one_stage.py), this pipeline is configured specifically for audio generation:

- **Model Configuration**: Uses `LTXAudioOnlyModelConfigurator` to load only audio weights, reducing memory footprint
- **Single-Stage Diffusion**: Performs one diffusion pass on audio latents using `DiffusionStage` with multimodal guidance
- **Guidance Mechanism**: Supports CFG (Classifier-Free Guidance), STG (Self-Guidance), and other modalities via `FactoryGuidedDenoiser` and `MultiModalGuiderFactory`

### Python API Implementation

Integrate the audio rephrasing pipeline programmatically:

```python
from ltx_pipelines.t2a_one_stage import T2AOneStagePipeline
from ltx_pipelines.utils.args import detect_checkpoint_path, default_1_stage_t2a_arg_parser
from ltx_pipelines.utils.constants import detect_params
from ltx_pipelines.utils.media_io import encode_audio

# Detect checkpoint and parameters

checkpoint = detect_checkpoint_path()
params = detect_params(checkpoint)

# Initialize pipeline with audio-only configuration

pipeline = T2AOneStagePipeline(
    checkpoint_path=checkpoint,
    gemma_root="/opt/gemma",
    loras=(),
)

# Generate rephrased audio

audio = pipeline(
    prompt="A calm voice says hello world",
    negative_prompt="",
    seed=1234,
    num_frames=500,
    frame_rate=24.0,
    num_inference_steps=50,
    audio_guider_params=MultiModalGuiderParams(
        cfg_scale=7.0,
        stg_scale=1.0,
        rescale_scale=0.0,
        modality_scale=1.0,
        skip_step=0,
        stg_blocks=0,
    ),
)

# Export to file

encode_audio(audio=audio, output_path="rephrased.wav")

```

The pipeline outputs raw audio through `AudioDecoder` (defined in [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py)), ready for encoding via `encode_audio`.

## Shared Components and Audio VAE

Both pipelines rely on common building blocks from [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py):

- **PromptEncoder**: Handles text conditioning for both pipelines
- **AudioDecoder**: Decodes audio latents into waveforms
- **DiffusionStage**: Core diffusion logic shared across modalities

The audio VAE implementation in [`packages/ltx-core/src/ltx_core/model/audio_vae/audio_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/audio_vae/audio_vae.py) provides the `vae_encode_audio` function used by `LipDubPipeline` to process reference audio, ensuring consistent latent representations across both pipelines.

## Summary

- **LipDubPipeline** ([`packages/ltx-pipelines/src/ltx_pipelines/lipdub.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/lipdub.py)) handles video+audio conditioning for lip-sync tasks using two-stage diffusion and frozen audio latents
- **T2AOneStagePipeline** ([`packages/ltx-pipelines/src/ltx_pipelines/t2a_one_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/t2a_one_stage.py)) provides lightweight text-to-audio generation using `LTXAudioOnlyModelConfigurator` and single-stage diffusion
- Both pipelines use `AudioConditionByReferenceLatent` for conditioning and share `AudioDecoder` from the utilities module
- The audio VAE in `ltx_core` standardizes audio latent encoding across the framework

## Frequently Asked Questions

### What is the difference between LipDubPipeline and T2AOneStagePipeline in LTX-2?

`LipDubPipeline` is a two-stage video generation pipeline that accepts video and audio references to produce lip-synced output, while `T2AOneStagePipeline` is an audio-only pipeline for pure text-to-audio generation without video components. The former uses `LTXVideoModelConfigurator` and spatial upsamplers, whereas the latter uses `LTXAudioOnlyModelConfigurator` to minimize memory usage.

### How does LTX-2 handle audio conditioning in the lip dubbing pipeline?

The pipeline extracts audio from the reference video using `_encode_reference_audio_vae_latent`, encodes it through the audio VAE, and patches it into `AudioConditionByReferenceLatent` tokens. These tokens are injected into the diffusion process alongside video conditionings, with the audio latent remaining frozen during the upsampling stage to preserve speech content.

### Can I use T2AOneStagePipeline for video generation?

No. `T2AOneStagePipeline` explicitly configures the model using `LTXAudioOnlyModelConfigurator`, which excludes video weights entirely. For video generation, use `LipDubPipeline` or other video-specific pipelines in the `ltx_pipelines` package that load full video model weights.

### What checkpoint configuration is required for audio rephrasing in LTX-2?

Audio rephrasing requires a checkpoint compatible with `LTXAudioOnlyModelConfigurator`. The pipeline uses `detect_checkpoint_path()` and `detect_params()` to automatically configure the correct audio VAE and model architecture, loading only the audio components necessary for the `DiffusionStage` to perform single-pass generation.