# Understanding the LTX-2 Two-Stage Pipeline Architecture: A Deep Dive into Text-to-Video Generation

> Explore the LTX-2 two-stage pipeline architecture for text-to-video generation. Learn how it creates high-quality videos through low-resolution generation and refinement.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-15

---

**The LTX-2 text-to-video pipeline uses a two-stage architecture that first generates a low-resolution video at half the target resolution, then upsamples and refines it using a distilled LoRA for high-quality output.**

The LTX-2 repository from Lightricks implements one of the most sophisticated open-source text-to-video generation systems available today. This article examines the **LTX-2 two-stage pipeline architecture** implemented in [`packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py), breaking down how it orchestrates prompt encoding, diffusion, upsampling, and decoding to transform text prompts into high-fidelity video with synchronized audio.

## Core Pipeline Class: TI2VidTwoStagesPipeline

The **entry point** for the entire system is the `TI2VidTwoStagesPipeline` class, defined at lines 61-68 of the main pipeline file. This class wires together all sub-modules and exposes a callable interface that accepts prompts, conditioning images, and generation parameters.

```python
from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_core.loader import LoraPathStrengthAndSDOps

pipeline = TI2VidTwoStagesPipeline(
    model_paths=ModelPaths(),
    distilled_lora=[LoraPathStrengthAndSDOps("distilled_lora.ckpt", 0.8, None)],
    spatial_upsampler_path="spatial_upsampler.ckpt",
    loras=[],
)

```

The pipeline is designed for **modularity** — each component operates as a self-contained class, enabling easy replacement or fine-tuning of individual parts without disrupting the entire generation flow.

## Stage 1: Low-Resolution Video Generation

Stage 1 produces the initial video latent at **half the target resolution**. This stage involves five key steps:

### 1. Prompt Encoding with PromptEncoder

The `PromptEncoder` (instantiated at lines 89-97) converts text prompts into latent video and audio contexts that condition the diffusion process.

### 2. Image Conditioning with ImageConditioner

The `ImageConditioner` (lines 98-105) prepares optional conditioning images for the video encoder, enabling image-to-video generation workflows.

### 3. Diffusion Stage Loading

`DiffusionStage.from_checkpoint` (lines 36-46) loads the main transformer model and any user-specified LoRAs. This establishes the backbone for the denoising process.

### 4. Guided Denoising with FactoryGuidedDenoiser

The `FactoryGuidedDenoiser` (used at lines 47-60) combines prompt, video, and audio contexts with configurable multimodal guiders including:

- **CFG** (Classifier-Free Guidance)
- **STG** (Self-Attention Guidance)
- **Rescale** (noise rescaling)

### 5. Stage 1 Execution

The pipeline calls `self.stage_1(...)` (lines 66-70) to produce `video_state` and `audio_state` latents.

```python
video, audio, n_frames, _ = pipeline(
    prompt="A sunrise over a mountain lake",
    negative_prompt="low quality",
    seed=42,
    height=720,
    width=1280,
    frame_rate=30.0,
    num_inference_steps=50,
    images=[],  # optional conditioning images

)

```

## Stage 2: Spatial Upsampling and Refinement

Stage 2 transforms the half-resolution output into the final high-resolution video using a **distilled LoRA** specialized for quality refinement.

### Spatial Upsampling with VideoUpsampler

The `VideoUpsampler` (lines 105-112) enlarges the Stage 1 latent video by **2×**, bringing it to the target resolution.

### Distilled LoRA for High-Quality Refinement

The second `DiffusionStage` (lines 47-57) reuses the same base transformer but adds a **distilled LoRA** (`distilled_lora`) optimized for refinement tasks. This architectural choice separates the coarse generation (Stage 1) from fine detail enhancement (Stage 2).

### Simple Denoising and Final Decoding

- `SimpleDenoiser` (lines 89-92) runs the second diffusion pass using the upscaled latent as initialization
- `VideoDecoder` (lines 13-21) and `AudioDecoder` (lines 21-28) convert refined latents to final video frames and audio waveform

## Key Architectural Components

| Component | Purpose | Location |
|-----------|---------|----------|
| `AllocatorTrimStrategy` | GPU memory trimming during inference | [`ltx_core/allocator_trim_strategy.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/allocator_trim_strategy.py) |
| `DiffVAEMode` | VAE execution mode (eager or chunked) | [`ltx_core/model/video_vae/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/transformer.py) |
| `Tiling & Auto-Tiling` | Memory-efficient handling of large frames | [`ltx_core/model/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae.py) |
| `Multimodal Guider Factory` | Per-modality CFG/STG/Rescale guiders | [`ltx_core/components/guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/components/guiders.py) |
| `DurationPredictor` | Automatic frame count prediction | [`ltx_core/duration_head/duration_head.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/duration_head/duration_head.py) |

## Complete Runtime Flow

1. **CLI parsing** — `resolve_cli_params()` builds the argument namespace
2. **Pipeline construction** — `TI2VidTwoStagesPipeline(...)` supplies model paths, LoRAs, device, and quantization settings
3. **Prompt & conditioning** — `prompt_encoder` and `image_conditioner` produce latent contexts
4. **Stage 1 diffusion** — `stage_1` runs guided denoising on half-resolution latents
5. **Upsampling** — `upsampler` enlarges the latent video
6. **Stage 2 diffusion** — `stage_2` refines using the distilled LoRA
7. **Decoding & encoding** — `video_decoder`, `audio_decoder`, and `encode_video` write the final MP4

## Command-Line Usage

The pipeline supports direct CLI invocation matching the `main()` function structure:

```bash
python -m ltx_pipelines.ti2vid_two_stages \
    --prompt "A futuristic city at night" \
    --negative_prompt "blurred" \
    --height 720 --width 1280 \
    --frame-rate 24 \
    --num-inference-steps 60 \
    --model-paths /path/to/models \
    --distilled-lora /path/to/distilled_lora.ckpt \
    --spatial-upsampler-path /path/to/upsampler.ckpt \
    --output-path result.mp4

```

## Source File Reference

| File | Purpose |
|------|---------|
| [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py) | Main two-stage pipeline implementation |
| [`utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/blocks.py) | Component definitions (`PromptEncoder`, `ImageConditioner`, `VideoUpsampler`, `DiffusionStage`, etc.) |
| [`utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/args.py) | CLI parsers and `MultiModalGuiderParams` |
| [`utils/media_io/encode.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/media_io/encode.py) | MP4 encoding utilities |
| [`utils/model_paths.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/model_paths.py) | Checkpoint location management |
| [`ltx_core/components/guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/components/guiders.py) | Multimodal guidance factory |
| [`ltx_core/duration_head/duration_head.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/duration_head/duration_head.py) | Automatic duration prediction |

## Summary

- The **LTX-2 two-stage pipeline** separates generation into coarse low-resolution synthesis (Stage 1) and high-quality refinement (Stage 2)
- **Half-resolution first pass** reduces computational cost; **2× upsampling plus distilled LoRA** recovers quality
- **Modular design** enables component substitution without full pipeline rewrites
- **Multimodal guiders** (CFG, STG, Rescale) operate independently on video and audio conditioning
- **Multi-GPU and auto-tiling support** scales from consumer GPUs to datacenter hardware

## Frequently Asked Questions

### What is the purpose of the distilled LoRA in Stage 2?

The distilled LoRA specializes in refining upscaled latents for high-quality output. According to the LTX-2 source code, this LoRA is loaded as an additional component on top of the base transformer in the second `DiffusionStage` (lines 47-57 of [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py)), distinguishing it from user-provided LoRAs used in Stage 1.

### How does the pipeline handle memory constraints for large videos?

The architecture implements multiple memory optimization strategies: `AllocatorTrimStrategy` controls GPU memory trimming, `DiffVAEMode` supports chunked VAE execution, and auto-tiling in [`ltx_core/model/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae.py) automatically adapts to hardware limits by splitting large frames into manageable tiles.

### Can I use custom images to condition the video generation?

Yes. The `ImageConditioner` component (lines 98-105) accepts optional conditioning images through the `images` parameter. Pass a list of images to the pipeline call, and they will be encoded and integrated into the video generation context.

### What happens when num_frames is set to "auto"?

The `DurationPredictor` in [`ltx_core/duration_head/duration_head.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/duration_head/duration_head.py) automatically predicts the optimal number of frames based on the prompt and other generation parameters. This eliminates manual frame count specification while maintaining temporal coherence appropriate to the content.