# How the Two-Stage Generation Process Works in LTX-2: A Technical Deep Dive

> Explore the LTX-2 two-stage generation process. Discover how it creates low-res video then refines it with distilled LoRA for impressive quality gains and modest inference time.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-06-20

---

**LTX-2 uses a two-stage diffusion workflow that first generates a low-resolution video and then refines it to full resolution using a distilled LoRA, trading modest inference time for significant quality gains.**

The two-stage generation process in LTX-2 is the core architecture behind Lightricks' text-to-video pipelines, implemented primarily in [`packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py). This approach splits generation into a fast coarse pass at half-resolution followed by an upsampled refinement stage, enabling high-fidelity output without the computational cost of full-resolution diffusion from scratch.

## Architecture Overview

The pipeline orchestration lives in `TI2VidTwoStagesPipeline.__call__` (lines 103-120 in [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py)), which coordinates conditioning preparation, dual-stage denoising, and final decoding. The workflow leverages reusable building blocks from [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py), including `DiffusionStage`, `PromptEncoder`, and `VideoUpsampler`.

## Stage 1: Half-Resolution Generation

The first stage produces a coarse video latent at half the target resolution, utilizing classifier-free guidance (CFG) for directional control.

### Prompt Encoding and Conditioning

Generation begins with text and image conditioning:

- **`PromptEncoder`** (lines 28-34) converts the user prompt and optional negative prompt into video and audio embeddings.
- **`ImageConditioner`** (lines 45-54) constructs image conditionings that are injected into the diffusion model during denoising.

### Denoising at Half Resolution

The pipeline computes the target half-resolution shape by dividing width and height by 2 (lines 138-144). Key steps include:

1. **Image conditioning preparation** – `combined_image_conditionings` creates resolution-specific conditionings for the reduced size (lines 145-152).
2. **Schedule generation** – `LTX2Scheduler` generates the noise schedule (`sigmas`) if not user-provided (lines 156-158).
3. **Guided denoising** – A **`FactoryGuidedDenoiser`** incorporates CFG and optional style-transfer guidance (lines 160-172).
4. **Diffusion execution** – The first **`DiffusionStage`** (defined in [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py) lines 91-105) runs the Euler sampler loop to produce the half-resolution latent video and audio (lines 160-182).

## Stage 2: Upsampling and Refinement

The second stage upsamples the coarse latent and applies a distilled LoRA to refine details at full resolution without CFG.

### Latent Upsampling

The half-resolution latent is upsampled 2× by the **`VideoUpsampler`**, which combines a video encoder with a spatial upsampler (lines 184-191 in [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py)).

### Refinement with Distilled LoRA

Full-resolution image conditionings are recomputed for the final target size (lines 192-200). This stage differs from Stage 1 in two critical ways:

- **Simplified denoising** – A **`SimpleDenoiser`** runs without classifier-free guidance (lines 199-212).
- **Distilled LoRA** – The `distilled_lora` parameter (passed to `self.stage_2`) applies lightweight refinement weights to add fine details without reloading the full base model.

The second **`DiffusionStage`** completes the remaining diffusion steps at the target resolution.

## Decoding and Output

Finally, the refined latents are converted to pixel-space outputs:

- **`VideoDecoder`** decodes the video latent into viewable frames.
- **`AudioDecoder`** decodes the audio latent into a waveform.

Both decoders are invoked at lines 220-222 in [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py).

## Implementation Example

The following example demonstrates running the complete two-stage pipeline:

```python
from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline
from ltx_pipelines.utils.args import detect_params, default_2_stage_arg_parser
from ltx_core.components.guiders import MultiModalGuiderParams

# 1️⃣ Detect checkpoint & parsing params

checkpoint = "models/ltx-2.3-22b-dev.safetensors"
params = detect_params(checkpoint)

# 2️⃣ Build the pipeline (paths to upsampler & Gemma text‑encoder root)

pipeline = TI2VidTwoStagesPipeline(
    checkpoint_path=checkpoint,
    distilled_lora=[("models/distilled_lora.safetensors", 1.0, None)],
    spatial_upsampler_path="models/ltx-2.3-spatial-upscaler-x2-1.1.safetensors",
    gemma_root="gemma-3",
    loras=[],
)

# 3️⃣ Run the two‑stage generation

video, audio = pipeline(
    prompt="A sunny beach with palm trees swaying in the wind",
    negative_prompt="low quality, blurry",
    seed=42,
    height=720,
    width=1280,
    num_frames=24,
    frame_rate=24.0,
    num_inference_steps=30,
    video_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
    audio_guider_params=MultiModalGuiderParams(cfg_scale=5.0),
    images=[],                # optional conditioning images

)

# 4️⃣ Save the result (requires ffmpeg)

from ltx_pipelines.utils.media_io import encode_video
encode_video(video, fps=24.0, audio=audio, output_path="output.mp4")

```

## Summary

- **Two-stage design** – Stage 1 generates at half-resolution (width/2, height/2) using full CFG, while Stage 2 upsamples 2× and refines with a distilled LoRA.
- **Performance trade-off** – The architecture adds modest inference time compared to single-stage generation but delivers substantial quality improvements.
- **Key files** – [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py) contains the orchestration, [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py) provides the `DiffusionStage` and `VideoUpsampler` utilities, and [`ti2vid_two_stages_hq.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages_hq.py) offers a high-quality variant using a second-order sampler.
- **Stage distinctions** – Stage 1 uses `FactoryGuidedDenoiser` with CFG; Stage 2 uses `SimpleDenoiser` without CFG and relies on the distilled LoRA for detail enhancement.

## Frequently Asked Questions

### What resolution does Stage 1 generate in LTX-2?

Stage 1 generates video at half the target resolution, computing the latent shape by dividing the requested width and height by 2 (see [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py) lines 138-144). This coarse latent is later upsampled 2× by the `VideoUpsampler` at the beginning of Stage 2.

### Why does Stage 2 use a distilled LoRA instead of the full model?

Stage 2 uses a **distilled LoRA** (passed via the `distilled_lora` parameter) to refine the upsampled latent because it is computationally cheaper than running the full base model at high resolution. The LoRA adds fine details and corrects artifacts introduced during upsampling without requiring the memory and compute of a full second diffusion pass.

### How does the two-stage process affect inference time?

The two-stage generation process trades a **modest increase in total inference time** for a **large quality boost**. While running two diffusion stages requires more steps than a single pass, the first stage operates at half-resolution (faster per-step), and the second stage uses a lightweight LoRA rather than the full model weights, keeping the overhead manageable compared to full-resolution generation from scratch.

### What is the difference between TI2VidTwoStagesPipeline and the HQ variant?

`TI2VidTwoStagesPipeline` (in [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py)) uses a standard Euler sampler for both stages. The HQ variant in [`ti2vid_two_stages_hq.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages_hq.py) implements a **second-order sampler** (`res_2s`) that can achieve higher quality with fewer inference steps, making it suitable for production environments where quality is prioritized over raw speed.