# How LTX‑2's Two‑Stage Pipeline Handles Latent Upsampling Between Stage I and Stage II

> Discover how LTX-2's two-stage pipeline manages latent upsampling between stages. Learn about encoder-guided normalization and pixel-shuffle upsampling for efficient refinement.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: internals
- Published: 2026-08-18

---

**LTX‑2 performs latent upsampling between Stage I and Stage II using a dedicated `VideoUpsampler` that applies encoder‑guided normalization, a 2× spatial pixel‑shuffle upsample, and renormalization before passing the upscaled latent to Stage II for full‑resolution refinement.**

The **two‑stage pipeline** in LTX‑2 is designed to generate high‑quality video efficiently: Stage I produces a low‑resolution draft, then a specialized upsampling module bridges the resolution gap before Stage II refines the result. This article explains exactly how the **latent upsampling** mechanism works, based on the implementation in `Lightricks/LTX-2`.

---

## Overview of the Two‑Stage Architecture

LTX‑2 provides two main two‑stage pipelines: `A2VidPipelineTwoStage` (audio‑to‑video) and `TI2VidTwoStagesPipeline` (text‑to‑video). Both follow the same fundamental pattern:

1. **Stage I**: Generate video at **half resolution** using full diffusion steps.
2. **Upsampling**: Scale the latent spatially by **2×** using `VideoUpsampler`.
3. **Stage II**: Refine the upscaled latent at **full resolution** with a short diffusion schedule and distilled LoRA.

This separation reduces computational cost—Stage I operates on fewer pixels—while preserving quality through the learned upsampler and Stage II refinement.

---

## Stage I: Generating the Low‑Resolution Latent

In `A2VidPipelineTwoStage.__call__` at [`packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py) (lines 24‑40), the pipeline first prepares Stage I inputs:

```python
stage_1_output_shape = OutputShapeParams(
    frame_rate=frame_rate,
    height=height // 2,   # Half target height

    width=width // 2,     # Half target width

    num_frames=num_frames,
)

video_state = self.stage_1(
    prompt=prompt,
    negative_prompt=negative_prompt,
    output_shape=stage_1_output_shape,
    seed=seed,
    sigmas=stage_1_sigmas,
    # ... additional parameters

)

```

The `video_state.latent` produced here has shape `[B, C, F, H/2, W/2]`—half the spatial resolution of the final output. This latent encapsulates the coarse structure of the generated video.

---

## The VideoUpsampler: Core Upsampling Component

The `VideoUpsampler` is instantiated once in the pipeline constructor (lines 126‑133 of the same file):

```python
self.upsampler = VideoUpsampler(
    video_vae_path=model_paths.video_vae(),
    spatial_upsampler_path=spatial_upsampler_path,
    # Loads encoder + LatentUpsampler internally

)

```

As implemented in [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py), `VideoUpsampler` combines two key components:

- **VideoEncoder**: Computes per‑channel statistics for normalization.
- **LatentUpsampler**: Performs the actual spatial upsampling via pixel‑shuffle.

### LatentUpsampler Implementation

The upsampling network resides in [`packages/ltx-core/src/ltx_core/model/upsampler/model.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/upsampler/model.py) (lines 11‑34). Key characteristics:

```python
class LatentUpsampler(nn.Module):
    def __init__(
        self,
        in_channels: int = 128,
        mid_channels: int = 512,
        num_blocks_per_stage: int = 4,
        dims: int = 2,                    # Spatial only

        spatial_upsample: bool = True,
        temporal_upsample: bool = False,  # Keep temporal dim fixed

        spatial_scale: float = 2.0,
    ):
        ...
        if spatial_upsample:
            self.spatial_upsampler = PixelShuffleND(2)  # 2× spatial

```

The `PixelShuffleND(2)` operation rearranges channels to achieve resolution doubling without checkerboard artifacts—critical for latent space upsampling.

---

## The Upsampling Sequence Between Stages

### Step‑by‑Step Flow

After Stage I completes, the pipeline executes the upsampling at lines 60‑62:

```python

# Extract single video from batch (batch size 1 assumed)

upscaled_video_latent = self.upsampler(video_state.latent[:1])

```

The `VideoUpsampler.__call__` method (via `upsample_video` at lines 30‑44 of [`model.py`](https://github.com/Lightricks/LTX-2/blob/main/model.py)) performs:

1. **Denormalization**: Scale latent using encoder's per‑channel mean/std.
2. **Upsampling**: Pass through `LatentUpsampler` (conv blocks + pixel‑shuffle).
3. **Renormalization**: Restore to VAE‑expected distribution.

This preserves the statistical properties required by the decoder while introducing high‑frequency details.

### Passing to Stage II

The upscaled latent feeds directly into Stage II (lines 86‑92):

```python
upscale_output_shape = OutputShapeParams(
    frame_rate=frame_rate,
    height=height,      # Full target resolution

    width=width,
    num_frames=num_frames,
)

final_video_state = self.stage_2(
    prompt=prompt,
    negative_prompt=negative_prompt,
    output_shape=upscale_output_shape,
    initial_latent=upscaled_video_latent,  # Upsampled starting point

    sigmas=stage_2_sigmas,                  # Short schedule

    # ... additional parameters

)

```

Stage II runs with a **distilled LoRA** and abbreviated diffusion steps, refining details while the audio conditioning remains frozen from Stage I.

---

## Why This Upsampling Design Works

The LTX‑2 latent upsampling succeeds through several carefully coordinated properties:

- **Distribution preservation**: Operating on denormalized latents prevents the upsampler from learning difficult residual corrections.
- **Exact scale matching**: The 2× spatial factor precisely inverts the VAE encoder's downsampling, ensuring decoder compatibility.
- **Temporal stability**: By setting `temporal_upsample=False`, motion coherence from Stage I is preserved—only spatial detail is enhanced.
- **Lightweight refinement**: Stage II's short schedule (fewer steps, distilled model) efficiently sharpens the upsampled result without full re‑generation cost.

These choices reflect the VAE architecture: the encoder reduces spatial dimensions by 2× per level, so the upsampler's output aligns with what the decoder expects at full resolution.

---

## Practical Code Examples

### Running the Complete Two‑Stage Pipeline

```python
from ltx_pipelines import A2VidPipelineTwoStage, ModelPaths, default_2_stage_arg_parser

# Configure via argument parser

parser = default_2_stage_arg_parser(params={})
args = parser.parse_args([
    "--model-paths", "path/to/monolith",
    "--distilled-lora", "path/to/distilled_lora.pt",
    "--spatial-upsampler-path", "path/to/upsampler.pt",
    "--prompt", "Ocean waves crashing on rocky cliffs",
    "--height", "512",
    "--width", "768",
    "--num-frames", "32",
    "--audio-path", "ambient_ocean.wav",
    "--output-path", "output.mp4"
])

# Initialize pipeline (upsampler loaded here)

pipeline = A2VidPipelineTwoStage(
    model_paths=args.model_paths,
    distilled_lora=args.distilled_lora,
    spatial_upsampler_path=args.spatial_upsampler_path,
)

# Run two-stage generation with automatic upsampling

video_iter, audio, tiling = pipeline(
    prompt=args.prompt,
    height=args.height,
    width=args.width,
    num_frames=args.num_frames,
    audio_path=args.audio_path,
    # Stage I at 256×384, Stage II at 512×768

)

```

### Manual Latent Upsampling

For inspection or custom pipelines, access the upsampler directly:

```python
import torch
from ltx_core.model.upsampler.model import LatentUpsampler, upsample_video
from ltx_core.model.video_vae import VideoEncoder

# Simulate Stage I output: [B, C, F, H, W] = [1, 128, 32, 256, 384]

stage_1_latent = torch.randn(1, 128, 32, 256, 384, device="cuda")

# Load components (normally managed by VideoUpsampler)

encoder = VideoEncoder.from_pretrained("path/to/video_vae").cuda().eval()

upsampler = LatentUpsampler(
    in_channels=128,
    mid_channels=512,
    num_blocks_per_stage=4,
    dims=2,
    spatial_upsample=True,
    temporal_upsample=False,
    spatial_scale=2.0,
).cuda().eval()

# Execute full upsampling sequence

with torch.no_grad():
    stage_2_latent = upsample_video(stage_1_latent, encoder, upsampler)

print(f"Stage I:  {stage_1_latent.shape}")   # [1, 128, 32, 256, 384]

print(f"Stage II: {stage_2_latent.shape}")   # [1, 128, 32, 512, 768]

```

---

## Key Implementation Files

| File | Location | Purpose |
|------|----------|---------|
| [`a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/a2vid_two_stage.py) | [`packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py) | Orchestrates Stage I, upsampling, and Stage II for audio‑conditioned video |
| [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py) | [`packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py) | Equivalent text‑to‑video two‑stage pipeline |
| [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py) | [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) | `VideoUpsampler` class definition and integration |
| [`model.py`](https://github.com/Lightricks/LTX-2/blob/main/model.py) (upsampler) | [`packages/ltx-core/src/ltx_core/model/upsampler/model.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/upsampler/model.py) | `LatentUpsampler` network with `PixelShuffleND` |
| [`video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/video_vae.py) | [`packages/ltx-core/src/ltx_core/model/video_vae/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/video_vae/video_vae.py) | Encoder/decoder providing normalization statistics |

---

## Summary

- **Stage I generates at half resolution** by halving height and width parameters before diffusion.
- **`VideoUpsampler` bridges stages** using encoder statistics for normalization-aware upsampling.
- **`LatentUpsampler` applies 2× spatial pixel‑shuffle** while preserving temporal dimensions.
- **Stage II receives the upscaled latent as `initial_latent`** and refines at full resolution with distilled sampling.
- The design ensures **VAE compatibility** by matching the encoder's spatial downsampling factor exactly.

---

## Frequently Asked Questions

### How does latent upsampling differ from image spatial upsampling in LTX‑2?

Latent upsampling operates in **compressed representation space** (128 channels) rather than RGB pixels. The `VideoUpsampler` uses encoder-derived statistics to denormalize before upsampling, ensuring the `LatentUpsampler` works on properly scaled values. Pixel‑shuffle rearranges spatial information across channels, then renormalization prepares the result for Stage II's diffusion process.

### Why does Stage I use half resolution instead of full resolution?

Half resolution reduces **compute and memory** during the lengthy initial generation. Stage I runs the full diffusion schedule (e.g., 30 steps), so operating on ¼ the pixels (½H × ½W) significantly accelerates this phase. The learned upsampler and efficient Stage II refinement recover quality without repeating full‑resolution diffusion.

### Can the upsampling factor be changed from 2×?

The current `LatentUpsampler` in [`ltx_core/model/upsampler/model.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/upsampler/model.py) hardcodes `spatial_scale=2.0` with `PixelShuffleND(2)`. The architecture supports configuration, but released checkpoints and pipelines assume **2× upsampling** to match the VAE's fixed downsampling ratios. Modifying this would require retraining the upsampler and adjusting Stage II expectations.

### Is temporal upsampling ever used between stages?

No—the two‑stage pipelines set `temporal_upsample=False`. Stage I generates the **final frame count**, and the upsampler preserves this dimension. Temporal extension would require a different pipeline architecture or iterative generation with conditioning on previous frames.