How LTX‑2's Two‑Stage Pipeline Handles Latent Upsampling Between Stage I and Stage II

LTX‑2 performs latent upsampling between Stage I and Stage II using a dedicated VideoUpsampler that applies encoder‑guided normalization, a 2× spatial pixel‑shuffle upsample, and renormalization before passing the upscaled latent to Stage II for full‑resolution refinement.

The two‑stage pipeline in LTX‑2 is designed to generate high‑quality video efficiently: Stage I produces a low‑resolution draft, then a specialized upsampling module bridges the resolution gap before Stage II refines the result. This article explains exactly how the latent upsampling mechanism works, based on the implementation in Lightricks/LTX-2.


Overview of the Two‑Stage Architecture

LTX‑2 provides two main two‑stage pipelines: A2VidPipelineTwoStage (audio‑to‑video) and TI2VidTwoStagesPipeline (text‑to‑video). Both follow the same fundamental pattern:

  1. Stage I: Generate video at half resolution using full diffusion steps.
  2. Upsampling: Scale the latent spatially by 2× using VideoUpsampler.
  3. Stage II: Refine the upscaled latent at full resolution with a short diffusion schedule and distilled LoRA.

This separation reduces computational cost—Stage I operates on fewer pixels—while preserving quality through the learned upsampler and Stage II refinement.


Stage I: Generating the Low‑Resolution Latent

In A2VidPipelineTwoStage.__call__ at packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py (lines 24‑40), the pipeline first prepares Stage I inputs:

stage_1_output_shape = OutputShapeParams(
    frame_rate=frame_rate,
    height=height // 2,   # Half target height

    width=width // 2,     # Half target width

    num_frames=num_frames,
)

video_state = self.stage_1(
    prompt=prompt,
    negative_prompt=negative_prompt,
    output_shape=stage_1_output_shape,
    seed=seed,
    sigmas=stage_1_sigmas,
    # ... additional parameters

)

The video_state.latent produced here has shape [B, C, F, H/2, W/2]—half the spatial resolution of the final output. This latent encapsulates the coarse structure of the generated video.


The VideoUpsampler: Core Upsampling Component

The VideoUpsampler is instantiated once in the pipeline constructor (lines 126‑133 of the same file):

self.upsampler = VideoUpsampler(
    video_vae_path=model_paths.video_vae(),
    spatial_upsampler_path=spatial_upsampler_path,
    # Loads encoder + LatentUpsampler internally

)

As implemented in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py, VideoUpsampler combines two key components:

  • VideoEncoder: Computes per‑channel statistics for normalization.
  • LatentUpsampler: Performs the actual spatial upsampling via pixel‑shuffle.

LatentUpsampler Implementation

The upsampling network resides in packages/ltx-core/src/ltx_core/model/upsampler/model.py (lines 11‑34). Key characteristics:

class LatentUpsampler(nn.Module):
    def __init__(
        self,
        in_channels: int = 128,
        mid_channels: int = 512,
        num_blocks_per_stage: int = 4,
        dims: int = 2,                    # Spatial only

        spatial_upsample: bool = True,
        temporal_upsample: bool = False,  # Keep temporal dim fixed

        spatial_scale: float = 2.0,
    ):
        ...
        if spatial_upsample:
            self.spatial_upsampler = PixelShuffleND(2)  # 2× spatial

The PixelShuffleND(2) operation rearranges channels to achieve resolution doubling without checkerboard artifacts—critical for latent space upsampling.


The Upsampling Sequence Between Stages

Step‑by‑Step Flow

After Stage I completes, the pipeline executes the upsampling at lines 60‑62:


# Extract single video from batch (batch size 1 assumed)

upscaled_video_latent = self.upsampler(video_state.latent[:1])

The VideoUpsampler.__call__ method (via upsample_video at lines 30‑44 of model.py) performs:

  1. Denormalization: Scale latent using encoder's per‑channel mean/std.
  2. Upsampling: Pass through LatentUpsampler (conv blocks + pixel‑shuffle).
  3. Renormalization: Restore to VAE‑expected distribution.

This preserves the statistical properties required by the decoder while introducing high‑frequency details.

Passing to Stage II

The upscaled latent feeds directly into Stage II (lines 86‑92):

upscale_output_shape = OutputShapeParams(
    frame_rate=frame_rate,
    height=height,      # Full target resolution

    width=width,
    num_frames=num_frames,
)

final_video_state = self.stage_2(
    prompt=prompt,
    negative_prompt=negative_prompt,
    output_shape=upscale_output_shape,
    initial_latent=upscaled_video_latent,  # Upsampled starting point

    sigmas=stage_2_sigmas,                  # Short schedule

    # ... additional parameters

)

Stage II runs with a distilled LoRA and abbreviated diffusion steps, refining details while the audio conditioning remains frozen from Stage I.


Why This Upsampling Design Works

The LTX‑2 latent upsampling succeeds through several carefully coordinated properties:

  • Distribution preservation: Operating on denormalized latents prevents the upsampler from learning difficult residual corrections.
  • Exact scale matching: The 2× spatial factor precisely inverts the VAE encoder's downsampling, ensuring decoder compatibility.
  • Temporal stability: By setting temporal_upsample=False, motion coherence from Stage I is preserved—only spatial detail is enhanced.
  • Lightweight refinement: Stage II's short schedule (fewer steps, distilled model) efficiently sharpens the upsampled result without full re‑generation cost.

These choices reflect the VAE architecture: the encoder reduces spatial dimensions by 2× per level, so the upsampler's output aligns with what the decoder expects at full resolution.


Practical Code Examples

Running the Complete Two‑Stage Pipeline

from ltx_pipelines import A2VidPipelineTwoStage, ModelPaths, default_2_stage_arg_parser

# Configure via argument parser

parser = default_2_stage_arg_parser(params={})
args = parser.parse_args([
    "--model-paths", "path/to/monolith",
    "--distilled-lora", "path/to/distilled_lora.pt",
    "--spatial-upsampler-path", "path/to/upsampler.pt",
    "--prompt", "Ocean waves crashing on rocky cliffs",
    "--height", "512",
    "--width", "768",
    "--num-frames", "32",
    "--audio-path", "ambient_ocean.wav",
    "--output-path", "output.mp4"
])

# Initialize pipeline (upsampler loaded here)

pipeline = A2VidPipelineTwoStage(
    model_paths=args.model_paths,
    distilled_lora=args.distilled_lora,
    spatial_upsampler_path=args.spatial_upsampler_path,
)

# Run two-stage generation with automatic upsampling

video_iter, audio, tiling = pipeline(
    prompt=args.prompt,
    height=args.height,
    width=args.width,
    num_frames=args.num_frames,
    audio_path=args.audio_path,
    # Stage I at 256×384, Stage II at 512×768

)

Manual Latent Upsampling

For inspection or custom pipelines, access the upsampler directly:

import torch
from ltx_core.model.upsampler.model import LatentUpsampler, upsample_video
from ltx_core.model.video_vae import VideoEncoder

# Simulate Stage I output: [B, C, F, H, W] = [1, 128, 32, 256, 384]

stage_1_latent = torch.randn(1, 128, 32, 256, 384, device="cuda")

# Load components (normally managed by VideoUpsampler)

encoder = VideoEncoder.from_pretrained("path/to/video_vae").cuda().eval()

upsampler = LatentUpsampler(
    in_channels=128,
    mid_channels=512,
    num_blocks_per_stage=4,
    dims=2,
    spatial_upsample=True,
    temporal_upsample=False,
    spatial_scale=2.0,
).cuda().eval()

# Execute full upsampling sequence

with torch.no_grad():
    stage_2_latent = upsample_video(stage_1_latent, encoder, upsampler)

print(f"Stage I:  {stage_1_latent.shape}")   # [1, 128, 32, 256, 384]

print(f"Stage II: {stage_2_latent.shape}")   # [1, 128, 32, 512, 768]

Key Implementation Files

File Location Purpose
a2vid_two_stage.py packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py Orchestrates Stage I, upsampling, and Stage II for audio‑conditioned video
ti2vid_two_stages.py packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py Equivalent text‑to‑video two‑stage pipeline
blocks.py packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py VideoUpsampler class definition and integration
model.py (upsampler) packages/ltx-core/src/ltx_core/model/upsampler/model.py LatentUpsampler network with PixelShuffleND
video_vae.py packages/ltx-core/src/ltx_core/model/video_vae/video_vae.py Encoder/decoder providing normalization statistics

Summary

  • Stage I generates at half resolution by halving height and width parameters before diffusion.
  • VideoUpsampler bridges stages using encoder statistics for normalization-aware upsampling.
  • LatentUpsampler applies 2× spatial pixel‑shuffle while preserving temporal dimensions.
  • Stage II receives the upscaled latent as initial_latent and refines at full resolution with distilled sampling.
  • The design ensures VAE compatibility by matching the encoder's spatial downsampling factor exactly.

Frequently Asked Questions

How does latent upsampling differ from image spatial upsampling in LTX‑2?

Latent upsampling operates in compressed representation space (128 channels) rather than RGB pixels. The VideoUpsampler uses encoder-derived statistics to denormalize before upsampling, ensuring the LatentUpsampler works on properly scaled values. Pixel‑shuffle rearranges spatial information across channels, then renormalization prepares the result for Stage II's diffusion process.

Why does Stage I use half resolution instead of full resolution?

Half resolution reduces compute and memory during the lengthy initial generation. Stage I runs the full diffusion schedule (e.g., 30 steps), so operating on ¼ the pixels (½H × ½W) significantly accelerates this phase. The learned upsampler and efficient Stage II refinement recover quality without repeating full‑resolution diffusion.

Can the upsampling factor be changed from 2×?

The current LatentUpsampler in ltx_core/model/upsampler/model.py hardcodes spatial_scale=2.0 with PixelShuffleND(2). The architecture supports configuration, but released checkpoints and pipelines assume 2× upsampling to match the VAE's fixed downsampling ratios. Modifying this would require retraining the upsampler and adjusting Stage II expectations.

Is temporal upsampling ever used between stages?

No—the two‑stage pipelines set temporal_upsample=False. Stage I generates the final frame count, and the upsampler preserves this dimension. Temporal extension would require a different pipeline architecture or iterative generation with conditioning on previous frames.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →