How LTX‑2's Two‑Stage Pipeline Handles Latent Upsampling Between Stage I and Stage II
LTX‑2 performs latent upsampling between Stage I and Stage II using a dedicated VideoUpsampler that applies encoder‑guided normalization, a 2× spatial pixel‑shuffle upsample, and renormalization before passing the upscaled latent to Stage II for full‑resolution refinement.
The two‑stage pipeline in LTX‑2 is designed to generate high‑quality video efficiently: Stage I produces a low‑resolution draft, then a specialized upsampling module bridges the resolution gap before Stage II refines the result. This article explains exactly how the latent upsampling mechanism works, based on the implementation in Lightricks/LTX-2.
Overview of the Two‑Stage Architecture
LTX‑2 provides two main two‑stage pipelines: A2VidPipelineTwoStage (audio‑to‑video) and TI2VidTwoStagesPipeline (text‑to‑video). Both follow the same fundamental pattern:
- Stage I: Generate video at half resolution using full diffusion steps.
- Upsampling: Scale the latent spatially by 2× using
VideoUpsampler. - Stage II: Refine the upscaled latent at full resolution with a short diffusion schedule and distilled LoRA.
This separation reduces computational cost—Stage I operates on fewer pixels—while preserving quality through the learned upsampler and Stage II refinement.
Stage I: Generating the Low‑Resolution Latent
In A2VidPipelineTwoStage.__call__ at packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py (lines 24‑40), the pipeline first prepares Stage I inputs:
stage_1_output_shape = OutputShapeParams(
frame_rate=frame_rate,
height=height // 2, # Half target height
width=width // 2, # Half target width
num_frames=num_frames,
)
video_state = self.stage_1(
prompt=prompt,
negative_prompt=negative_prompt,
output_shape=stage_1_output_shape,
seed=seed,
sigmas=stage_1_sigmas,
# ... additional parameters
)
The video_state.latent produced here has shape [B, C, F, H/2, W/2]—half the spatial resolution of the final output. This latent encapsulates the coarse structure of the generated video.
The VideoUpsampler: Core Upsampling Component
The VideoUpsampler is instantiated once in the pipeline constructor (lines 126‑133 of the same file):
self.upsampler = VideoUpsampler(
video_vae_path=model_paths.video_vae(),
spatial_upsampler_path=spatial_upsampler_path,
# Loads encoder + LatentUpsampler internally
)
As implemented in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py, VideoUpsampler combines two key components:
- VideoEncoder: Computes per‑channel statistics for normalization.
- LatentUpsampler: Performs the actual spatial upsampling via pixel‑shuffle.
LatentUpsampler Implementation
The upsampling network resides in packages/ltx-core/src/ltx_core/model/upsampler/model.py (lines 11‑34). Key characteristics:
class LatentUpsampler(nn.Module):
def __init__(
self,
in_channels: int = 128,
mid_channels: int = 512,
num_blocks_per_stage: int = 4,
dims: int = 2, # Spatial only
spatial_upsample: bool = True,
temporal_upsample: bool = False, # Keep temporal dim fixed
spatial_scale: float = 2.0,
):
...
if spatial_upsample:
self.spatial_upsampler = PixelShuffleND(2) # 2× spatial
The PixelShuffleND(2) operation rearranges channels to achieve resolution doubling without checkerboard artifacts—critical for latent space upsampling.
The Upsampling Sequence Between Stages
Step‑by‑Step Flow
After Stage I completes, the pipeline executes the upsampling at lines 60‑62:
# Extract single video from batch (batch size 1 assumed)
upscaled_video_latent = self.upsampler(video_state.latent[:1])
The VideoUpsampler.__call__ method (via upsample_video at lines 30‑44 of model.py) performs:
- Denormalization: Scale latent using encoder's per‑channel mean/std.
- Upsampling: Pass through
LatentUpsampler(conv blocks + pixel‑shuffle). - Renormalization: Restore to VAE‑expected distribution.
This preserves the statistical properties required by the decoder while introducing high‑frequency details.
Passing to Stage II
The upscaled latent feeds directly into Stage II (lines 86‑92):
upscale_output_shape = OutputShapeParams(
frame_rate=frame_rate,
height=height, # Full target resolution
width=width,
num_frames=num_frames,
)
final_video_state = self.stage_2(
prompt=prompt,
negative_prompt=negative_prompt,
output_shape=upscale_output_shape,
initial_latent=upscaled_video_latent, # Upsampled starting point
sigmas=stage_2_sigmas, # Short schedule
# ... additional parameters
)
Stage II runs with a distilled LoRA and abbreviated diffusion steps, refining details while the audio conditioning remains frozen from Stage I.
Why This Upsampling Design Works
The LTX‑2 latent upsampling succeeds through several carefully coordinated properties:
- Distribution preservation: Operating on denormalized latents prevents the upsampler from learning difficult residual corrections.
- Exact scale matching: The 2× spatial factor precisely inverts the VAE encoder's downsampling, ensuring decoder compatibility.
- Temporal stability: By setting
temporal_upsample=False, motion coherence from Stage I is preserved—only spatial detail is enhanced. - Lightweight refinement: Stage II's short schedule (fewer steps, distilled model) efficiently sharpens the upsampled result without full re‑generation cost.
These choices reflect the VAE architecture: the encoder reduces spatial dimensions by 2× per level, so the upsampler's output aligns with what the decoder expects at full resolution.
Practical Code Examples
Running the Complete Two‑Stage Pipeline
from ltx_pipelines import A2VidPipelineTwoStage, ModelPaths, default_2_stage_arg_parser
# Configure via argument parser
parser = default_2_stage_arg_parser(params={})
args = parser.parse_args([
"--model-paths", "path/to/monolith",
"--distilled-lora", "path/to/distilled_lora.pt",
"--spatial-upsampler-path", "path/to/upsampler.pt",
"--prompt", "Ocean waves crashing on rocky cliffs",
"--height", "512",
"--width", "768",
"--num-frames", "32",
"--audio-path", "ambient_ocean.wav",
"--output-path", "output.mp4"
])
# Initialize pipeline (upsampler loaded here)
pipeline = A2VidPipelineTwoStage(
model_paths=args.model_paths,
distilled_lora=args.distilled_lora,
spatial_upsampler_path=args.spatial_upsampler_path,
)
# Run two-stage generation with automatic upsampling
video_iter, audio, tiling = pipeline(
prompt=args.prompt,
height=args.height,
width=args.width,
num_frames=args.num_frames,
audio_path=args.audio_path,
# Stage I at 256×384, Stage II at 512×768
)
Manual Latent Upsampling
For inspection or custom pipelines, access the upsampler directly:
import torch
from ltx_core.model.upsampler.model import LatentUpsampler, upsample_video
from ltx_core.model.video_vae import VideoEncoder
# Simulate Stage I output: [B, C, F, H, W] = [1, 128, 32, 256, 384]
stage_1_latent = torch.randn(1, 128, 32, 256, 384, device="cuda")
# Load components (normally managed by VideoUpsampler)
encoder = VideoEncoder.from_pretrained("path/to/video_vae").cuda().eval()
upsampler = LatentUpsampler(
in_channels=128,
mid_channels=512,
num_blocks_per_stage=4,
dims=2,
spatial_upsample=True,
temporal_upsample=False,
spatial_scale=2.0,
).cuda().eval()
# Execute full upsampling sequence
with torch.no_grad():
stage_2_latent = upsample_video(stage_1_latent, encoder, upsampler)
print(f"Stage I: {stage_1_latent.shape}") # [1, 128, 32, 256, 384]
print(f"Stage II: {stage_2_latent.shape}") # [1, 128, 32, 512, 768]
Key Implementation Files
| File | Location | Purpose |
|---|---|---|
a2vid_two_stage.py |
packages/ltx-pipelines/src/ltx_pipelines/a2vid_two_stage.py |
Orchestrates Stage I, upsampling, and Stage II for audio‑conditioned video |
ti2vid_two_stages.py |
packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py |
Equivalent text‑to‑video two‑stage pipeline |
blocks.py |
packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py |
VideoUpsampler class definition and integration |
model.py (upsampler) |
packages/ltx-core/src/ltx_core/model/upsampler/model.py |
LatentUpsampler network with PixelShuffleND |
video_vae.py |
packages/ltx-core/src/ltx_core/model/video_vae/video_vae.py |
Encoder/decoder providing normalization statistics |
Summary
- Stage I generates at half resolution by halving height and width parameters before diffusion.
VideoUpsamplerbridges stages using encoder statistics for normalization-aware upsampling.LatentUpsamplerapplies 2× spatial pixel‑shuffle while preserving temporal dimensions.- Stage II receives the upscaled latent as
initial_latentand refines at full resolution with distilled sampling. - The design ensures VAE compatibility by matching the encoder's spatial downsampling factor exactly.
Frequently Asked Questions
How does latent upsampling differ from image spatial upsampling in LTX‑2?
Latent upsampling operates in compressed representation space (128 channels) rather than RGB pixels. The VideoUpsampler uses encoder-derived statistics to denormalize before upsampling, ensuring the LatentUpsampler works on properly scaled values. Pixel‑shuffle rearranges spatial information across channels, then renormalization prepares the result for Stage II's diffusion process.
Why does Stage I use half resolution instead of full resolution?
Half resolution reduces compute and memory during the lengthy initial generation. Stage I runs the full diffusion schedule (e.g., 30 steps), so operating on ¼ the pixels (½H × ½W) significantly accelerates this phase. The learned upsampler and efficient Stage II refinement recover quality without repeating full‑resolution diffusion.
Can the upsampling factor be changed from 2×?
The current LatentUpsampler in ltx_core/model/upsampler/model.py hardcodes spatial_scale=2.0 with PixelShuffleND(2). The architecture supports configuration, but released checkpoints and pipelines assume 2× upsampling to match the VAE's fixed downsampling ratios. Modifying this would require retraining the upsampler and adjusting Stage II expectations.
Is temporal upsampling ever used between stages?
No—the two‑stage pipelines set temporal_upsample=False. Stage I generates the final frame count, and the upsampler preserves this dimension. Temporal extension would require a different pipeline architecture or iterative generation with conditioning on previous frames.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →