Using LTX-2 Spatial and Temporal Upsamplers for Higher Resolution Video Generation
LTX-2 generates higher-resolution video by first producing a low-resolution latent, then applying a spatial upsampler to increase frame size and an optional temporal upsampler to multiply frame count.
This two-stage architecture lets you scale both resolution and temporal fidelity without retraining the base model. The Lightricks/LTX-2 repository provides modular, swappable checkpoint-based upsamplers you can invoke via Python API or command line.
How LTX-2 Upsampling Works
LTX-2 separates video generation into distinct stages:
- Stage 1: Generate a low-resolution latent video using the base model.
- Stage 2: Apply upsamplers to increase quality before final decoding.
The pipeline uses standalone checkpoint files for upsamplers, meaning you can swap different scaling factors or custom-trained models without modifying code.
Spatial vs. Temporal Upsamplers
| Upsampler | Effect | When to Use |
|---|---|---|
| Spatial | Doubles width and height (×4 pixels per frame) | Every time you want sharper frames |
| Temporal | Doubles frame count per round | When you need smoother motion or longer clips |
Spatial upsampling runs once. Temporal upsampling can run multiple rounds—each round doubles the frame count.
Core Components from the LTX-2 Source Code
VideoUpsampler Class
The VideoUpsampler class in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py (lines 99–141) orchestrates spatial upsampling. It:
- Loads a video encoder from your base checkpoint
- Builds a latent upsampling model from your spatial upsampler checkpoint
- Calls
upsample_video()to merge encoder features with the upsampler output - Cleans up resources automatically
from ltx_pipelines.utils.blocks import VideoUpsampler
upsampler = VideoUpsampler(
checkpoint_path="/models/ltx-2.3-video-x2-1.0.safetensors",
upsampler_path="/models/ltx-2.3-spatial-upscaler-x2-1.0.safetensors",
dtype=torch.float16,
device=torch.device("cuda"),
)
high_res_latent = upsampler(latent) # 2× width, 2× height
TemporalUpsampler Integration
Temporal upsampling is handled separately in packages/ltx-pipelines/src/ltx_pipelines/dfr_pipeline.py (lines 220–250). The pipeline conditionally builds a temporal upsampler when --temporal-upsampler-path is provided and --temporal-upsample-rounds is greater than zero.
The LatentUpsamplerConfigurator in packages/ltx-core/src/ltx_core/model/upsampler/model_configurator.py provides the generic configuration logic used for both upsampler types.
Implementing Spatial Upsampling with VideoUpsampler
The minimal API for spatial upsampling requires only the base checkpoint and spatial upsampler checkpoint:
import torch
from ltx_pipelines.utils.blocks import VideoUpsampler
# Configure paths to your safetensor checkpoints
video_ckpt = "/models/ltx-2.3-video-x2-1.0.safetensors"
spatial_upsampler_ckpt = "/models/ltx-2.3-spatial-upscaler-x2-1.0.safetensors"
# Initialize upsampler with desired precision and device
upsampler = VideoUpsampler(
checkpoint_path=video_ckpt,
upsampler_path=spatial_upsampler_ckpt,
dtype=torch.float16,
device=torch.device("cuda"),
)
# Pass Stage 1 latent tensor; receive 2x spatial resolution
high_res_latent = upsampler(latent)
The upsample_video function—imported from ltx_core.model.upsampler—performs the actual latent-space operation by combining encoder-derived features with the upsampler's learned transformations.
Adding Temporal Upsampling for Higher Frame Rates
Temporal upsampling multiplies frame count. Set temporal_upsample_rounds=2 to get 4× frames, 3 for 8×, etc.
Python API: DFRPipeline with Both Upsamplers
from ltx_pipelines.dfr_pipeline import DFRPipeline
pipeline = DFRPipeline(
video_ckpt="/models/ltx-2.3-video-x2-1.0.safetensors",
spatial_upsampler_path="/models/ltx-2.3-spatial-upscaler-x2-1.0.safetensors",
temporal_upsampler_path="/models/ltx-2.3-temporal-upscaler-x2-1.0.safetensors",
temporal_upsample_rounds=1, # 1 round = 2× frames; 2 rounds = 4×
dtype=torch.float16,
device=torch.device("cuda"),
)
final_video = pipeline(prompt="A soaring eagle over a mountain range")
What Happens Internally
According to the source in dfr_pipeline.py, the pipeline executes:
# Spatial upsample first
latent = self.upsampler(video_state.latent[:1])
# Then temporal upsample in a loop
for _ in range(self.temporal_upsample_rounds):
latent = self.temporal_upsampler(latent)
Each temporal upsampler forward pass doubles the temporal dimension of the latent.
Command-Line Usage for LTX-2 Upsampling
The same functionality is exposed via CLI flags defined in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py:
| Flag | Purpose |
|---|---|
--spatial-upsampler-path |
Path to spatial upsampler checkpoint |
--temporal-upsampler-path |
Path to temporal upsampler checkpoint |
--temporal-upsample-rounds |
Number of temporal doubling rounds (default: 0) |
python -m ltx_pipelines.dfr_pipeline \
--prompt "A bustling city at night" \
--spatial-upsampler-path /models/ltx-2.3-spatial-upscaler-x2-1.0.safetensors \
--temporal-upsampler-path /models/ltx-2.3-temporal-upscaler-x2-1.0.safetensors \
--temporal-upsample-rounds 2
With rounds=2, you get 4× the original frame count after spatial upsampling completes.
Performance and Practical Considerations
Memory: Each upsampler loads its own checkpoint. Spatial and temporal models are distinct—plan GPU memory accordingly.
Checkpoint flexibility: Because upsamplers are standalone, you can mix:
- Different spatial scaling factors (×2, ×4)
- Custom-trained upsamplers on specific domains
- Temporal upsamplers trained for different motion characteristics
Stage ordering: Always spatial first, then temporal. The pipeline enforces this in dfr_pipeline.py.
Summary
- LTX-2 spatial and temporal upsamplers operate as separate, swappable checkpoints in a two-stage pipeline.
- VideoUpsampler in
blocks.pyhandles spatial doubling viaupsample_video(). - Temporal upsampling runs multiple rounds through a distinct checkpoint, doubling frames each round.
- Configure via Python API (
DFRPipeline,VideoUpsampler) or CLI flags--spatial-upsampler-path,--temporal-upsampler-path, and--temporal-upsample-rounds. - The modular design lets you upgrade resolution or frame rate without retraining the base video model.
Frequently Asked Questions
Can I use spatial upsampling without temporal upsampling?
Yes. Omit --temporal-upsampler-path or set temporal_upsample_rounds=0. The VideoUpsampler alone doubles width and height while preserving the original frame count. This is the default behavior when only spatial checkpoint paths are provided.
How many temporal upsample rounds should I use?
Each round doubles frame count. One round gives 2× frames, two rounds give 4×, three rounds give 8×. Most use cases need 1–2 rounds. Beyond that, diminishing returns and memory costs increase. The dfr_pipeline.py implementation places no hard limit, but practical constraints apply.
Are spatial and temporal upsampler checkpoints interchangeable?
No. They are trained for different latent transformations. Spatial upsamplers increase spatial dimensions (H, W). Temporal upsamplers increase the time dimension (T). Both use LatentUpsamplerConfigurator for instantiation, but their weights are not interchangeable. Load each from its designated checkpoint path.
Where is the core upsampling logic implemented?
The upsample_video function in packages/ltx-core/src/ltx_core/model/upsampler/__init__.py performs the actual latent-space upsampling by merging video encoder features with upsampler model outputs. VideoUpsampler in blocks.py wraps this with checkpoint loading and resource management.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →