Understanding the LTX-2 Two-Stage Pipeline Architecture: A Deep Dive into Text-to-Video Generation
The LTX-2 text-to-video pipeline uses a two-stage architecture that first generates a low-resolution video at half the target resolution, then upsamples and refines it using a distilled LoRA for high-quality output.
The LTX-2 repository from Lightricks implements one of the most sophisticated open-source text-to-video generation systems available today. This article examines the LTX-2 two-stage pipeline architecture implemented in packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py, breaking down how it orchestrates prompt encoding, diffusion, upsampling, and decoding to transform text prompts into high-fidelity video with synchronized audio.
Core Pipeline Class: TI2VidTwoStagesPipeline
The entry point for the entire system is the TI2VidTwoStagesPipeline class, defined at lines 61-68 of the main pipeline file. This class wires together all sub-modules and exposes a callable interface that accepts prompts, conditioning images, and generation parameters.
from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_core.loader import LoraPathStrengthAndSDOps
pipeline = TI2VidTwoStagesPipeline(
model_paths=ModelPaths(),
distilled_lora=[LoraPathStrengthAndSDOps("distilled_lora.ckpt", 0.8, None)],
spatial_upsampler_path="spatial_upsampler.ckpt",
loras=[],
)
The pipeline is designed for modularity — each component operates as a self-contained class, enabling easy replacement or fine-tuning of individual parts without disrupting the entire generation flow.
Stage 1: Low-Resolution Video Generation
Stage 1 produces the initial video latent at half the target resolution. This stage involves five key steps:
1. Prompt Encoding with PromptEncoder
The PromptEncoder (instantiated at lines 89-97) converts text prompts into latent video and audio contexts that condition the diffusion process.
2. Image Conditioning with ImageConditioner
The ImageConditioner (lines 98-105) prepares optional conditioning images for the video encoder, enabling image-to-video generation workflows.
3. Diffusion Stage Loading
DiffusionStage.from_checkpoint (lines 36-46) loads the main transformer model and any user-specified LoRAs. This establishes the backbone for the denoising process.
4. Guided Denoising with FactoryGuidedDenoiser
The FactoryGuidedDenoiser (used at lines 47-60) combines prompt, video, and audio contexts with configurable multimodal guiders including:
- CFG (Classifier-Free Guidance)
- STG (Self-Attention Guidance)
- Rescale (noise rescaling)
5. Stage 1 Execution
The pipeline calls self.stage_1(...) (lines 66-70) to produce video_state and audio_state latents.
video, audio, n_frames, _ = pipeline(
prompt="A sunrise over a mountain lake",
negative_prompt="low quality",
seed=42,
height=720,
width=1280,
frame_rate=30.0,
num_inference_steps=50,
images=[], # optional conditioning images
)
Stage 2: Spatial Upsampling and Refinement
Stage 2 transforms the half-resolution output into the final high-resolution video using a distilled LoRA specialized for quality refinement.
Spatial Upsampling with VideoUpsampler
The VideoUpsampler (lines 105-112) enlarges the Stage 1 latent video by 2×, bringing it to the target resolution.
Distilled LoRA for High-Quality Refinement
The second DiffusionStage (lines 47-57) reuses the same base transformer but adds a distilled LoRA (distilled_lora) optimized for refinement tasks. This architectural choice separates the coarse generation (Stage 1) from fine detail enhancement (Stage 2).
Simple Denoising and Final Decoding
SimpleDenoiser(lines 89-92) runs the second diffusion pass using the upscaled latent as initializationVideoDecoder(lines 13-21) andAudioDecoder(lines 21-28) convert refined latents to final video frames and audio waveform
Key Architectural Components
| Component | Purpose | Location |
|---|---|---|
AllocatorTrimStrategy |
GPU memory trimming during inference | ltx_core/allocator_trim_strategy.py |
DiffVAEMode |
VAE execution mode (eager or chunked) | ltx_core/model/video_vae/transformer.py |
Tiling & Auto-Tiling |
Memory-efficient handling of large frames | ltx_core/model/video_vae.py |
Multimodal Guider Factory |
Per-modality CFG/STG/Rescale guiders | ltx_core/components/guiders.py |
DurationPredictor |
Automatic frame count prediction | ltx_core/duration_head/duration_head.py |
Complete Runtime Flow
- CLI parsing —
resolve_cli_params()builds the argument namespace - Pipeline construction —
TI2VidTwoStagesPipeline(...)supplies model paths, LoRAs, device, and quantization settings - Prompt & conditioning —
prompt_encoderandimage_conditionerproduce latent contexts - Stage 1 diffusion —
stage_1runs guided denoising on half-resolution latents - Upsampling —
upsamplerenlarges the latent video - Stage 2 diffusion —
stage_2refines using the distilled LoRA - Decoding & encoding —
video_decoder,audio_decoder, andencode_videowrite the final MP4
Command-Line Usage
The pipeline supports direct CLI invocation matching the main() function structure:
python -m ltx_pipelines.ti2vid_two_stages \
--prompt "A futuristic city at night" \
--negative_prompt "blurred" \
--height 720 --width 1280 \
--frame-rate 24 \
--num-inference-steps 60 \
--model-paths /path/to/models \
--distilled-lora /path/to/distilled_lora.ckpt \
--spatial-upsampler-path /path/to/upsampler.ckpt \
--output-path result.mp4
Source File Reference
| File | Purpose |
|---|---|
ti2vid_two_stages.py |
Main two-stage pipeline implementation |
utils/blocks.py |
Component definitions (PromptEncoder, ImageConditioner, VideoUpsampler, DiffusionStage, etc.) |
utils/args.py |
CLI parsers and MultiModalGuiderParams |
utils/media_io/encode.py |
MP4 encoding utilities |
utils/model_paths.py |
Checkpoint location management |
ltx_core/components/guiders.py |
Multimodal guidance factory |
ltx_core/duration_head/duration_head.py |
Automatic duration prediction |
Summary
- The LTX-2 two-stage pipeline separates generation into coarse low-resolution synthesis (Stage 1) and high-quality refinement (Stage 2)
- Half-resolution first pass reduces computational cost; 2× upsampling plus distilled LoRA recovers quality
- Modular design enables component substitution without full pipeline rewrites
- Multimodal guiders (CFG, STG, Rescale) operate independently on video and audio conditioning
- Multi-GPU and auto-tiling support scales from consumer GPUs to datacenter hardware
Frequently Asked Questions
What is the purpose of the distilled LoRA in Stage 2?
The distilled LoRA specializes in refining upscaled latents for high-quality output. According to the LTX-2 source code, this LoRA is loaded as an additional component on top of the base transformer in the second DiffusionStage (lines 47-57 of ti2vid_two_stages.py), distinguishing it from user-provided LoRAs used in Stage 1.
How does the pipeline handle memory constraints for large videos?
The architecture implements multiple memory optimization strategies: AllocatorTrimStrategy controls GPU memory trimming, DiffVAEMode supports chunked VAE execution, and auto-tiling in ltx_core/model/video_vae.py automatically adapts to hardware limits by splitting large frames into manageable tiles.
Can I use custom images to condition the video generation?
Yes. The ImageConditioner component (lines 98-105) accepts optional conditioning images through the images parameter. Pass a list of images to the pipeline call, and they will be encoded and integrated into the video generation context.
What happens when num_frames is set to "auto"?
The DurationPredictor in ltx_core/duration_head/duration_head.py automatically predicts the optimal number of frames based on the prompt and other generation parameters. This eliminates manual frame count specification while maintaining temporal coherence appropriate to the content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →