Understanding the LTX-2 Two-Stage Pipeline Architecture: A Deep Dive into Text-to-Video Generation

The LTX-2 text-to-video pipeline uses a two-stage architecture that first generates a low-resolution video at half the target resolution, then upsamples and refines it using a distilled LoRA for high-quality output.

The LTX-2 repository from Lightricks implements one of the most sophisticated open-source text-to-video generation systems available today. This article examines the LTX-2 two-stage pipeline architecture implemented in packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py, breaking down how it orchestrates prompt encoding, diffusion, upsampling, and decoding to transform text prompts into high-fidelity video with synchronized audio.

Core Pipeline Class: TI2VidTwoStagesPipeline

The entry point for the entire system is the TI2VidTwoStagesPipeline class, defined at lines 61-68 of the main pipeline file. This class wires together all sub-modules and exposes a callable interface that accepts prompts, conditioning images, and generation parameters.

from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_core.loader import LoraPathStrengthAndSDOps

pipeline = TI2VidTwoStagesPipeline(
    model_paths=ModelPaths(),
    distilled_lora=[LoraPathStrengthAndSDOps("distilled_lora.ckpt", 0.8, None)],
    spatial_upsampler_path="spatial_upsampler.ckpt",
    loras=[],
)

The pipeline is designed for modularity — each component operates as a self-contained class, enabling easy replacement or fine-tuning of individual parts without disrupting the entire generation flow.

Stage 1: Low-Resolution Video Generation

Stage 1 produces the initial video latent at half the target resolution. This stage involves five key steps:

1. Prompt Encoding with PromptEncoder

The PromptEncoder (instantiated at lines 89-97) converts text prompts into latent video and audio contexts that condition the diffusion process.

2. Image Conditioning with ImageConditioner

The ImageConditioner (lines 98-105) prepares optional conditioning images for the video encoder, enabling image-to-video generation workflows.

3. Diffusion Stage Loading

DiffusionStage.from_checkpoint (lines 36-46) loads the main transformer model and any user-specified LoRAs. This establishes the backbone for the denoising process.

4. Guided Denoising with FactoryGuidedDenoiser

The FactoryGuidedDenoiser (used at lines 47-60) combines prompt, video, and audio contexts with configurable multimodal guiders including:

  • CFG (Classifier-Free Guidance)
  • STG (Self-Attention Guidance)
  • Rescale (noise rescaling)

5. Stage 1 Execution

The pipeline calls self.stage_1(...) (lines 66-70) to produce video_state and audio_state latents.

video, audio, n_frames, _ = pipeline(
    prompt="A sunrise over a mountain lake",
    negative_prompt="low quality",
    seed=42,
    height=720,
    width=1280,
    frame_rate=30.0,
    num_inference_steps=50,
    images=[],  # optional conditioning images

)

Stage 2: Spatial Upsampling and Refinement

Stage 2 transforms the half-resolution output into the final high-resolution video using a distilled LoRA specialized for quality refinement.

Spatial Upsampling with VideoUpsampler

The VideoUpsampler (lines 105-112) enlarges the Stage 1 latent video by 2×, bringing it to the target resolution.

Distilled LoRA for High-Quality Refinement

The second DiffusionStage (lines 47-57) reuses the same base transformer but adds a distilled LoRA (distilled_lora) optimized for refinement tasks. This architectural choice separates the coarse generation (Stage 1) from fine detail enhancement (Stage 2).

Simple Denoising and Final Decoding

  • SimpleDenoiser (lines 89-92) runs the second diffusion pass using the upscaled latent as initialization
  • VideoDecoder (lines 13-21) and AudioDecoder (lines 21-28) convert refined latents to final video frames and audio waveform

Key Architectural Components

Component Purpose Location
AllocatorTrimStrategy GPU memory trimming during inference ltx_core/allocator_trim_strategy.py
DiffVAEMode VAE execution mode (eager or chunked) ltx_core/model/video_vae/transformer.py
Tiling & Auto-Tiling Memory-efficient handling of large frames ltx_core/model/video_vae.py
Multimodal Guider Factory Per-modality CFG/STG/Rescale guiders ltx_core/components/guiders.py
DurationPredictor Automatic frame count prediction ltx_core/duration_head/duration_head.py

Complete Runtime Flow

  1. CLI parsing — resolve_cli_params() builds the argument namespace
  2. Pipeline construction — TI2VidTwoStagesPipeline(...) supplies model paths, LoRAs, device, and quantization settings
  3. Prompt & conditioning — prompt_encoder and image_conditioner produce latent contexts
  4. Stage 1 diffusion — stage_1 runs guided denoising on half-resolution latents
  5. Upsampling — upsampler enlarges the latent video
  6. Stage 2 diffusion — stage_2 refines using the distilled LoRA
  7. Decoding & encoding — video_decoder, audio_decoder, and encode_video write the final MP4

Command-Line Usage

The pipeline supports direct CLI invocation matching the main() function structure:

python -m ltx_pipelines.ti2vid_two_stages \
    --prompt "A futuristic city at night" \
    --negative_prompt "blurred" \
    --height 720 --width 1280 \
    --frame-rate 24 \
    --num-inference-steps 60 \
    --model-paths /path/to/models \
    --distilled-lora /path/to/distilled_lora.ckpt \
    --spatial-upsampler-path /path/to/upsampler.ckpt \
    --output-path result.mp4

Source File Reference

File Purpose
ti2vid_two_stages.py Main two-stage pipeline implementation
utils/blocks.py Component definitions (PromptEncoder, ImageConditioner, VideoUpsampler, DiffusionStage, etc.)
utils/args.py CLI parsers and MultiModalGuiderParams
utils/media_io/encode.py MP4 encoding utilities
utils/model_paths.py Checkpoint location management
ltx_core/components/guiders.py Multimodal guidance factory
ltx_core/duration_head/duration_head.py Automatic duration prediction

Summary

  • The LTX-2 two-stage pipeline separates generation into coarse low-resolution synthesis (Stage 1) and high-quality refinement (Stage 2)
  • Half-resolution first pass reduces computational cost; 2× upsampling plus distilled LoRA recovers quality
  • Modular design enables component substitution without full pipeline rewrites
  • Multimodal guiders (CFG, STG, Rescale) operate independently on video and audio conditioning
  • Multi-GPU and auto-tiling support scales from consumer GPUs to datacenter hardware

Frequently Asked Questions

What is the purpose of the distilled LoRA in Stage 2?

The distilled LoRA specializes in refining upscaled latents for high-quality output. According to the LTX-2 source code, this LoRA is loaded as an additional component on top of the base transformer in the second DiffusionStage (lines 47-57 of ti2vid_two_stages.py), distinguishing it from user-provided LoRAs used in Stage 1.

How does the pipeline handle memory constraints for large videos?

The architecture implements multiple memory optimization strategies: AllocatorTrimStrategy controls GPU memory trimming, DiffVAEMode supports chunked VAE execution, and auto-tiling in ltx_core/model/video_vae.py automatically adapts to hardware limits by splitting large frames into manageable tiles.

Can I use custom images to condition the video generation?

Yes. The ImageConditioner component (lines 98-105) accepts optional conditioning images through the images parameter. Pass a list of images to the pipeline call, and they will be encoded and integrated into the video generation context.

What happens when num_frames is set to "auto"?

The DurationPredictor in ltx_core/duration_head/duration_head.py automatically predicts the optimal number of frames based on the prompt and other generation parameters. This eliminates manual frame count specification while maintaining temporal coherence appropriate to the content.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →