Which LTX-2 Pipeline Should You Use for Production Text/Image-to-Video Generation?

Use TI2VidTwoStagesHQPipeline for production deployments, as it combines a two-stage coarse-to-fine generation strategy with the Res2S sampler and distilled LoRA refinement to deliver the highest visual quality with optimal inference speed.

The Lightricks LTX-2 repository provides three distinct inference pipelines for text-to-video and image-to-video generation. Selecting the right LTX-2 pipeline for production text/image-to-video generation depends on your specific requirements for quality, latency, and GPU memory constraints. This guide breaks down each option based on the actual source implementation to help you choose the optimal configuration for deployment.

Overview of the Three LTX-2 Pipelines

LTX-2 ships with three ready-to-use pipelines in the packages/ltx-pipelines/src/ltx_pipelines/ directory. Each implements a different trade-off between inference speed and output fidelity.

TI2VidOneStagePipeline

The TI2VidOneStagePipeline (implemented in ti2vid_one_stage.py) generates the final video in a single diffusion pass at the target resolution. It requires the full (non-distilled) model checkpoint and runs the classic Euler-based sampler.

  • Best for: Quick prototyping or scenarios where inference latency is the primary concern and the full model can be loaded in memory.
  • Trade-off: Fastest execution but lacks the refinement capabilities of multi-stage approaches.

TI2VidTwoStagesPipeline

The TI2VidTwoStagesPipeline (implemented in ti2vid_two_stages.py) uses a two-stage approach: stage 1 creates a half-resolution video with classifier-free guidance (CFG), and stage 2 upsamples ×2 and refines it with a distilled LoRA using the standard Euler sampler.

  • Best for: Production workloads that need a good trade-off between quality and compute.
  • Key feature: The distilled LoRA in stage 2 improves visual fidelity while keeping memory usage manageable compared to full-resolution generation.

TI2VidTwoStagesHQPipeline

The TI2VidTwoStagesHQPipeline (implemented in ti2vid_two_stages_hq.py) follows the same two-stage structure but replaces the Euler sampler with the second-order Res2S sampler. This yields comparable quality with fewer inference steps, making it the most efficient high-quality option.

  • Best for: Production deployment where you want the best visual quality and speed.
  • Key advantage: The Res2sDiffusionStep implementation (found in the pipeline's denoising logic) reduces the number of diffusion steps required while preserving detail, and the distilled LoRA further sharpens the output.

Why TI2VidTwoStagesHQPipeline is the Production Choice

According to the LTX-2 source code, the HQ two-stage pipeline provides four critical advantages for production environments:

Higher Visual Fidelity

Stage 2 applies a distilled LoRA that has been fine-tuned on high-resolution data, producing sharper frames and more coherent motion. The separation of coarse generation (stage 1) and refinement (stage 2) allows each phase to optimize for different aspects of video quality.

Faster Inference with Res2S

The Res2S sampler converges in fewer steps than the standard Euler method, cutting GPU time without sacrificing quality. While the standard two-stage pipeline requires more steps to achieve similar results, the HQ variant can operate effectively with approximately 30 inference steps due to the second-order sampling implementation.

Scalable Memory Usage

By generating a half-resolution video first, the pipeline fits comfortably on a single GPU even for long sequences. The system then upsamples only the latent representation in stage 2, rather than processing full-resolution tensors throughout the entire generation process.

Flexible Conditioning

Both two-stage pipelines accept an optional list of image conditioning inputs through the ImageConditioningInput type, allowing you to guide generation with reference images alongside text prompts.

Implementation Guide

Below is a minimal example showing how to launch the production-grade pipeline from a Python script. Replace the placeholder paths with your actual checkpoint, LoRA, and upsampler files.

Prerequisites

Ensure you have the following model resources available:

  • Full LTX-2 checkpoint (non-distilled)
  • Distilled LoRA checkpoint for stage 2 refinement
  • Spatial upsampler checkpoint
  • Gemma root directory for text encoding

Loading the Pipeline

import torch
from ltx_pipelines.ti2vid_two_stages_hq import TI2VidTwoStagesHQPipeline
from ltx_core.loader import LoraPathStrengthAndSDOps
from ltx_pipelines.utils.args import ImageConditioningInput

# Define model resources

checkpoint_path = "/path/to/full_ltx2_checkpoint"
distilled_lora = [
    LoraPathStrengthAndSDOps(
        path="/path/to/distilled_lora.ckpt",
        strength=0.8,                # strength for stage 2

        sd_ops=None,                 # optional extra ops

    )
]
spatial_upsampler_path = "/path/to/spatial_upsampler.ckpt"
gemma_root = "/path/to/gemma_root"

# Optional image conditioning (list of (image_tensor, weight) tuples)

images: list[ImageConditioningInput] = []   # No image conditioning

# Build the pipeline

pipeline = TI2VidTwoStagesHQPipeline(
    checkpoint_path=checkpoint_path,
    distilled_lora=distilled_lora,
    distilled_lora_strength_stage_1=0.5,   # weaker guidance in stage 1

    distilled_lora_strength_stage_2=0.8,   # stronger refinement in stage 2

    spatial_upsampler_path=spatial_upsampler_path,
    gemma_root=gemma_root,
    loras=(),                               # additional LoRAs if needed

    quantization=None,                      # or QuantizationPolicy for TRT-LLM

    offload_mode=torch.device("cpu")        # GPU residence with CPU offload option

)

Running Inference


# Execute generation

video, audio = pipeline(
    prompt="A sunset over a bustling futuristic city",
    negative_prompt="low-resolution, blurry",
    seed=42,
    height=720,
    width=1280,
    num_frames=48,
    frame_rate=24.0,
    num_inference_steps=30,                 # Res2S works well with ~30 steps

    video_guider_params=dict(
        cfg_scale=7.0,
        stg_scale=0.0,
        rescale_scale=0.0,
        modality_scale=0.0,
        skip_step=0,
        stg_blocks=0,
    ),
    audio_guider_params=dict(
        cfg_scale=7.0,
        stg_scale=0.0,
        rescale_scale=0.0,
        modality_scale=0.0,
        skip_step=0,
        stg_blocks=0,
    ),
    images=images,
    tiling_config=None,                     # use default tiling

    enhance_prompt=False,
    max_batch_size=1,
)

# Encode to MP4

from ltx_pipelines.utils.media_io import encode_video
encode_video(
    video=video,
    fps=24.0,
    audio=audio,
    output_path="output.mp4",
    video_chunks_number=1,
)

Key configuration details:

  • Set distilled_lora_strength_stage_1 lower (0.5) to avoid over-constraining the coarse generation, and higher for stage 2 (0.8) to sharpen the final output.
  • The video_guider_params and audio_guider_params dictionaries map directly to MultiModalGuiderParams classes defined in the pipeline utilities.

Key Source Files

Understanding these implementation files helps with advanced customization:

  • ti2vid_one_stage.py: Implements the single-stage pipeline for fast prototyping.
  • ti2vid_two_stages.py: Implements the standard two-stage pipeline using the Euler sampler.
  • ti2vid_two_stages_hq.py: Implements the high-quality two-stage pipeline with Res2S sampling.
  • utils/blocks.py: Contains core building blocks including PromptEncoder, ImageConditioner, DiffusionStage, and VideoUpsampler.
  • utils/denoisers.py: Provides FactoryGuidedDenoiser, SimpleDenoiser, and the Res2S-compatible GuidedDenoiser.
  • utils/media_io.py: Helper functions to encode generated video and audio into common container formats.
  • utils/args.py: Argument parsers for command-line tools, exposing the same programmatic options.

Summary

  • Use TI2VidTwoStagesHQPipeline for production deployments requiring the best balance of quality and speed.
  • The Res2S sampler in the HQ pipeline reduces inference steps compared to standard Euler sampling while maintaining visual fidelity.
  • Two-stage generation halves memory requirements by processing at reduced resolution initially, then upsampling with LoRA refinement.
  • Configure stage-specific LoRA strengths (0.5 for stage 1, 0.8 for stage 2) to optimize coarse generation without over-constraining detail.
  • Reference ti2vid_two_stages_hq.py directly when implementing custom inference workflows or integrating with existing ML pipelines.

Frequently Asked Questions

What is the difference between the standard and HQ two-stage pipelines?

The standard TI2VidTwoStagesPipeline uses the Euler sampler in both stages, requiring more diffusion steps to achieve quality results. The TI2VidTwoStagesHQPipeline replaces the stage 2 sampler with Res2S (second-order sampling), which converges faster and produces comparable or better quality with approximately 30 steps instead of 50+. Both use the same distilled LoRA refinement, but the HQ variant completes inference significantly faster.

Can I use image conditioning with the production pipeline?

Yes. Both two-stage pipelines accept an optional images parameter containing a list of ImageConditioningInput tuples. Each tuple consists of an image tensor and a weight value that controls the conditioning strength. This allows you to guide video generation with reference images alongside text prompts, implemented through the ImageConditioner class in utils/blocks.py.

How much GPU memory does the HQ two-stage pipeline require?

The pipeline generates video at half resolution during stage 1, then upsamples by 2× in stage 2. This approach keeps memory usage manageable even for 720p or 1080p outputs, typically fitting on a single high-end GPU (e.g., A100 or H100) for sequences up to 48 frames. The offload_mode parameter also supports CPU offloading for models that exceed VRAM capacity.

When should I use the single-stage pipeline instead?

Use TI2VidOneStagePipeline only when latency is the absolute priority and you have sufficient VRAM to load the full non-distilled checkpoint at target resolution. It skips the LoRA refinement and upsampling stages, producing results faster but with noticeably lower temporal consistency and detail compared to the two-stage approaches.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →