# Which LTX-2 Pipeline Should You Use for Production Text/Image-to-Video Generation?

> Choose TI2VidTwoStagesHQPipeline for production text to video generation. Leverage its two-stage coarse-to-fine strategy and Res2S sampler for superior quality and speed.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: architecture
- Published: 2026-06-21

---

**Use `TI2VidTwoStagesHQPipeline` for production deployments, as it combines a two-stage coarse-to-fine generation strategy with the Res2S sampler and distilled LoRA refinement to deliver the highest visual quality with optimal inference speed.**

The Lightricks LTX-2 repository provides three distinct inference pipelines for text-to-video and image-to-video generation. Selecting the right LTX-2 pipeline for production text/image-to-video generation depends on your specific requirements for quality, latency, and GPU memory constraints. This guide breaks down each option based on the actual source implementation to help you choose the optimal configuration for deployment.

## Overview of the Three LTX-2 Pipelines

LTX-2 ships with three ready-to-use pipelines in the `packages/ltx-pipelines/src/ltx_pipelines/` directory. Each implements a different trade-off between inference speed and output fidelity.

### TI2VidOneStagePipeline

The **`TI2VidOneStagePipeline`** (implemented in [`ti2vid_one_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_one_stage.py)) generates the final video in a single diffusion pass at the target resolution. It requires the full (non-distilled) model checkpoint and runs the classic Euler-based sampler.

- **Best for**: Quick prototyping or scenarios where inference latency is the primary concern and the full model can be loaded in memory.
- **Trade-off**: Fastest execution but lacks the refinement capabilities of multi-stage approaches.

### TI2VidTwoStagesPipeline

The **`TI2VidTwoStagesPipeline`** (implemented in [`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py)) uses a two-stage approach: stage 1 creates a half-resolution video with classifier-free guidance (CFG), and stage 2 upsamples ×2 and refines it with a distilled LoRA using the standard Euler sampler.

- **Best for**: Production workloads that need a good trade-off between quality and compute.
- **Key feature**: The distilled LoRA in stage 2 improves visual fidelity while keeping memory usage manageable compared to full-resolution generation.

### TI2VidTwoStagesHQPipeline

The **`TI2VidTwoStagesHQPipeline`** (implemented in [`ti2vid_two_stages_hq.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages_hq.py)) follows the same two-stage structure but replaces the Euler sampler with the second-order **Res2S** sampler. This yields comparable quality with fewer inference steps, making it the most efficient high-quality option.

- **Best for**: Production deployment where you want the best visual quality and speed.
- **Key advantage**: The `Res2sDiffusionStep` implementation (found in the pipeline's denoising logic) reduces the number of diffusion steps required while preserving detail, and the distilled LoRA further sharpens the output.

## Why TI2VidTwoStagesHQPipeline is the Production Choice

According to the LTX-2 source code, the HQ two-stage pipeline provides four critical advantages for production environments:

### Higher Visual Fidelity

Stage 2 applies a distilled LoRA that has been fine-tuned on high-resolution data, producing sharper frames and more coherent motion. The separation of coarse generation (stage 1) and refinement (stage 2) allows each phase to optimize for different aspects of video quality.

### Faster Inference with Res2S

The **Res2S sampler** converges in fewer steps than the standard Euler method, cutting GPU time without sacrificing quality. While the standard two-stage pipeline requires more steps to achieve similar results, the HQ variant can operate effectively with approximately 30 inference steps due to the second-order sampling implementation.

### Scalable Memory Usage

By generating a half-resolution video first, the pipeline fits comfortably on a single GPU even for long sequences. The system then upsamples only the latent representation in stage 2, rather than processing full-resolution tensors throughout the entire generation process.

### Flexible Conditioning

Both two-stage pipelines accept an optional list of image conditioning inputs through the `ImageConditioningInput` type, allowing you to guide generation with reference images alongside text prompts.

## Implementation Guide

Below is a minimal example showing how to launch the production-grade pipeline from a Python script. Replace the placeholder paths with your actual checkpoint, LoRA, and upsampler files.

### Prerequisites

Ensure you have the following model resources available:
- Full LTX-2 checkpoint (non-distilled)
- Distilled LoRA checkpoint for stage 2 refinement
- Spatial upsampler checkpoint
- Gemma root directory for text encoding

### Loading the Pipeline

```python
import torch
from ltx_pipelines.ti2vid_two_stages_hq import TI2VidTwoStagesHQPipeline
from ltx_core.loader import LoraPathStrengthAndSDOps
from ltx_pipelines.utils.args import ImageConditioningInput

# Define model resources

checkpoint_path = "/path/to/full_ltx2_checkpoint"
distilled_lora = [
    LoraPathStrengthAndSDOps(
        path="/path/to/distilled_lora.ckpt",
        strength=0.8,                # strength for stage 2

        sd_ops=None,                 # optional extra ops

    )
]
spatial_upsampler_path = "/path/to/spatial_upsampler.ckpt"
gemma_root = "/path/to/gemma_root"

# Optional image conditioning (list of (image_tensor, weight) tuples)

images: list[ImageConditioningInput] = []   # No image conditioning

# Build the pipeline

pipeline = TI2VidTwoStagesHQPipeline(
    checkpoint_path=checkpoint_path,
    distilled_lora=distilled_lora,
    distilled_lora_strength_stage_1=0.5,   # weaker guidance in stage 1

    distilled_lora_strength_stage_2=0.8,   # stronger refinement in stage 2

    spatial_upsampler_path=spatial_upsampler_path,
    gemma_root=gemma_root,
    loras=(),                               # additional LoRAs if needed

    quantization=None,                      # or QuantizationPolicy for TRT-LLM

    offload_mode=torch.device("cpu")        # GPU residence with CPU offload option

)

```

### Running Inference

```python

# Execute generation

video, audio = pipeline(
    prompt="A sunset over a bustling futuristic city",
    negative_prompt="low-resolution, blurry",
    seed=42,
    height=720,
    width=1280,
    num_frames=48,
    frame_rate=24.0,
    num_inference_steps=30,                 # Res2S works well with ~30 steps

    video_guider_params=dict(
        cfg_scale=7.0,
        stg_scale=0.0,
        rescale_scale=0.0,
        modality_scale=0.0,
        skip_step=0,
        stg_blocks=0,
    ),
    audio_guider_params=dict(
        cfg_scale=7.0,
        stg_scale=0.0,
        rescale_scale=0.0,
        modality_scale=0.0,
        skip_step=0,
        stg_blocks=0,
    ),
    images=images,
    tiling_config=None,                     # use default tiling

    enhance_prompt=False,
    max_batch_size=1,
)

# Encode to MP4

from ltx_pipelines.utils.media_io import encode_video
encode_video(
    video=video,
    fps=24.0,
    audio=audio,
    output_path="output.mp4",
    video_chunks_number=1,
)

```

**Key configuration details:**
- Set `distilled_lora_strength_stage_1` lower (0.5) to avoid over-constraining the coarse generation, and higher for stage 2 (0.8) to sharpen the final output.
- The `video_guider_params` and `audio_guider_params` dictionaries map directly to `MultiModalGuiderParams` classes defined in the pipeline utilities.

## Key Source Files

Understanding these implementation files helps with advanced customization:

- **[`ti2vid_one_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_one_stage.py)**: Implements the single-stage pipeline for fast prototyping.
- **[`ti2vid_two_stages.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages.py)**: Implements the standard two-stage pipeline using the Euler sampler.
- **[`ti2vid_two_stages_hq.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages_hq.py)**: Implements the high-quality two-stage pipeline with Res2S sampling.
- **[`utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/blocks.py)**: Contains core building blocks including `PromptEncoder`, `ImageConditioner`, `DiffusionStage`, and `VideoUpsampler`.
- **[`utils/denoisers.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/denoisers.py)**: Provides `FactoryGuidedDenoiser`, `SimpleDenoiser`, and the Res2S-compatible `GuidedDenoiser`.
- **[`utils/media_io.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/media_io.py)**: Helper functions to encode generated video and audio into common container formats.
- **[`utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/args.py)**: Argument parsers for command-line tools, exposing the same programmatic options.

## Summary

- **Use `TI2VidTwoStagesHQPipeline`** for production deployments requiring the best balance of quality and speed.
- **The Res2S sampler** in the HQ pipeline reduces inference steps compared to standard Euler sampling while maintaining visual fidelity.
- **Two-stage generation** halves memory requirements by processing at reduced resolution initially, then upsampling with LoRA refinement.
- **Configure stage-specific LoRA strengths** (0.5 for stage 1, 0.8 for stage 2) to optimize coarse generation without over-constraining detail.
- **Reference [`ti2vid_two_stages_hq.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stages_hq.py)** directly when implementing custom inference workflows or integrating with existing ML pipelines.

## Frequently Asked Questions

### What is the difference between the standard and HQ two-stage pipelines?

The standard `TI2VidTwoStagesPipeline` uses the Euler sampler in both stages, requiring more diffusion steps to achieve quality results. The `TI2VidTwoStagesHQPipeline` replaces the stage 2 sampler with **Res2S** (second-order sampling), which converges faster and produces comparable or better quality with approximately 30 steps instead of 50+. Both use the same distilled LoRA refinement, but the HQ variant completes inference significantly faster.

### Can I use image conditioning with the production pipeline?

Yes. Both two-stage pipelines accept an optional `images` parameter containing a list of `ImageConditioningInput` tuples. Each tuple consists of an image tensor and a weight value that controls the conditioning strength. This allows you to guide video generation with reference images alongside text prompts, implemented through the `ImageConditioner` class in [`utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/blocks.py).

### How much GPU memory does the HQ two-stage pipeline require?

The pipeline generates video at half resolution during stage 1, then upsamples by 2× in stage 2. This approach keeps memory usage manageable even for 720p or 1080p outputs, typically fitting on a single high-end GPU (e.g., A100 or H100) for sequences up to 48 frames. The `offload_mode` parameter also supports CPU offloading for models that exceed VRAM capacity.

### When should I use the single-stage pipeline instead?

Use `TI2VidOneStagePipeline` only when latency is the absolute priority and you have sufficient VRAM to load the full non-distilled checkpoint at target resolution. It skips the LoRA refinement and upsampling stages, producing results faster but with noticeably lower temporal consistency and detail compared to the two-stage approaches.