DistilledPipeline vs TI2VidTwoStagesPipeline: Speed and Quality Comparison in LTX-2

DistilledPipeline delivers the fastest inference with 12 total diffusion steps, while TI2VidTwoStagesPipeline prioritizes higher output quality through 44 steps, classifier-free guidance, and LoRA refinement.

Both pipelines in the Lightricks/LTX-2 repository implement two-stage video generation, but they diverge significantly in model architecture, denoising strategy, and guidance mechanisms. Your choice depends on whether latency or visual fidelity matters more for your use case.

Stage 1: Foundation Model and Denoising Schedule

The architectural difference begins with the transformer checkpoint each pipeline loads.

DistilledPipeline: Fixed Eight-Step Schedule

In packages/ltx-pipelines/src/ltx_pipelines/distilled.py, the DistilledPipeline initializes with a distilled transformer checkpoint (ltx-2.5-22b-distilled-transformer-bf16.safetensors). This model has been fine-tuned to operate on a rigid eight-step sigma schedule defined in DISTILLED_SIGMAS:


# From distilled.py lines 200-202

DISTILLED_SIGMAS = torch.tensor([14.6146, 6.4744, 2.6850, 0.8623, 0.1674, 0.0100, 0.0010, 0.0001])

This fixed schedule eliminates runtime step configuration. The pipeline detects whether to use an ancestral sampler through should_use_ancestral_sampler (lines 76-84), but critically, no classifier-free guidance (CFG) is applied—avoiding the memory and computation overhead of dual forward passes.

TI2VidTwoStagesPipeline: Configurable 40-Step Generation

The TI2VidTwoStagesPipeline in packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py loads the full non-distilled transformer and delegates step scheduling to LTX2Scheduler (lines 44-48). By default, this yields 40 denoising steps, though users can override via num_inference_steps.

More significantly, this pipeline wraps its denoiser with FactoryGuidedDenoiser (lines 48-60), enabling classifier-free guidance through separate positive and negative prompt conditioning with per-modality scaling:

video_guider_params=MultiModalGuiderParams(cfg_scale=3.0),
audio_guider_params=MultiModalGuiderParams(cfg_scale=7.0),

Stage 2: Refinement Approach

Both pipelines share the same spatial upsampler architecture but differ in how they refine upscaled latents.

Pipeline Stage 2 Steps Refinement Method
DistilledPipeline 4 steps (STAGE_2_DISTILLED_SIGMAS) Deterministic refinement; no LoRA needed
TI2VidTwoStagesPipeline 4 steps (same sigmas) + distilled LoRA Learned detail injection via LoRA

The DistilledPipeline relies entirely on the distilled checkpoint's embedded knowledge. The TI2VidTwoStagesPipeline accepts a distilled_lora parameter—typically ltx-2.5-22b-distilled-lora-450-bf16.safetensors—to restore high-frequency details lost during upscaling.

Inference Speed Comparison

DistilledPipeline is the fastest option in the LTX-2 repository. With 12 total diffusion steps (8 + 4) and no CFG loop, it typically executes in 0.5–1× the time of the full model on identical hardware. The repository README explicitly identifies it as "Fastest inference with 8 predefined sigmas."

TI2VidTwoStagesPipeline runs 3–5× slower due to:

  • 40 + 4 = 44 total denoising steps
  • CFG evaluation requiring dual forward passes per step
  • LoRA weight application during stage 2

The computational gap widens with higher-resolution outputs where memory bandwidth becomes constrained.

Output Quality Comparison

Quality trade-offs map directly to architectural choices:

DistilledPipeline limitations:

  • Fixed schedule restricts iterative refinement
  • No negative prompt guidance for error correction
  • Deterministic stage 2 cannot hallucinate missing detail

TI2VidTwoStagesPipeline advantages:

  • CFG steering improves prompt adherence and texture richness
  • Negative prompt filtering reduces artifacts
  • Distilled LoRA recovers fine details post-upscaling

For maximum quality, the ti2vid_two_stages_hq.py variant replaces the default sampler with res_2s—a second-order method achieving superior results with fewer effective steps.

Practical Usage Examples

Fast Generation with DistilledPipeline

from ltx_pipelines.distilled import DistilledPipeline
from ltx_pipelines.utils.model_paths import ModelPaths

pipeline = DistilledPipeline(
    model_paths=ModelPaths(
        transformer_path="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
        text_encoder_path="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
        video_vae_path="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
        audio_vae_path="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
        spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
    ),
    spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
    loras=[],
)

video, audio, num_frames, tiling_cfg = pipeline(
    prompt="A sunrise over a calm lake, with gentle mist rising.",
    seed=42,
    height=512,
    width=768,
    frame_rate=24.0,
    images=[],
)

Note the absence of negative_prompt or guider parameters—CFG is not supported.

High-Quality Generation with TI2VidTwoStagesPipeline

from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_core.components.guiders import MultiModalGuiderParams

pipeline = TI2VidTwoStagesPipeline(
    model_paths=ModelPaths(
        transformer_path="models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors",
        text_encoder_path="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
        video_vae_path="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
        audio_vae_path="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
        spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
    ),
    distilled_lora=[
        "models/ltx-2.5/loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors"
    ],
    spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
    loras=[],
)

video, audio, num_frames, tiling_cfg = pipeline(
    prompt="A bustling futuristic city at night, neon lights reflecting on rain-slick streets.",
    negative_prompt="low quality, blurry, noise",
    seed=42,
    height=1088,
    width=1920,
    frame_rate=24.0,
    num_inference_steps=40,
    video_guider_params=MultiModalGuiderParams(cfg_scale=3.0),
    audio_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
    images=[],
)

Source Code Reference

Key files implementing these pipelines:

Summary

  • Speed winner: DistilledPipeline — 12 steps, no CFG, fastest LTX-2 inference
  • Quality winner: TI2VidTwoStagesPipeline — 44 steps, CFG guidance, LoRA refinement
  • Architecture difference: Distilled transformer vs. full transformer with learned upscaling adapter
  • Use distilled for: Prototyping, real-time demos, latency-sensitive applications
  • Use TI2Vid for: Production video, fine detail preservation, precise prompt control

Frequently Asked Questions

How much faster is DistilledPipeline compared to TI2VidTwoStagesPipeline?

DistilledPipeline typically runs 3–5× faster on identical hardware. The speedup comes from 12 total diffusion steps versus 44, elimination of CFG dual-forward overhead, and optimized ancestral sampling. On high-resolution generations (1080p+), the gap widens due to memory bandwidth savings from fewer transformer evaluations.

Can I use classifier-free guidance with DistilledPipeline?

No. The DistilledPipeline in distilled.py explicitly does not implement CFG. The distilled transformer checkpoint was trained to generate directly without negative prompt conditioning. Attempting to add CFG would break the fixed sigma schedule assumptions baked into the model weights.

Does TI2VidTwoStagesPipeline always require 40 inference steps?

No—num_inference_steps is configurable. However, reducing steps below 30 typically degrades quality noticeably since the non-distilled transformer relies on iterative refinement. The HQ variant in ti2vid_two_stages_hq.py can achieve comparable quality with fewer effective steps through its second-order res_2s sampler.

Which pipeline should I choose for production video generation?

Choose TI2VidTwoStagesPipeline or its HQ variant when visual fidelity matters. The combination of CFG steering, negative prompt filtering, and distilled LoRA refinement produces more detailed, prompt-faithful output. Reserve DistilledPipeline for previews, rapid iteration, or bandwidth-constrained deployments where quality trade-offs are acceptable.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →