DistilledPipeline vs TI2VidTwoStagesPipeline: Speed and Quality Comparison in LTX-2
DistilledPipeline delivers the fastest inference with 12 total diffusion steps, while TI2VidTwoStagesPipeline prioritizes higher output quality through 44 steps, classifier-free guidance, and LoRA refinement.
Both pipelines in the Lightricks/LTX-2 repository implement two-stage video generation, but they diverge significantly in model architecture, denoising strategy, and guidance mechanisms. Your choice depends on whether latency or visual fidelity matters more for your use case.
Stage 1: Foundation Model and Denoising Schedule
The architectural difference begins with the transformer checkpoint each pipeline loads.
DistilledPipeline: Fixed Eight-Step Schedule
In packages/ltx-pipelines/src/ltx_pipelines/distilled.py, the DistilledPipeline initializes with a distilled transformer checkpoint (ltx-2.5-22b-distilled-transformer-bf16.safetensors). This model has been fine-tuned to operate on a rigid eight-step sigma schedule defined in DISTILLED_SIGMAS:
# From distilled.py lines 200-202
DISTILLED_SIGMAS = torch.tensor([14.6146, 6.4744, 2.6850, 0.8623, 0.1674, 0.0100, 0.0010, 0.0001])
This fixed schedule eliminates runtime step configuration. The pipeline detects whether to use an ancestral sampler through should_use_ancestral_sampler (lines 76-84), but critically, no classifier-free guidance (CFG) is applied—avoiding the memory and computation overhead of dual forward passes.
TI2VidTwoStagesPipeline: Configurable 40-Step Generation
The TI2VidTwoStagesPipeline in packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py loads the full non-distilled transformer and delegates step scheduling to LTX2Scheduler (lines 44-48). By default, this yields 40 denoising steps, though users can override via num_inference_steps.
More significantly, this pipeline wraps its denoiser with FactoryGuidedDenoiser (lines 48-60), enabling classifier-free guidance through separate positive and negative prompt conditioning with per-modality scaling:
video_guider_params=MultiModalGuiderParams(cfg_scale=3.0),
audio_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
Stage 2: Refinement Approach
Both pipelines share the same spatial upsampler architecture but differ in how they refine upscaled latents.
| Pipeline | Stage 2 Steps | Refinement Method |
|---|---|---|
| DistilledPipeline | 4 steps (STAGE_2_DISTILLED_SIGMAS) |
Deterministic refinement; no LoRA needed |
| TI2VidTwoStagesPipeline | 4 steps (same sigmas) + distilled LoRA | Learned detail injection via LoRA |
The DistilledPipeline relies entirely on the distilled checkpoint's embedded knowledge. The TI2VidTwoStagesPipeline accepts a distilled_lora parameter—typically ltx-2.5-22b-distilled-lora-450-bf16.safetensors—to restore high-frequency details lost during upscaling.
Inference Speed Comparison
DistilledPipeline is the fastest option in the LTX-2 repository. With 12 total diffusion steps (8 + 4) and no CFG loop, it typically executes in 0.5–1× the time of the full model on identical hardware. The repository README explicitly identifies it as "Fastest inference with 8 predefined sigmas."
TI2VidTwoStagesPipeline runs 3–5× slower due to:
- 40 + 4 = 44 total denoising steps
- CFG evaluation requiring dual forward passes per step
- LoRA weight application during stage 2
The computational gap widens with higher-resolution outputs where memory bandwidth becomes constrained.
Output Quality Comparison
Quality trade-offs map directly to architectural choices:
DistilledPipeline limitations:
- Fixed schedule restricts iterative refinement
- No negative prompt guidance for error correction
- Deterministic stage 2 cannot hallucinate missing detail
TI2VidTwoStagesPipeline advantages:
- CFG steering improves prompt adherence and texture richness
- Negative prompt filtering reduces artifacts
- Distilled LoRA recovers fine details post-upscaling
For maximum quality, the ti2vid_two_stages_hq.py variant replaces the default sampler with res_2s—a second-order method achieving superior results with fewer effective steps.
Practical Usage Examples
Fast Generation with DistilledPipeline
from ltx_pipelines.distilled import DistilledPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
pipeline = DistilledPipeline(
model_paths=ModelPaths(
transformer_path="models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors",
text_encoder_path="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
video_vae_path="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
audio_vae_path="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
),
spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
loras=[],
)
video, audio, num_frames, tiling_cfg = pipeline(
prompt="A sunrise over a calm lake, with gentle mist rising.",
seed=42,
height=512,
width=768,
frame_rate=24.0,
images=[],
)
Note the absence of negative_prompt or guider parameters—CFG is not supported.
High-Quality Generation with TI2VidTwoStagesPipeline
from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_core.components.guiders import MultiModalGuiderParams
pipeline = TI2VidTwoStagesPipeline(
model_paths=ModelPaths(
transformer_path="models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors",
text_encoder_path="models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors",
video_vae_path="models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors",
audio_vae_path="models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors",
spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
),
distilled_lora=[
"models/ltx-2.5/loras/ltx-2.5-22b-distilled-lora-450-bf16.safetensors"
],
spatial_upsampler_path="models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors",
loras=[],
)
video, audio, num_frames, tiling_cfg = pipeline(
prompt="A bustling futuristic city at night, neon lights reflecting on rain-slick streets.",
negative_prompt="low quality, blurry, noise",
seed=42,
height=1088,
width=1920,
frame_rate=24.0,
num_inference_steps=40,
video_guider_params=MultiModalGuiderParams(cfg_scale=3.0),
audio_guider_params=MultiModalGuiderParams(cfg_scale=7.0),
images=[],
)
Source Code Reference
Key files implementing these pipelines:
packages/ltx-pipelines/src/ltx_pipelines/distilled.py—DistilledPipelinewith fixed sigma schedulespackages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py—TI2VidTwoStagesPipelinewith CFG and LoRA supportpackages/ltx-pipelines/src/ltx_pipelines/utils/constants.py—DISTILLED_SIGMASandSTAGE_2_DISTILLED_SIGMASdefinitionspackages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages_hq.py— HQ variant withres_2ssampler
Summary
- Speed winner:
DistilledPipeline— 12 steps, no CFG, fastest LTX-2 inference - Quality winner:
TI2VidTwoStagesPipeline— 44 steps, CFG guidance, LoRA refinement - Architecture difference: Distilled transformer vs. full transformer with learned upscaling adapter
- Use distilled for: Prototyping, real-time demos, latency-sensitive applications
- Use TI2Vid for: Production video, fine detail preservation, precise prompt control
Frequently Asked Questions
How much faster is DistilledPipeline compared to TI2VidTwoStagesPipeline?
DistilledPipeline typically runs 3–5× faster on identical hardware. The speedup comes from 12 total diffusion steps versus 44, elimination of CFG dual-forward overhead, and optimized ancestral sampling. On high-resolution generations (1080p+), the gap widens due to memory bandwidth savings from fewer transformer evaluations.
Can I use classifier-free guidance with DistilledPipeline?
No. The DistilledPipeline in distilled.py explicitly does not implement CFG. The distilled transformer checkpoint was trained to generate directly without negative prompt conditioning. Attempting to add CFG would break the fixed sigma schedule assumptions baked into the model weights.
Does TI2VidTwoStagesPipeline always require 40 inference steps?
No—num_inference_steps is configurable. However, reducing steps below 30 typically degrades quality noticeably since the non-distilled transformer relies on iterative refinement. The HQ variant in ti2vid_two_stages_hq.py can achieve comparable quality with fewer effective steps through its second-order res_2s sampler.
Which pipeline should I choose for production video generation?
Choose TI2VidTwoStagesPipeline or its HQ variant when visual fidelity matters. The combination of CFG steering, negative prompt filtering, and distilled LoRA refinement produces more detailed, prompt-faithful output. Reserve DistilledPipeline for previews, rapid iteration, or bandwidth-constrained deployments where quality trade-offs are acceptable.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →