LTX-2 DistilledPipeline vs. Full-Model Pipelines: When to Use Which for Video Generation
Use the LTX-2 DistilledPipeline when you need faster inference with lower VRAM (roughly half the memory), and switch to full-model pipelines when maximum quality is your priority.
The LTX-2 video generation framework from Lightricks offers two distinct execution paths: a memory-efficient DistilledPipeline that runs a compressed transformer with fewer denoising steps, and full-model pipelines that preserve the complete diffusion process. Understanding their architectural differences helps you pick the right tool for your hardware constraints and quality requirements.
Core Architectural Differences
Checkpoint Types and Model Size
The fundamental split begins with which weights you load. In utils/args.py, the detect_checkpoint_path function (lines 460-470) inspects metadata to determine whether you've provided a distilled checkpoint:
- DistilledPipeline: Loads
--distilled-checkpoint-pathcontaining compressed transformer weights (~½ the size of full weights) - Full-model pipelines: Loads
--checkpoint-pathwith the complete non-distilled transformer
This size reduction directly translates to lower memory pressure during inference.
Noise Schedule: Fixed Short vs. Configurable Long
The sigma schedules are hardcoded differently between the two approaches. utils/constants.py (lines 15-24) defines the distilled schedule:
# From utils/constants.py
DISTILLED_SIGMA_VALUES = [14.615, 6.315, 2.865, 1.340, 0.615] # 5 values, typically 9→4 steps
STAGE_2_DISTILLED_SIGMA_VALUES = [0.615, 0.291, 0.125, 0.029] # Even shorter for refinement
- DistilledPipeline: Uses fixed
DISTILLED_SIGMAS— typically 9 steps reduced to 4 effective steps - Full-model pipelines: Inherit 30-40 step schedules from checkpoint metadata via
DiffusionStagedefaults
Stage Layout: Two-Stage with Upsampling
DistilledPipeline Structure
The DistilledPipeline in distilled.py implements a specific two-stage strategy:
- Stage 1: Generate video at ½ target resolution using the distilled sigma schedule
- Stage 2: Upsample 2× and refine with
STAGE_2_DISTILLED_SIGMAS
This resolution-cascade approach reduces compute in the heavy diffusion steps.
Full-Model Pipeline Variants
| Pipeline | Stages | Resolution Strategy |
|---|---|---|
TI2VidOneStagePipeline (ti2vid_one_stage.py) |
Single | Generate at full target resolution directly |
Two-stage HQ (ti2vid_two_stages.py) |
Two | Both stages use full-resolution diffusion, no distilled schedule |
The full-model two-stage variant still upsamples, but runs the complete transformer and full sigma schedule at both scales — unlike the distilled version's lightweight Stage 1.
LoRA Compatibility: Distilled-Specific Weights
LoRA loading differs meaningfully. From ti2vid_two_stages_hq.py (lines 64-77):
- DistilledPipeline: Requires
--distilled-lorawith--distilled-lora-strength-stage-1/2— these LoRAs are trained specifically for the shortened distilled schedule - Full-model pipelines: Use standard
--loraarguments with regular LoRA weights trained for full diffusion schedules
You cannot interchange these LoRA types — the schedule mismatch causes degraded output.
Sampler Behavior: Ancestral vs. Deterministic
The distilled path includes conditional sampler logic. In distilled.py (lines 76-84), should_use_ancestral_sampler selects an ancestral Euler sampler for newer checkpoints (version ≥ 2.5), adding controlled stochasticity that can improve perceptual quality with fewer steps.
Full-model pipelines default to the deterministic LTX2Scheduler unless explicitly overridden.
VRAM and Performance Trade-offs
According to type hints in utils/types.py (lines 133-134), the full model requires ~28 GB VRAM when generating at target resolution.
The DistilledPipeline achieves roughly half the memory footprint through:
- Smaller checkpoint size
- Stage 1 operating at reduced resolution
- Fewer active diffusion steps
Performance summary:
| Metric | DistilledPipeline | Full-Model Pipelines |
|---|---|---|
| VRAM | ~14 GB | ~28 GB |
| Inference steps | 4-9 | 30-40 |
| Speed | Faster | Slower |
| Quality | Good (mitigated by Stage 2) | Best |
Practical Code Examples
DistilledPipeline: Two-Stage Generation
from ltx_pipelines.model_paths import ModelPaths
from ltx_pipelines.distilled import DistilledPipeline
from ltx_pipelines.utils.args import resolve_cli_params
# Enable distilled mode
args = resolve_cli_params(distilled=True)
paths = ModelPaths(
transformer=args.distilled_checkpoint_path,
video_vae=args.video_vae_path,
audio_vae=args.audio_vae_path,
spatial_upsampler=args.spatial_upsampler_path,
duration_head=args.duration_head_path,
)
pipeline = DistilledPipeline(
model_paths=paths,
spatial_upsampler_path=args.spatial_upsampler_path,
loras=args.distilled_lora, # Distilled-specific LoRA
)
video_iter, audio, tiling_cfg = pipeline(
prompt="A sunrise over a misty forest",
seed=42,
height=720,
width=1280,
frame_rate=24.0,
images=[],
)
Full-Model One-Stage Pipeline
from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from ltx_pipelines.utils.args import resolve_cli_params
# Standard (non-distilled) mode
args = resolve_cli_params(distilled=False)
paths = ModelPaths(
transformer=args.checkpoint_path, # Full checkpoint
video_vae=args.video_vae_path,
audio_vae=args.audio_vae_path,
spatial_upsampler=args.spatial_upsampler_path,
duration_head=args.duration_head_path,
)
pipeline = TI2VidOneStagePipeline(
model_paths=paths,
loras=args.lora, # Standard LoRA
)
video_iter, audio, tiling_cfg = pipeline(
prompt="A futuristic city at night",
negative_prompt="low quality, blurry",
seed=123,
height=720,
width=1280,
frame_rate=30.0,
num_inference_steps=40, # Full schedule
video_guider_params=args.video_guider_params,
audio_guider_params=args.audio_guider_params,
images=[],
)
Decision Framework: When to Choose Each
Choose DistilledPipeline when:
- VRAM is constrained (16 GB GPUs, consumer hardware)
- Iteration speed matters (prototyping, batch generation)
- Slight quality trade-offs are acceptable
Choose full-model pipelines when:
- Maximum fidelity is required (final production renders)
- Hardware budget allows 28+ GB VRAM
- You're using LoRAs trained for standard diffusion schedules
Key Source Files Reference
| File | Purpose |
|---|---|
packages/ltx-pipelines/src/ltx_pipelines/distilled.py |
DistilledPipeline implementation, ancestral sampler logic |
packages/ltx-pipelines/src/ltx_pipelines/ti2vid_one_stage.py |
Single-stage full-model pipeline |
packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py |
Two-stage full-model pipeline |
packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py |
DISTILLED_SIGMA_VALUES, STAGE_2_DISTILLED_SIGMA_VALUES |
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py |
detect_checkpoint_path, CLI argument routing |
packages/ltx-pipelines/src/ltx_pipelines/utils/types.py |
VRAM requirement documentation |
Summary
- DistilledPipeline uses compressed checkpoints, fixed short sigma schedules, and a two-stage resolution cascade to deliver ~2× memory efficiency and faster inference
- Full-model pipelines preserve complete transformer weights and full diffusion schedules for maximum quality at ~28 GB VRAM cost
- LoRA weights are not interchangeable — distilled pipelines require schedule-matched distilled LoRAs
- Source implementation lives in
distilled.pyvs.ti2vid_one_stage.py/ti2vid_two_stages.py, with configuration handled throughutils/args.py
Frequently Asked Questions
Can I switch between distilled and full-model pipelines without changing checkpoints?
No. The --distilled-checkpoint-path and --checkpoint-path arguments load fundamentally different weight formats. The detect_checkpoint_path function in utils/args.py validates checkpoint metadata and routes to the appropriate pipeline class. Attempting to load a distilled checkpoint into a full-model pipeline (or vice versa) will fail at initialization.
Why does the DistilledPipeline use an ancestral sampler for newer checkpoints?
The should_use_ancestral_sampler check in distilled.py (lines 76-84) enables stochastic sampling for checkpoint versions ≥ 2.5. This injects controlled noise during generation, which helps compensate for the reduced step count by improving sample diversity and perceptual richness that deterministic schedules might flatten with fewer iterations.
How much quality do I actually lose with the DistilledPipeline?
The trade-off is modest and context-dependent. Stage 2's 2× upsampling with STAGE_2_DISTILLED_SIGMA_VALUES recovers significant detail. In practice, the distilled output suits most preview and production use cases, while full-model pipelines excel for high-stakes final renders where diffusion artifacts must be minimized. Benchmark against your specific quality bar using identical seeds.
Can I use the same LoRA on both pipeline types?
No. Distilled LoRAs are trained on the shortened sigma schedule and expect the fixed DISTILLED_SIGMAS distribution. Standard LoRAs assume the full 30-40 step diffusion trajectory. The distilled_lora and lora argument namespaces in utils/args.py exist precisely because these training regimes produce incompatible weight adaptations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →