LTX-2 DistilledPipeline vs. Full-Model Pipelines: When to Use Which for Video Generation

Use the LTX-2 DistilledPipeline when you need faster inference with lower VRAM (roughly half the memory), and switch to full-model pipelines when maximum quality is your priority.

The LTX-2 video generation framework from Lightricks offers two distinct execution paths: a memory-efficient DistilledPipeline that runs a compressed transformer with fewer denoising steps, and full-model pipelines that preserve the complete diffusion process. Understanding their architectural differences helps you pick the right tool for your hardware constraints and quality requirements.

Core Architectural Differences

Checkpoint Types and Model Size

The fundamental split begins with which weights you load. In utils/args.py, the detect_checkpoint_path function (lines 460-470) inspects metadata to determine whether you've provided a distilled checkpoint:

  • DistilledPipeline: Loads --distilled-checkpoint-path containing compressed transformer weights (~½ the size of full weights)
  • Full-model pipelines: Loads --checkpoint-path with the complete non-distilled transformer

This size reduction directly translates to lower memory pressure during inference.

Noise Schedule: Fixed Short vs. Configurable Long

The sigma schedules are hardcoded differently between the two approaches. utils/constants.py (lines 15-24) defines the distilled schedule:


# From utils/constants.py

DISTILLED_SIGMA_VALUES = [14.615, 6.315, 2.865, 1.340, 0.615]  # 5 values, typically 9→4 steps

STAGE_2_DISTILLED_SIGMA_VALUES = [0.615, 0.291, 0.125, 0.029]   # Even shorter for refinement
  • DistilledPipeline: Uses fixed DISTILLED_SIGMAS — typically 9 steps reduced to 4 effective steps
  • Full-model pipelines: Inherit 30-40 step schedules from checkpoint metadata via DiffusionStage defaults

Stage Layout: Two-Stage with Upsampling

DistilledPipeline Structure

The DistilledPipeline in distilled.py implements a specific two-stage strategy:

  1. Stage 1: Generate video at ½ target resolution using the distilled sigma schedule
  2. Stage 2: Upsample 2× and refine with STAGE_2_DISTILLED_SIGMAS

This resolution-cascade approach reduces compute in the heavy diffusion steps.

Full-Model Pipeline Variants

Pipeline Stages Resolution Strategy
TI2VidOneStagePipeline (ti2vid_one_stage.py) Single Generate at full target resolution directly
Two-stage HQ (ti2vid_two_stages.py) Two Both stages use full-resolution diffusion, no distilled schedule

The full-model two-stage variant still upsamples, but runs the complete transformer and full sigma schedule at both scales — unlike the distilled version's lightweight Stage 1.

LoRA Compatibility: Distilled-Specific Weights

LoRA loading differs meaningfully. From ti2vid_two_stages_hq.py (lines 64-77):

  • DistilledPipeline: Requires --distilled-lora with --distilled-lora-strength-stage-1/2 — these LoRAs are trained specifically for the shortened distilled schedule
  • Full-model pipelines: Use standard --lora arguments with regular LoRA weights trained for full diffusion schedules

You cannot interchange these LoRA types — the schedule mismatch causes degraded output.

Sampler Behavior: Ancestral vs. Deterministic

The distilled path includes conditional sampler logic. In distilled.py (lines 76-84), should_use_ancestral_sampler selects an ancestral Euler sampler for newer checkpoints (version ≥ 2.5), adding controlled stochasticity that can improve perceptual quality with fewer steps.

Full-model pipelines default to the deterministic LTX2Scheduler unless explicitly overridden.

VRAM and Performance Trade-offs

According to type hints in utils/types.py (lines 133-134), the full model requires ~28 GB VRAM when generating at target resolution.

The DistilledPipeline achieves roughly half the memory footprint through:

  • Smaller checkpoint size
  • Stage 1 operating at reduced resolution
  • Fewer active diffusion steps

Performance summary:

Metric DistilledPipeline Full-Model Pipelines
VRAM ~14 GB ~28 GB
Inference steps 4-9 30-40
Speed Faster Slower
Quality Good (mitigated by Stage 2) Best

Practical Code Examples

DistilledPipeline: Two-Stage Generation

from ltx_pipelines.model_paths import ModelPaths
from ltx_pipelines.distilled import DistilledPipeline
from ltx_pipelines.utils.args import resolve_cli_params

# Enable distilled mode

args = resolve_cli_params(distilled=True)

paths = ModelPaths(
    transformer=args.distilled_checkpoint_path,
    video_vae=args.video_vae_path,
    audio_vae=args.audio_vae_path,
    spatial_upsampler=args.spatial_upsampler_path,
    duration_head=args.duration_head_path,
)

pipeline = DistilledPipeline(
    model_paths=paths,
    spatial_upsampler_path=args.spatial_upsampler_path,
    loras=args.distilled_lora,  # Distilled-specific LoRA

)

video_iter, audio, tiling_cfg = pipeline(
    prompt="A sunrise over a misty forest",
    seed=42,
    height=720,
    width=1280,
    frame_rate=24.0,
    images=[],
)

Full-Model One-Stage Pipeline

from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from ltx_pipelines.utils.args import resolve_cli_params

# Standard (non-distilled) mode

args = resolve_cli_params(distilled=False)

paths = ModelPaths(
    transformer=args.checkpoint_path,  # Full checkpoint

    video_vae=args.video_vae_path,
    audio_vae=args.audio_vae_path,
    spatial_upsampler=args.spatial_upsampler_path,
    duration_head=args.duration_head_path,
)

pipeline = TI2VidOneStagePipeline(
    model_paths=paths,
    loras=args.lora,  # Standard LoRA

)

video_iter, audio, tiling_cfg = pipeline(
    prompt="A futuristic city at night",
    negative_prompt="low quality, blurry",
    seed=123,
    height=720,
    width=1280,
    frame_rate=30.0,
    num_inference_steps=40,  # Full schedule

    video_guider_params=args.video_guider_params,
    audio_guider_params=args.audio_guider_params,
    images=[],
)

Decision Framework: When to Choose Each

Choose DistilledPipeline when:

  • VRAM is constrained (16 GB GPUs, consumer hardware)
  • Iteration speed matters (prototyping, batch generation)
  • Slight quality trade-offs are acceptable

Choose full-model pipelines when:

  • Maximum fidelity is required (final production renders)
  • Hardware budget allows 28+ GB VRAM
  • You're using LoRAs trained for standard diffusion schedules

Key Source Files Reference

File Purpose
packages/ltx-pipelines/src/ltx_pipelines/distilled.py DistilledPipeline implementation, ancestral sampler logic
packages/ltx-pipelines/src/ltx_pipelines/ti2vid_one_stage.py Single-stage full-model pipeline
packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py Two-stage full-model pipeline
packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py DISTILLED_SIGMA_VALUES, STAGE_2_DISTILLED_SIGMA_VALUES
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py detect_checkpoint_path, CLI argument routing
packages/ltx-pipelines/src/ltx_pipelines/utils/types.py VRAM requirement documentation

Summary

  • DistilledPipeline uses compressed checkpoints, fixed short sigma schedules, and a two-stage resolution cascade to deliver ~2× memory efficiency and faster inference
  • Full-model pipelines preserve complete transformer weights and full diffusion schedules for maximum quality at ~28 GB VRAM cost
  • LoRA weights are not interchangeable — distilled pipelines require schedule-matched distilled LoRAs
  • Source implementation lives in distilled.py vs. ti2vid_one_stage.py/ti2vid_two_stages.py, with configuration handled through utils/args.py

Frequently Asked Questions

Can I switch between distilled and full-model pipelines without changing checkpoints?

No. The --distilled-checkpoint-path and --checkpoint-path arguments load fundamentally different weight formats. The detect_checkpoint_path function in utils/args.py validates checkpoint metadata and routes to the appropriate pipeline class. Attempting to load a distilled checkpoint into a full-model pipeline (or vice versa) will fail at initialization.

Why does the DistilledPipeline use an ancestral sampler for newer checkpoints?

The should_use_ancestral_sampler check in distilled.py (lines 76-84) enables stochastic sampling for checkpoint versions ≥ 2.5. This injects controlled noise during generation, which helps compensate for the reduced step count by improving sample diversity and perceptual richness that deterministic schedules might flatten with fewer iterations.

How much quality do I actually lose with the DistilledPipeline?

The trade-off is modest and context-dependent. Stage 2's 2× upsampling with STAGE_2_DISTILLED_SIGMA_VALUES recovers significant detail. In practice, the distilled output suits most preview and production use cases, while full-model pipelines excel for high-stakes final renders where diffusion artifacts must be minimized. Benchmark against your specific quality bar using identical seeds.

Can I use the same LoRA on both pipeline types?

No. Distilled LoRAs are trained on the shortened sigma schedule and expect the fixed DISTILLED_SIGMAS distribution. Standard LoRAs assume the full 30-40 step diffusion trajectory. The distilled_lora and lora argument namespaces in utils/args.py exist precisely because these training regimes produce incompatible weight adaptations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →