DistilledPipeline vs Full-Model Pipelines in LTX-2: Architecture, Performance, and Use Cases Explained

DistilledPipeline in LTX-2 uses a compressed transformer checkpoint with a reduced sigma schedule and two-stage upsampling for faster, lower-VRAM inference, while full-model pipelines run the complete diffusion process at full resolution with higher quality but greater computational cost.

LTX-2, the open-source video generation model from Lightricks, ships with two distinct pipeline families. Understanding the difference between DistilledPipeline and full-model pipelines helps you choose the right trade-off between generation speed, hardware requirements, and output quality.

What is the DistilledPipeline in LTX-2?

The DistilledPipeline is a specialized inference path designed for efficient video generation. It operates on a distilled checkpoint—compressed transformer weights that preserve most of the model's capabilities in a smaller footprint.

The pipeline is implemented in packages/ltx-pipelines/src/ltx_pipelines/distilled.py. Key characteristics include:

  • Checkpoint selection: Uses --distilled-checkpoint-path instead of the standard --checkpoint-path
  • Fixed sigma schedules: Hardcoded DISTILLED_SIGMAS and STAGE_2_DISTILLED_SIGMAS arrays (typically 9 → 4 steps)
  • Two-stage architecture: Stage 1 generates at half resolution, Stage 2 upsamples 2× with refinement

The checkpoint detection happens in utils.args.detect_checkpoint_path (lines 460-470), which routes to DistilledPipeline when distilled=True.

How Full-Model Pipelines Work in LTX-2

Full-model pipelines preserve the original LTX-2 transformer architecture without compression. Two variants exist:

TI2VidOneStagePipeline (ti2vid_one_stage.py)

  • Single diffusion pass at target resolution
  • Uses standard LTX2Scheduler with 30-40 inference steps
  • Loads weights via --checkpoint-path

Two-stage full-model pipeline (ti2vid_two_stages.py)

  • Also upsamples across two stages
  • Both stages use the full transformer and complete sigma schedule
  • No distilled-specific optimizations

Key Differences: DistilledPipeline vs Full-Model Pipelines

Factor DistilledPipeline Full-Model Pipelines
Checkpoint size ~50% of full model Original uncompressed weights
Inference steps 9 (Stage 1) + 4 (Stage 2) = 13 total 30-40 steps typical
Stage 1 resolution ½ target height/width Full target resolution (one-stage) or full model at ½ res (two-stage)
VRAM requirement Significantly lower ~28 GB for full model
Quality Slight trade-off, mitigated by Stage 2 refinement Maximum fidelity
LoRA compatibility Requires distilled-trained LoRA Standard LoRA weights

Noise Schedules and Sampler Behavior

The sigma schedule is where the pipelines diverge most sharply. In packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py (lines 15-24):


# Distilled sigma schedule - Stage 1

DISTILLED_SIGMA_VALUES = [
    14.6150, 6.3150, 2.8500, 1.2500,
    0.5610, 0.2500, 0.1100, 0.0480, 0.0020
]

# Distilled sigma schedule - Stage 2

STAGE_2_DISTILLED_SIGMA_VALUES = [
    0.4400, 0.2200, 0.0880, 0.0020
]

These fixed arrays replace the learned or computed schedules used in full-model inference. For newer distilled checkpoints (version ≥ 2.5), DistilledPipeline may switch to an ancestral Euler sampler via should_use_ancestral_sampler (lines 76-84 in distilled.py):

def should_use_ancestral_sampler(self) -> bool:
    """Determines if ancestral sampling should be used based on checkpoint version."""
    # Implementation checks checkpoint metadata for version ≥ 2.5

    ...

Full-model pipelines always use the deterministic LTX2Scheduler unless explicitly overridden.

Practical Code Examples

Running the DistilledPipeline

from ltx_pipelines.model_paths import ModelPaths
from ltx_pipelines.distilled import DistilledPipeline
from ltx_pipelines.utils.args import resolve_cli_params

# Enable distilled mode

args = resolve_cli_params(distilled=True)

paths = ModelPaths(
    transformer=args.distilled_checkpoint_path,  # Distilled checkpoint

    video_vae=args.video_vae_path,
    audio_vae=args.audio_vae_path,
    spatial_upsampler=args.spatial_upsampler_path,
    duration_head=args.duration_head_path,
)

pipeline = DistilledPipeline(
    model_paths=paths,
    spatial_upsampler_path=args.spatial_upsampler_path,
    loras=args.distilled_lora,  # Must be distilled-compatible

    lora_strength_stage_1=args.distilled_lora_strength_stage_1,
    lora_strength_stage_2=args.distilled_lora_strength_stage_2,
)

video_iter, audio, tiling_cfg = pipeline(
    prompt="A drone shot sweeping across mountain peaks",
    seed=42,
    height=720,
    width=1280,
    frame_rate=24.0,
)

Running a Full-Model Pipeline

from ltx_pipelines.ti2vid_one_stage import TI2VidOneStagePipeline
from ltx_pipelines.utils.args import resolve_cli_params

# Standard (non-distilled) configuration

args = resolve_cli_params(distilled=False)

paths = ModelPaths(
    transformer=args.checkpoint_path,  # Full checkpoint

    video_vae=args.video_vae_path,
    audio_vae=args.audio_vae_path,
    spatial_upsampler=args.spatial_upsampler_path,
    duration_head=args.duration_head_path,
)

pipeline = TI2VidOneStagePipeline(
    model_paths=paths,
    loras=args.lora,  # Standard LoRA weights

)

video_iter, audio, tiling_cfg = pipeline(
    prompt="Crystal clear ocean waves at sunset",
    negative_prompt="artifact, blur, low quality",
    seed=123,
    height=1080,
    width=1920,
    frame_rate=30.0,
    num_inference_steps=40,  # Full diffusion schedule

)

LoRA Compatibility: Critical Distinction

DistilledPipeline requires specially trained distilled LoRAs. The argument parser in ti2vid_two_stages_hq.py (lines 64-77) exposes:

  • --distilled-lora: Path to distilled LoRA weights
  • --distilled-lora-strength-stage-1: Blend factor for Stage 1
  • --distilled-lora-strength-stage-2: Blend factor for Stage 2

Standard LoRAs trained for full-model inference will not work correctly with the distilled schedule. The reduced step count and modified noise distribution require LoRA adaptation during training.

VRAM and Performance Characteristics

Memory requirements differ substantially. The type hint comment in packages/ltx-pipelines/src/ltx_pipelines/utils/types.py (lines 133-134) notes that full-model pipelines "require enough VRAM for the full model (~28 GB)".

DistilledPipeline reduces this through:

  1. Smaller checkpoint: ~50% weight reduction
  2. Half-resolution Stage 1: Reduced activation memory
  3. Fewer total steps: Less intermediate state to maintain

Exact VRAM savings depend on batch size and resolution, but distilled inference typically runs comfortably on 16-24 GB GPUs where full models require 28 GB+.

When to Choose Each Pipeline

Choose DistilledPipeline when:

  • GPU memory is constrained (16-24 GB VRAM)
  • Generation speed is critical
  • Slight quality degradation is acceptable
  • Using newer checkpoints with ancestral sampler support

Choose full-model pipelines when:

  • Maximum visual fidelity is required
  • Hardware has 28+ GB VRAM available
  • Working with standard (non-distilled) LoRA checkpoints
  • Generating at very high resolutions where every detail matters

Summary

  • DistilledPipeline uses compressed transformer weights, fixed short sigma schedules, and two-stage upsampling for efficient inference with modest quality trade-offs

  • Full-model pipelines preserve original model fidelity through complete diffusion schedules and uncompressed checkpoints, demanding more VRAM and compute time

  • Key implementation files: distilled.py (distilled logic), ti2vid_one_stage.py / ti2vid_two_stages.py (full models), utils/constants.py (sigma schedules), utils/args.py (checkpoint routing)

  • LoRA incompatibility: Distilled and full-model pipelines require separately trained LoRA weights due to schedule differences

  • Hardware planning: DistilledPipeline typically runs on 16-24 GB GPUs; full models need ~28 GB

Frequently Asked Questions

Can I switch between distilled and full-model pipelines without changing checkpoints?

No. The checkpoint formats are fundamentally different. DistilledPipeline requires a distilled checkpoint loaded via --distilled-checkpoint-path, while full-model pipelines use --checkpoint-path. The detect_checkpoint_path function in utils/args.py routes to the appropriate pipeline based on this argument.

Does the two-stage full-model pipeline offer the same speed benefits as DistilledPipeline?

No. While both architectures use two stages, the full-model two-stage pipeline runs the complete transformer and full sigma schedule at each stage. DistilledPipeline achieves its speedup through the compressed checkpoint, reduced step count, and half-resolution Stage 1 processing.

Why does DistilledPipeline need a special LoRA?

The distilled sigma schedule (9 → 4 steps) has a different noise distribution and denoising trajectory than the standard 30-40 step schedule. LoRA weights adapt the transformer's behavior to specific noise levels, so weights trained for full-model schedules perform poorly when the inference steps are compressed. Use --distilled-lora with matching distilled training.

Is the ancestral sampler in newer distilled checkpoints mandatory?

The should_use_ancestral_sampler method in distilled.py (lines 76-84) checks checkpoint metadata to determine sampler selection. For checkpoints version 2.5 and above, it enables ancestral Euler sampling automatically. This is not mandatory—it's an optimization tuned for newer distilled weights—but overriding it requires modifying the pipeline initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →