# Distilled vs. Full Model Checkpoints in LTX-2: A Complete Technical Comparison

> Explore LTX-2's distilled vs full model checkpoints. Discover 22B parameter full models for fidelity and 3-4x faster distilled models with optimized inference.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-14

---

**LTX-2 provides two checkpoint families: full-model checkpoints with 22B parameters and highest visual fidelity, and distilled checkpoints that deliver 3-4× faster inference through a fixed 8-step sigma schedule, no CFG guidance, and simplified denoising.**

The Lightricks LTX-2 video generation model ships with dual checkpoint architectures designed for different deployment scenarios. Understanding the differences between distilled and full model checkpoints helps you optimize for quality, speed, or GPU memory constraints. This guide breaks down the technical implementation in the open-source `Lightricks/LTX-2` repository.

## Checkout Identification and File Naming

The first visible difference appears in checkpoint filenames:

| Checkpoint Type | Filename Pattern |
|---------------|----------------|
| **Full model** | `...-transformer-bf16.safetensors` (no "distilled" marker) |
| **Distilled** | `...-distilled-transformer-bf16.safetensors` |

The `ModelPaths` class in [`packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py) parses these paths to determine which pipeline logic to activate. Pass `--checkpoint-path` for full checkpoints or `--distilled-checkpoint-path` for distilled variants.

## Architecture and Parameter Count

Both checkpoint families share the **same underlying 22B parameter transformer architecture**. The distilled checkpoint does not reduce model size—instead, it reduces the diffusion schedule complexity through knowledge distillation training.

**Full checkpoint:**
- Maximum generation fidelity
- Supports all guidance modalities
- Highest VRAM requirements during sampling

**Distilled checkpoint:**
- Identical transformer weights, optimized inference path
- Pre-tuned sigma schedules eliminate dynamic scheduler computation
- Recommended default for most production deployments

## Sigma Schedule: The Core Technical Difference

The most significant implementation divergence lives in [`packages/ltx-pipelines/src/ltx_pipelines/distilled.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/distilled.py) and [`packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py).

### Full-Model Dynamic Scheduling

Full checkpoints generate sigmas at runtime through the sampler. The scheduler constructs a full-length diffusion timeline with dozens of denoising steps. You can customize the step count via standard sampling parameters.

### Distilled Fixed Schedules

Distilled checkpoints use **hardcoded sigma sequences** imported at module load:

```python
from ltx_pipelines.utils.constants import DISTILLED_SIGMAS, STAGE_2_DISTILLED_SIGMAS

```

These constants define:
- **Stage 1:** 8 denoising steps (`DISTILLED_SIGMAS`)
- **Stage 2:** 3 denoising steps (`STAGE_2_DISTILLED_SIGMAS`)

This 3-4× reduction in steps directly translates to proportional inference speedup.

## Guidance Mechanisms: CFG and STG Availability

### Full Checkpoint Guidance

Full checkpoints support advanced guidance techniques:

- **CFG (Classifier-Free Guidance):** Compare conditional and unconditional predictions
- **STG (Style Guidance):** Fine-grained style control through negative prompting

The `TI2VidTwoStagesPipeline` expects `cfg_scale` and negative prompt arguments, applying them through standard diffusion guidance equations.

### Distilled Pipeline: SimpleDenoiser Only

The `DistilledPipeline` ([`packages/ltx-pipelines/src/ltx_pipelines/distilled.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/distilled.py)) replaces guided sampling with `SimpleDenoiser`:

```python

# Distilled pipeline ignores CFG parameters

result = pipeline(
    prompt="A robot dancing in a neon city",
    seed=42,
    height=720,
    width=1280,
    # cfg_scale=7.0  # ← silently ignored

)

```

Each step performs a **single forward pass** without conditional branching. The negative prompt and guidance scale arguments exist for API compatibility but have no effect on output.

## Sampler Selection Logic

### Version-Dependent Ancestral Sampling

For checkpoints version 2.5 and above, the distilled pipeline makes version-aware sampler decisions. The function `should_use_ancestral_sampler()` in [`distilled.py`](https://github.com/Lightricks/LTX-2/blob/main/distilled.py) handles this:

```python
def should_use_ancestral_sampler(transformer_path: str) -> bool:
    return detect_model_version(transformer_path) >= ANCESTRAL_SAMPLER_SINCE_VERSION

# ANCESTRAL_SAMPLER_SINCE_VERSION = (2, 5)

```

**Sampling behavior by checkpoint version:**

| Stage | Version < 2.5 | Version ≥ 2.5 |
|-------|--------------|---------------|
| Stage 1 | Plain Euler | **EulerAncestralDiffusionStep** |
| Stage 2 | Plain Euler | Plain Euler (short schedule always) |

The ancestral sampler for stage 1 adds stochasticity that improves motion coherence in distilled outputs. Full checkpoints can optionally enable ancestral sampling through manual configuration.

## Checkpoint Contents and Metadata

| Component | Full Checkpoint | Distilled Checkpoint |
|-----------|---------------|----------------------|
| Transformer weights | Full 22B weights | Distilled-compatible weights |
| VAE | Standard VAE | Optional distilled DiffVAE |
| Audio VAE | Included | Excluded or distilled variant |
| Duration head | Optional | Optional |

The `ModelPaths` loader reads checkpoint metadata to auto-select compatible VAE types, ensuring the distilled pipeline pairs with optimized decoder weights when available.

## Practical Usage Examples

### Full Checkpoint with CFG

```python
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines import ti2vid_two_stages

model_paths = ModelPaths.from_monolith(
    checkpoint_path="models/ltx-2.5/ltx-2.5-22b-transformer-bf16.safetensors"
)

pipeline = ti2vid_two_stages.TI2VidTwoStagesPipeline(
    model_paths=model_paths,
    spatial_upsampler_path="models/ltx-2.5/spatial_upscaler.safetensors",
    loras=[],
    distilled=False,  # Explicit full-mode flag

)

result = pipeline(
    prompt="Aerial view of ocean waves crashing on cliffs",
    seed=42,
    height=720,
    width=1280,
    frame_rate=30.0,
    images=[],
    cfg_scale=7.0,  # Active guidance

    negative_prompt="blurry, low quality"  # Active negative conditioning

)

```

### Distilled Checkpoint for Fast Inference

```python
from ltx_pipelines.utils.model_paths import ModelPaths
from ltx_pipelines import distilled

model_paths = ModelPaths.from_monolith(
    distilled_checkpoint_path="models/ltx-2.5/ltx-2.5-22b-distilled-transformer-bf16.safetensors"
)

pipeline = distilled.DistilledPipeline(
    model_paths=model_paths,
    spatial_upsampler_path="models/ltx-2.5/spatial_upscaler.safetensors",
    loras=[],  # Optional: add distilled LoRA for stage-2 refinement

)

result = pipeline(
    prompt="Aerial view of ocean waves crashing on cliffs",
    seed=42,
    height=720,
    width=1280,
    frame_rate=30.0,
    images=[],
    # No cfg_scale—distilled runs guidance-free

)

# Optional: load distilled LoRA for quality refinement

# pipeline.load_distilled_lora("path/to/distilled-lora.safetensors")

```

## CLI Flag Reference

| Purpose | Full Checkpoint | Distilled Checkpoint |
|---------|---------------|----------------------|
| Main path | `--checkpoint-path` | `--distilled-checkpoint-path` (or `--transformer-path` in unified CLI) |
| LoRA loading | `--lora` | `--distilled-lora` (stage-2 specific) |
| Guidance scale | `--cfg-scale` (functional) | `--cfg-scale` (ignored, API compatibility only) |

## Performance and Quality Trade-offs

| Metric | Full Checkpoint | Distilled Checkpoint |
|--------|---------------|----------------------|
| Inference speed | Baseline | **3-4× faster** |
| VRAM during sampling | Higher (full schedule state) | Lower (fixed short schedule) |
| Visual fidelity | Highest | Minor degradation, often imperceptible |
| Motion smoothness | Excellent | Excellent (with ancestral sampler ≥v2.5) |
| Control precision | Maximum (CFG + STG) | Limited (prompt-only) |

The LTX-2 documentation ([`docs/pipelines.md`](https://github.com/Lightricks/LTX-2/blob/main/docs/pipelines.md)) officially recommends distilled checkpoints as the default for most users, reserving full checkpoints for maximum quality requirements or research scenarios needing guidance manipulation.

## Summary

- **File naming:** Distilled checkpoints contain "distilled" in the transformer filename; full checkpoints do not
- **Speed:** Distilled delivers 3-4× faster inference through fixed 8-step (stage 1) and 3-step (stage 2) schedules defined in [`utils/constants.py`](https://github.com/Lightricks/LTX-2/blob/main/utils/constants.py)
- **Guidance:** Full checkpoints support CFG and STG; distilled pipelines use `SimpleDenoiser` with no guidance
- **Sampling:** Distilled pipelines auto-detect checkpoint version and apply `EulerAncestralDiffusionStep` for stage 1 when version ≥ 2.5
- **Implementation:** Core logic resides in [`packages/ltx-pipelines/src/ltx_pipelines/distilled.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/distilled.py) with version detection and sigma schedule imports

## Frequently Asked Questions

### What happens if I pass a full checkpoint to the distilled pipeline?

The pipeline will attempt to load it, but behavior is undefined. The [`distilled.py`](https://github.com/Lightricks/LTX-2/blob/main/distilled.py) implementation expects specific checkpoint metadata and will fail or produce garbled output when full-model weights are supplied. Always match pipeline class to checkpoint type—use `TI2VidTwoStagesPipeline` for full checkpoints and `DistilledPipeline` for distilled checkpoints.

### Can I use CFG with distilled checkpoints by modifying the code?

No—this would require fundamental retraining. The distilled checkpoint was trained without CFG conditioning; the unconditional prediction path does not exist in the distilled model weights. The `SimpleDenoiser` class assumes single forward passes per step. Re-enabling CFG would need the full 22B training pipeline with distillation-aware guidance objectives.

### Why does the distilled pipeline use different step counts for stage 1 versus stage 2?

Stage 1 generates the coarse video structure where temporal coherence matters most; the 8-step `DISTILLED_SIGMAS` schedule with optional ancestral noise maintains motion quality. Stage 2 refines spatial details with `STAGE_2_DISTILLED_SIGMAS` (3 steps)—sufficient for detail enhancement since temporal structure is already established. This asymmetric design optimizes the speed/quality trade-off.

### How do I detect whether my checkpoint supports the ancestral sampler?

The `detect_model_version()` function parses checkpoint paths to extract version tuples. Checkpoints with "2.5" or higher in the filename return `(2, 5)` or above, triggering `should_use_ancestral_sampler()` to return `True`. You can verify manually: `python -c "from ltx_pipelines.distilled import detect_model_version; print(detect_model_version('your/path.safetensors'))"`