How 8-Step Distilled Inference Works in LTX-2

LTX-2 replaces the traditional 30–40 step diffusion process with a fixed 8-step distilled inference schedule that runs eight denoising steps at low resolution followed by three refinement steps after upsampling, using hardcoded sigma tensors that eliminate the need for variable inference configuration.

LTX-2 ships with a distilled variant that trades iterative sampling for deterministic speed. This 8-step distilled inference pipeline, implemented in packages/ltx-pipelines/src/ltx_pipelines/distilled.py, relies on pre-computed sigma schedules defined in packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py to generate video in exactly two stages without exposing a num_inference_steps parameter.

Fixed Sigma Schedules in LTX-2 Distilled Inference

The distilled pipeline operates on hard-coded sigma tensors created at import time. These tensors define the noise levels for each timestep and cannot be modified at runtime, ensuring deterministic latency across generations.

Stage 1 Low-Resolution Denoising (8 Steps)

The first stage uses the DISTILLED_SIGMAS tensor containing nine values that define eight denoising intervals:

DISTILLED_SIGMAS = [1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0]

According to the source code at lines 14–22 of packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py, this schedule generates video at half the target resolution. The diffusion loop iterates exactly eight times, with each step corresponding to the interval between consecutive sigma values.

Stage 2 Upsampling Refinement (3 Steps)

After upsampling with VideoUpsampler, the second stage applies the STAGE_2_DISTILLED_SIGMAS tensor for final refinement:

STAGE_2_DISTILLED_SIGMAS = [0.909375, 0.725, 0.421875, 0.0]

This creates three additional denoising steps, bringing the total to eleven steps (eight plus three), though the architecture is referred to as 8-step distilled inference based on the primary stage. The stage 2 sigmas begin at 0.909375 to match the tail end of the stage 1 schedule, ensuring smooth transitions between resolutions.

Two-Stage Generation Pipeline Architecture

The DistilledPipeline class orchestrates generation through two distinct diffusion stages that share the same DiffusionStage object. Both stages utilize the SimpleDenoiser class from packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py, which performs non-guided denoising without classifier-free guidance (CFG) or style conditioning.

The process flows as follows:

  1. Prompt encoding converts text inputs into video and audio contexts.
  2. Stage 1 passes stage_1_sigmas to the denoiser, executing eight diffusion steps to produce a low-resolution latent.
  3. Spatial upsampling via VideoUpsampler doubles the latent resolution.
  4. Stage 2 runs stage_2_sigmas through three additional diffusion steps for refinement.
  5. Decoding converts the final latent to video and audio via VideoDecoder and AudioDecoder.

As implemented in packages/ltx-pipelines/src/ltx_pipelines/distilled.py (lines 99–101), the pipeline accepts the fixed sigma tensors as arguments to the __call__ method, ensuring reproducible behavior across runs.

Why Eight Steps? The Distillation Strategy

The eight-step schedule represents the output of a distillation process where the full 30–40 step diffusion model is compressed into a minimal step count while preserving output quality. The nine sigma values in DISTILLED_SIGMAS are specifically tuned so that each step corresponds to a meaningful denoising jump that matches the latent distribution of the full model.

This design eliminates the num_inference_steps parameter entirely for distilled models. The schedule is documented in the repository's CLAUDE.md file (lines 39–41), which confirms that distilled pipelines use a fixed schedule rather than exposing step configuration to users.

Implementing 8-Step Distilled Inference in Python

To run the distilled pipeline, import DistilledPipeline from the distilled module and initialize it with the required checkpoint paths:

from ltx_pipelines.distilled import DistilledPipeline
from ltx_pipelines.utils.args import detect_checkpoint_path, default_2_stage_distilled_arg_parser

# Locate distilled model weights

checkpoint = detect_checkpoint_path(distilled=True)

# Configure arguments

parser = default_2_stage_distilled_arg_parser()
args = parser.parse_args([
    "--distilled_checkpoint_path", checkpoint,
    "--gemma_root", "/path/to/gemma",
    "--spatial_upsampler_path", "/path/to/upsampler",
    "--prompt", "A sunrise over a calm lake",
    "--seed", "42",
    "--height", "720",
    "--width", "1280",
    "--num_frames", "24",
    "--frame_rate", "24",
])

# Initialize pipeline

pipeline = DistilledPipeline(
    distilled_checkpoint_path=args.distilled_checkpoint_path,
    gemma_root=args.gemma_root,
    spatial_upsampler_path=args.spatial_upsampler_path,
    loras=(),
)

# Execute 8-step stage 1 + 3-step stage 2

video, audio = pipeline(
    prompt=args.prompt,
    seed=args.seed,
    height=args.height,
    width=args.width,
    num_frames=args.num_frames,
    frame_rate=args.frame_rate,
)

For command-line execution, use the module entry point:

python -m ltx_pipelines.distilled \
    --distilled_checkpoint_path <ckpt> \
    --gemma_root <gemma_dir> \
    --spatial_upsampler_path <upsampler_ckpt> \
    --prompt "A sunrise over a calm lake" \
    --height 720 \
    --width 1280 \
    --num_frames 24 \
    --output_path result.mp4

Summary

  • 8-step distilled inference in LTX-2 uses fixed sigma tensors (DISTILLED_SIGMAS and STAGE_2_DISTILLED_SIGMAS) defined in packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py.
  • The pipeline runs a two-stage generation process: eight denoising steps at low resolution followed by three refinement steps after upsampling.
  • The DistilledPipeline class in packages/ltx-pipelines/src/ltx_pipelines/distilled.py does not expose num_inference_steps; the schedule is hard-coded for deterministic latency.
  • Both stages use SimpleDenoiser without classifier-free guidance, making inference significantly faster than the full model.
  • Memory is released after each call, allowing pipeline reuse without GPU state accumulation.

Frequently Asked Questions

Can I change the number of inference steps in LTX-2 distilled mode?

No. The 8-step distilled inference pipeline uses hard-coded sigma schedules that are baked into the model weights. Unlike the full diffusion model, the distilled variant in packages/ltx-pipelines/src/ltx_pipelines/distilled.py does not accept a num_inference_steps argument. The DISTILLED_SIGMAS tensor contains exactly nine values defining eight intervals, and this schedule is fixed at import time.

What is the difference between Stage 1 and Stage 2 in the distilled pipeline?

Stage 1 generates video at half the target resolution using the full DISTILLED_SIGMAS schedule (eight steps). Stage 2 begins after VideoUpsampler doubles the spatial resolution, applying the STAGE_2_DISTILLED_SIGMAS schedule (three steps) to refine details and reduce artifacts introduced during upsampling. Both stages share the same DiffusionStage object but operate on different sigma tensors.

Why does the distilled pipeline use SimpleDenoiser instead of CFG?

The distilled pipeline uses SimpleDenoiser (defined in packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py) because the distillation process encodes guidance directly into the model weights. This eliminates the need for classifier-free guidance (CFG), which typically requires doubling the UNet forward passes. The result is faster inference with fewer memory allocations while maintaining output quality comparable to guided sampling.

How do the sigma values translate to actual denoising steps?

Each sigma tensor defines the noise levels at the boundaries of denoising steps. The DISTILLED_SIGMAS tensor contains nine floating-point values, which create eight denoising intervals (steps). Similarly, STAGE_2_DISTILLED_SIGMAS contains four values creating three intervals. The diffusion loop iterates once per interval, with the model predicting noise to move from the current sigma level to the next.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →