FP8 Quantization Backends in LTX-2: FP8-Cast and FP8-Scaled-MM Explained

LTX-2 supports two distinct FP8 quantization backends—FP8-Cast and FP8-Scaled-MM—exposed through the QuantizationKind enum in quantization_factory.py, each optimized for different checkpoint formats and inference workflows.

The LTX-2 video generation framework from Lightricks provides efficient low-precision inference through FP8 quantization. Understanding the supported FP8 quantization backends is essential for optimizing memory usage and computational performance when deploying or training models. Both backends implement the QuantizationPolicy interface but differ in how they store weights and handle scale tensors at runtime.

Supported FP8 Quantization Backends

LTX-2 defines its FP8 quantization backends as entries in the QuantizationKind enum located in packages/ltx-pipelines/src/ltx_pipelines/utils/quantization_factory.py. Each backend returns a QuantizationPolicy via the to_policy(checkpoint_path) method, determining how model weights are loaded and how FP8 arithmetic executes.

FP8-Cast Backend

The FP8-Cast backend (fp8-cast) stores linear weights in the FP8 (float8-e4m3) format and up-casts them to the input dtype during inference. This implementation folds any pre-quantized scale tensors (*_scale) into the weight matrix during model loading, eliminating the need for separate scale storage at runtime.

According to the LTX-2 source code in packages/ltx-core/src/ltx_core/quantization/fp8_cast.py, this backend is ideal for BF16 checkpoints that contain pre-quantized scale tensors but require dynamic up-casting for compatibility with mixed-precision pipelines.

FP8-Scaled-MM Backend

The FP8-Scaled-MM backend (fp8-scaled-mm) utilizes a scaled matrix-multiply kernel where FP8 weights remain paired with per-tensor weight_scale tensors. Unlike the cast backend, this approach performs de-quantization on-the-fly within the kernel, requiring checkpoints that already contain native FP8 weights and matching .weight_scale tensors.

This implementation lives in packages/ltx-core/src/ltx_core/quantization/fp8_scaled_mm.py and is optimized for scenarios where the checkpoint has been fully quantized to FP8 with explicit scale metadata.

Configuration and Usage

Both backends share a unified API through the QuantizationKind dispatcher. You can instantiate either policy programmatically or via the LTX-2 CLI.

Python API Configuration

To select a backend in Python, import QuantizationKind and call to_policy() with your checkpoint path:

from ltx_pipelines.utils.quantization_factory import QuantizationKind

checkpoint = "/path/to/my_fp8_checkpoint.safetensors"

# FP8-Cast: For BF16 checkpoints with scale tensors

policy_cast = QuantizationKind.FP8_CAST.to_policy(checkpoint)

# FP8-Scaled-MM: For native FP8 checkpoints with weight_scale tensors

policy_scaled = QuantizationKind.FP8_SCALED_MM.to_policy(checkpoint)

# Use with pipeline runners

# ltx_pipelines.ti2vid_one_stage.run(..., quantization_policy=policy_cast)

CLI Configuration

The ltx-trainer CLI accepts the backend selection via the --quantization-backend flag, which maps directly to the QuantizationKind enum strings:


# Select FP8-Cast backend

ltx-trainer train \
    --checkpoint /path/to/checkpoint.safetensors \
    --quantization-backend fp8-cast

# Select FP8-Scaled-MM backend

ltx-trainer train \
    --checkpoint /path/to/checkpoint.safetensors \
    --quantization-backend fp8-scaled-mm

Key Implementation Files

The FP8 quantization system spans several packages in the LTX-2 repository:

Summary

  • LTX-2 supports two FP8 quantization backends: FP8-Cast (up-casting) and FP8-Scaled-MM (scaled matrix-multiply).
  • FP8-Cast (fp8-cast) folds scales into weights during loading and converts FP8 to input dtype at inference, suitable for BF16 checkpoints with *_scale tensors.
  • FP8-Scaled-MM (fp8-scaled-mm) maintains separate weight_scale tensors and de-quantizes during the matrix operation, requiring native FP8 checkpoints.
  • Configuration is unified through QuantizationKind in quantization_factory.py, accessible via both Python API (QuantizationKind.FP8_CAST.to_policy()) and CLI (--quantization-backend).

Frequently Asked Questions

What is the difference between FP8-Cast and FP8-Scaled-MM?

FP8-Cast stores weights in FP8 but up-casts them to the input dtype (e.g., BF16) during inference, having folded any scale tensors into the weights at load time. FP8-Scaled-MM keeps weights in FP8 and performs a scaled matrix multiplication using separate weight_scale tensors, de-quantizing on-the-fly within the kernel. The cast backend requires less specialized kernel support, while the scaled-mm backend offers potentially faster inference on hardware with native FP8 support.

Which checkpoint format is required for each FP8 backend?

The FP8-Cast backend works with BF16 checkpoints that contain pre-quantized *_scale tensors, which it folds into the weights during initialization. The FP8-Scaled-MM backend requires checkpoints where weights are already stored as FP8 and include corresponding .weight_scale tensors for de-quantization. Attempting to use FP8-Scaled-MM with non-FP8 checkpoints will result in runtime errors.

How do I select a backend using the LTX-2 CLI?

Pass the --quantization-backend flag to ltx-trainer with the string value fp8-cast or fp8-scaled-mm. The CLI forwards this string to the QuantizationKind dispatcher in quantization_factory.py, which instantiates the appropriate QuantizationPolicy for the training or inference pipeline.

Where is the FP8 quantization logic implemented in the source code?

The backend implementations reside in packages/ltx-core/src/ltx_core/quantization/fp8_cast.py and packages/ltx-core/src/ltx_core/quantization/fp8_scaled_mm.py. The factory pattern and enum definitions are located in packages/ltx-pipelines/src/ltx_pipelines/utils/quantization_factory.py, while the policy interface is defined in packages/ltx-core/src/ltx_core/quantization/policy.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →