# How to Enable and Configure FP8 Quantization in LTX-2 for Lower Memory Usage

> Learn how to enable and configure FP8 quantization in LTX-2 to significantly reduce memory usage and preserve output quality. Save up to 31 GiB on large models with these simple steps.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-06-20

---

**LTX-2 supports FP8 (8-bit floating-point) quantization through two distinct backends—`fp8-cast` and `fp8-scaled-mm`—which can reduce the memory footprint of the 19-billion-parameter model by approximately 31 GiB while preserving output quality.**

The Lightricks/LTX-2 repository provides native support for FP8 quantization via the `ltx_trainer` and `ltx_pipelines` packages. By converting `torch.nn.Linear` parameters from bfloat16 or float32 to `torch.float8_e4m3fn`, you can run inference or training on GPUs with significantly less VRAM. This guide covers the internal mechanics, configuration options, and practical implementation methods based on the latest source code.

## Understanding FP8 Quantization Backends in LTX-2

LTX-2 implements two specialized FP8 backends that handle weight compression differently. Your choice depends on whether your checkpoint contains pre-computed scale tensors and how aggressively you need to minimize memory usage.

### FP8-Cast Backend (On-the-Fly Upcasting)

The **`fp8-cast`** backend stores model weights in FP8 format but casts them back to bfloat16 during the forward pass. This approach is implemented in `ltx_core.quantization.fp8_cast` (lines 140-210) and is exposed through the policy factory in [`quantization_factory.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization_factory.py).

**When to use:** Select this mode when working with standard BF16 checkpoints that do not contain FP8-specific scale tensors. It provides a modest memory reduction with minimal code changes.

### FP8-Scaled-MM Backend (Fused Kernels with Scale Tensors)

The **`fp8-scaled-mm`** backend utilizes FP8 weights alongside separate scale tensors and custom matrix-multiply kernels. This implementation lives in `ltx_core.quantization.fp8_scaled_mm` (lines 150-215) and requires checkpoints that contain `*.weight_scale` tensors.

**When to use:** Choose this for maximum memory efficiency (≈31 GiB savings on the 19B model) and when you have checkpoints specifically exported with FP8 scale tensors. The custom kernel provides optimized FP8 matrix multiplication without upcasting overhead.

## How FP8 Quantization Works Internally

### Precision Selection and Dtype Mapping

FP8 quantization begins with string-based precision selection. In the trainer, the `quantize_model()` function in [`packages/ltx-trainer/src/ltx_trainer/quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/quantization.py) (lines 52-70) accepts arguments like `"fp8-quanto"` or `"fp8uz-quanto"`. The helper `_get_quanto_dtype()` (lines 71-88) maps these strings to specific `quanto` dtypes:

- `"fp8-quanto"` → `qfloat8` (standard FP8)
- `"fp8uz-quanto"` → `qfloat8_e4m3fnuz` (FP8 with no zero, optimized for specific range distributions)

In the pipelines, the `QuantizationKind` enum in [`quantization_factory.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization_factory.py) (lines 17-36) dispatches to either `_build_fp8_cast_policy` or `_build_fp8_scaled_mm_policy`.

### Block-Wise GPU Quantization for Memory Efficiency

To prevent VRAM spikes during quantization, LTX-2 implements **block-wise quantization** on the GPU. The `quantize_model()` function detects the `transformer_blocks` attribute and processes each block individually via `_quantize_blockwise()`:

1. Move the transformer block to CUDA
2. Apply `optimum.quanto.quantize` to convert `Linear` layers to FP8
3. Freeze quantized weights
4. Move the block back to CPU

This technique ensures that the full model is never materialized in high precision on the GPU simultaneously, maintaining a low peak memory profile during the conversion process.

### Exclusion Patterns for Numerical Stability

Certain layers remain in higher precision to prevent numerical instability. The `EXCLUDE_PATTERNS` list in [`quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization.py) (lines 22-38) defines modules excluded from quantization, typically including projection layers and normalization layers. This list is passed directly to `quanto.quantize` to protect sensitive operations.

### Device Compatibility and Safety Checks

FP8 quantization is **not supported on Apple MPS devices**. The code explicitly checks for MPS in [`quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization.py) (lines 88-92) and raises a `ValueError` if detected, directing users to alternative quantization schemes like INT8 or INT4.

## Enabling FP8 Quantization in LTX-2

### Method 1: Trainer Python API

Import the quantization utility from `ltx_trainer` and apply it to your loaded model:

```python
import torch
from ltx_trainer.quantization import quantize_model

# Load your LTX-2 model (e.g., via ltx_trainer.model_loader)

model = load_model()

# Configure FP8 precision

quantized_model = quantize_model(
    model,
    precision="fp8-quanto",          # or "fp8uz-quanto" for the fnuz variant

    quantize_activations=False,        # Set True to quantize activations as well

    device=torch.device("cuda")        # Defaults to CUDA if available

)

```

This invokes the block-wise quantization routine in [`quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization.py) (lines 73-99), ensuring efficient memory usage during conversion.

### Method 2: Trainer CLI

Pass the `--quantization` flag when running trainer scripts:

```bash
uv run python packages/ltx-trainer/scripts/train.py \
    --quantization fp8-quanto \
    --checkpoint-path /path/to/ltx2-19b-dev.safetensors \
    --output-dir ./quantized_output

```

The script handles the flag in [`serve_captioner.py`](https://github.com/Lightricks/LTX-2/blob/main/serve_captioner.py) (lines 104-110), forwarding it to `quantize_model()`.

### Method 3: Pipeline CLI

For inference pipelines, use the high-level API with specific backend selection:

**Using FP8-Cast:**

```bash
uv run python -m ltx_pipelines.ti2vid_two_stages \
    --checkpoint-path /path/to/checkpoint.safetensors \
    --quantization fp8-cast \
    --prompt "your prompt here"

```

**Using FP8-Scaled-MM:**

```bash
uv run python -m ltx_pipelines.ti2vid_two_stages \
    --checkpoint-path /path/to/checkpoint.safetensors \
    --quantization fp8-scaled-mm \
    --prompt "your prompt here"

```

The `--quantization` argument is parsed in [`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py) (lines 374-383) and mapped to a `QuantizationPolicy` via [`quantization_factory.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization_factory.py) (lines 21-35).

### Selecting the Optimal FP8 Mode

| Mode | Checkpoint Requirements | Memory Impact | Best For |
|------|---------------------------|---------------|----------|
| **`fp8-quanto`** | Standard BF16 checkpoint | ~31 GiB reduction | Quick start, maximum compatibility |
| **`fp8uz-quanto`** | Standard BF16 checkpoint | ~31 GiB reduction | Scenarios requiring the *fnuz* (no zero) variant to avoid overflow |
| **`fp8-cast`** | BF16 checkpoint | Significant reduction, runtime BF16 | Minimal code changes, on-the-fly casting |
| **`fp8-scaled-mm`** | Must contain `*.weight_scale` tensors | Maximum reduction (custom kernels) | Production inference with pre-exported FP8 checkpoints |

## Troubleshooting FP8 Quantization Errors

**Missing Scale Tensors:** If you select `fp8-scaled-mm` but the checkpoint lacks the required `*.weight_scale` files, the loader raises an error at [`fp8_scaled_mm.py`](https://github.com/Lightricks/LTX-2/blob/main/fp8_scaled_mm.py) line 164. Resolve this by either converting the checkpoint with the `fp8-cast` pipeline first or exporting a proper FP8-scaled checkpoint.

**MPS Device Errors:** Running FP8 on Apple Silicon raises a `ValueError` (see [`quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization.py) lines 88-92). Use INT8 or INT4 quantization instead for MPS compatibility.

**Quality Degradation:** If output quality drops, verify that critical layers are excluded from quantization. Check the `EXCLUDE_PATTERNS` constant in [`quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization.py) (lines 22-38) to see which modules are protected by default, and adjust if necessary for your specific use case.

## Summary

- LTX-2 provides two FP8 backends: **`fp8-cast`** (upcasts to BF16 during inference) and **`fp8-scaled-mm`** (uses scale tensors with fused kernels).
- Enable FP8 via the trainer API (`quantize_model()`), trainer CLI (`--quantization fp8-quanto`), or pipeline CLI (`--quantization fp8-cast` or `fp8-scaled-mm`).
- Block-wise GPU quantization in [`quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization.py) ensures low peak memory during conversion by processing transformer blocks individually.
- The **`fp8-scaled-mm`** mode requires checkpoints with `*.weight_scale` tensors and provides the highest memory savings.
- FP8 is **not supported on MPS devices**; use alternative quantization methods for Apple Silicon.

## Frequently Asked Questions

### What is the difference between fp8-cast and fp8-scaled-mm in LTX-2?

**`fp8-cast`** stores weights in FP8 but casts them to bfloat16 during the forward pass, making it compatible with standard BF16 checkpoints. **`fp8-scaled-mm`** keeps weights in FP8 throughout inference and uses separate scale tensors with custom fused kernels, requiring checkpoints that contain `*.weight_scale` files but providing higher memory efficiency.

### How much memory does FP8 quantization save in LTX-2?

FP8 quantization reduces the memory footprint of the 19-billion-parameter LTX-2 model by approximately **31 GiB**. The exact savings depend on whether you use `fp8-cast` or `fp8-scaled-mm`, with the latter providing the maximum reduction by avoiding upcasting overhead.

### Can I use FP8 quantization on Apple Silicon (MPS)?

No, FP8 quantization is **not supported on MPS devices**. The `quantize_model()` function in [`quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization.py) explicitly checks for MPS devices (lines 88-92) and raises a `ValueError` if detected. For Apple Silicon, use INT8 or INT4 quantization methods instead.

### Why am I getting an error about missing weight_scale tensors?

This error occurs when you select `fp8-scaled-mm` but your checkpoint was not exported with FP8 scale tensors. According to [`fp8_scaled_mm.py`](https://github.com/Lightricks/LTX-2/blob/main/fp8_scaled_mm.py) line 164, this mode requires `*.weight_scale` files to perform scaled matrix multiplication. Convert your checkpoint using the `fp8-cast` pipeline first, or obtain a checkpoint specifically exported with FP8 scaling support.