# LTX-2 Quantization Policies and GPU Support: Complete Implementation Guide

> Explore LTX-2 quantization policies fp8-cast fp8-scaled-mm nvfp4-cast and nvfp4-prequant. Discover GPU support for Ada Lovelace Hopper and BF16-capable CUDA GPUs.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-20

---

**LTX-2 offers four quantization policies—`fp8-cast`, `fp8-scaled-mm`, `nvfp4-cast`, and `nvfp4-prequant`—with FP8 variants requiring NVIDIA Ada Lovelace or Hopper GPUs (RTX 4090, A100, H100), while NVFP4 policies work on any BF16-capable CUDA GPU.**

LTX-2 quantization policies reduce memory bandwidth and accelerate inference through low-precision compute formats. This guide covers all available policies in `ltx‑pipelines/src/ltx_pipelines/utils/quantization_factory.py`, their hardware requirements, and how to configure them via Python or CLI.

## Available Quantization Policies in LTX-2

The `QuantizationKind` enum defines four distinct policies. Each trades numerical precision for speed differently and carries specific checkpoint and GPU requirements.

### FP8-Cast: On-the-Fly FP8-to-FP16/FP32 Conversion

**`fp8-cast`** casts activations from FP8 to higher precision formats during inference. This policy minimizes accuracy loss while still leveraging FP8 memory bandwidth.

- **Requires checkpoint**: Yes
- **GPU support**: NVIDIA GPUs with native FP8 hardware (Ada Lovelace RTX 4090, A100, H100)

The casting happens in the `QuantizationPolicy` returned by `QuantizationKind.FP8_CAST.to_policy(checkpoint_path=...)`.

### FP8-Scaled-MM: FP8 Matrix Multiplication with Scaling

**`fp8-scaled-mm`** performs matrix multiplication directly in FP8 using scaling factors for numerical stability. This provides maximum throughput for FP8-capable hardware.

- **Requires checkpoint**: Yes
- **GPU support**: Same as `fp8-cast` (native FP8 hardware required)

Both FP8 policies rely on the same hardware capability detection in [`ltx_trainer/src/ltx_trainer/quantization.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/src/ltx_trainer/quantization.py).

### NVFP4-Cast: Runtime 4-Bit Emulation

**`nvfp4-cast`** emulates 4-bit precision by casting BF16 tensors to NVFP4 format on the GPU. Unlike FP8 policies, this requires no pre-quantized weights.

- **Requires checkpoint**: No (operates on-the-fly)
- **GPU support**: Any CUDA GPU supporting BF16 (most RTX 30xx/40xx series)

This policy suits rapid experimentation without model conversion steps.

### NVFP4-Prequant: Static 4-Bit Weight Quantization

**`nvfp4-prequant`** pre-quantizes model weights to NVFP4 before inference begins. This eliminates runtime quantization overhead.

- **Requires checkpoint**: Yes (weights quantized once at load time)
- **GPU support**: Same BF16-compatible GPUs as `nvfp4-cast`

## GPU Compatibility and Hardware Validation

### How LTX-2 Validates GPU Capabilities

The quantization factory itself does not query hardware. Validation occurs in two locations:

1. **[`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py)** — CLI argument parsing raises `ValueError` for incompatible devices
2. **[`ltx_trainer/src/ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/src/ltx_trainer/trainer.py)** — Training runtime enforces policy-device alignment

Both use `torch.cuda.get_device_capability()` and `torch.cuda.is_available()` to check requirements. Attempting FP8 on unsupported GPUs (RTX 20xx, non-CUDA devices) fails immediately with:

```

ValueError: fp8-cast quantization requires a CUDA device with FP8 support.

```

### GPU Support Matrix

| Policy | Minimum GPU | Native FP8 Required | Checkpoint Needed |
|--------|-------------|---------------------|-------------------|
| `fp8-cast` | RTX 4090, A100, H100 | Yes | Yes |
| `fp8-scaled-mm` | RTX 4090, A100, H100 | Yes | Yes |
| `nvfp4-cast` | RTX 30xx/40xx (BF16-capable) | No | No |
| `nvfp4-prequant` | RTX 30xx/40xx (BF16-capable) | No | Yes |

## Using LTX-2 Quantization Policies Programmatically

Import `QuantizationKind` from the factory and convert to a policy with optional checkpoint path:

```python
from ltx_pipelines.utils.quantization_factory import QuantizationKind
from ltx_pipelines.ti2vid_two_stages import ti2vid_two_stages

# FP8-cast with checkpoint (requires RTX 4090/A100/H100)

fp8_policy = QuantizationKind.FP8_CAST.to_policy(
    checkpoint_path="checkpoints/model.ckpt"
)

# NVFP4-cast without checkpoint (works on most modern GPUs)

nvfp4_policy = QuantizationKind.NVFP4_CAST.to_policy()

# Attach to pipeline

pipeline = ti2vid_two_stages(
    model_paths={"base": "checkpoints/base.safetensors"},
    quantization=fp8_policy,  # or nvfp4_policy

    batch_size=1,
)
output = pipeline.run(input_video)

```

The `to_policy()` method returns a `QuantizationPolicy` object consumed by pipeline builders throughout `ltx_pipelines`.

## Command-Line Usage

Pass `--quantization` with the enum value name (hyphenated lowercase):

```bash

# FP8-scaled-mm: fastest on H100/A100/RTX 4090

uv run python -m ltx_pipelines.ti2vid_two_stages \
    --checkpoint-path checkpoints/model.ckpt \
    --quantization fp8-scaled-mm \
    --output out.mp4 \
    input.mp4

# NVFP4-cast: no checkpoint, broader GPU support

uv run python -m ltx_pipelines.ti2vid_two_stages \
    --quantization nvfp4-cast \
    --output out.mp4 \
    input.mp4

# NVFP4-prequant: faster inference, requires checkpoint

uv run python -m ltx_pipelines.ti2vid_two_stages \
    --checkpoint-path checkpoints/model.ckpt \
    --quantization nvfp4-prequant \
    --output out.mp4 \
    input.mp4

```

The CLI parser in [`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py) validates checkpoint presence and raises descriptive errors for mismatched configurations.

## Key Source Files for Quantization

| File | Purpose |
|------|---------|
| `ltx‑pipelines/src/ltx_pipelines/utils/quantization_factory.py` | `QuantizationKind` enum and policy dispatch |
| `ltx‑pipelines/src/ltx_pipelines/utils/args.py` | CLI parsing and validation logic |
| `ltx‑trainer/src/ltx_trainer/trainer.py` | Runtime policy enforcement |
| `ltx‑trainer/src/ltx_trainer/quantization.py` | Low-level quantization routines and device checks |
| `ltx‑pipelines/src/ltx_pipelines/utils/gpu_model.py` | GPU memory management context |

## Summary

- **Four policies**: `fp8-cast`, `fp8-scaled-mm`, `nvfp4-cast`, `nvfp4-prequant` in [`quantization_factory.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization_factory.py)
- **FP8 policies** require NVIDIA GPUs with native FP8 (RTX 4090, A100, H100) and a checkpoint
- **NVFP4 policies** work on any BF16-capable CUDA GPU; `nvfp4-cast` needs no checkpoint
- **Hardware validation** occurs in [`args.py`](https://github.com/Lightricks/LTX-2/blob/main/args.py) and [`trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/trainer.py) using `torch.cuda` APIs
- **API design**: `QuantizationKind.to_policy()` produces objects for programmatic use; `--quantization` flag handles CLI

## Frequently Asked Questions

### Can I use LTX-2 FP8 quantization on an RTX 3080?

No. FP8 policies require GPUs with native FP8 hardware support. The RTX 3080 lacks this capability. Use `nvfp4-cast` or `nvfp4-prequant` instead—these work on any BF16-capable GPU including the RTX 3080.

### Why does `nvfp4-cast` not require a checkpoint while `nvfp4-prequant` does?

`nvfp4-cast` performs 4-bit casting dynamically during inference, so original weights remain in higher precision. `nvfp4-prequant` converts weights to 4-bit once at load time, requiring the checkpoint to store and reload quantized parameters.

### How do I detect which quantization policies my GPU supports at runtime?

Check `torch.cuda.get_device_capability()` before policy selection. FP8 requires capability ≥8.9 (Hopper/Ada). Alternatively, attempt instantiation and catch the `ValueError` raised by [`args.py`](https://github.com/Lightricks/LTX-2/blob/main/args.py) or [`trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/trainer.py) validation.

### What is the performance difference between `fp8-cast` and `fp8-scaled-mm`?

Both use FP8 hardware, but `fp8-scaled-mm` executes matrix multiplication natively in FP8 with scaling factors, maximizing throughput. `fp8-cast` converts to FP16/FP32 for compute, trading some speed for broader compatibility within FP8-capable hardware.