LTX-2 Quantization Policies and GPU Support: Complete Implementation Guide

LTX-2 offers four quantization policies—fp8-cast, fp8-scaled-mm, nvfp4-cast, and nvfp4-prequant—with FP8 variants requiring NVIDIA Ada Lovelace or Hopper GPUs (RTX 4090, A100, H100), while NVFP4 policies work on any BF16-capable CUDA GPU.

LTX-2 quantization policies reduce memory bandwidth and accelerate inference through low-precision compute formats. This guide covers all available policies in ltx‑pipelines/src/ltx_pipelines/utils/quantization_factory.py, their hardware requirements, and how to configure them via Python or CLI.

Available Quantization Policies in LTX-2

The QuantizationKind enum defines four distinct policies. Each trades numerical precision for speed differently and carries specific checkpoint and GPU requirements.

FP8-Cast: On-the-Fly FP8-to-FP16/FP32 Conversion

fp8-cast casts activations from FP8 to higher precision formats during inference. This policy minimizes accuracy loss while still leveraging FP8 memory bandwidth.

  • Requires checkpoint: Yes
  • GPU support: NVIDIA GPUs with native FP8 hardware (Ada Lovelace RTX 4090, A100, H100)

The casting happens in the QuantizationPolicy returned by QuantizationKind.FP8_CAST.to_policy(checkpoint_path=...).

FP8-Scaled-MM: FP8 Matrix Multiplication with Scaling

fp8-scaled-mm performs matrix multiplication directly in FP8 using scaling factors for numerical stability. This provides maximum throughput for FP8-capable hardware.

  • Requires checkpoint: Yes
  • GPU support: Same as fp8-cast (native FP8 hardware required)

Both FP8 policies rely on the same hardware capability detection in ltx_trainer/src/ltx_trainer/quantization.py.

NVFP4-Cast: Runtime 4-Bit Emulation

nvfp4-cast emulates 4-bit precision by casting BF16 tensors to NVFP4 format on the GPU. Unlike FP8 policies, this requires no pre-quantized weights.

  • Requires checkpoint: No (operates on-the-fly)
  • GPU support: Any CUDA GPU supporting BF16 (most RTX 30xx/40xx series)

This policy suits rapid experimentation without model conversion steps.

NVFP4-Prequant: Static 4-Bit Weight Quantization

nvfp4-prequant pre-quantizes model weights to NVFP4 before inference begins. This eliminates runtime quantization overhead.

  • Requires checkpoint: Yes (weights quantized once at load time)
  • GPU support: Same BF16-compatible GPUs as nvfp4-cast

GPU Compatibility and Hardware Validation

How LTX-2 Validates GPU Capabilities

The quantization factory itself does not query hardware. Validation occurs in two locations:

  1. ltx_pipelines/utils/args.py — CLI argument parsing raises ValueError for incompatible devices
  2. ltx_trainer/src/ltx_trainer/trainer.py — Training runtime enforces policy-device alignment

Both use torch.cuda.get_device_capability() and torch.cuda.is_available() to check requirements. Attempting FP8 on unsupported GPUs (RTX 20xx, non-CUDA devices) fails immediately with:


ValueError: fp8-cast quantization requires a CUDA device with FP8 support.

GPU Support Matrix

Policy Minimum GPU Native FP8 Required Checkpoint Needed
fp8-cast RTX 4090, A100, H100 Yes Yes
fp8-scaled-mm RTX 4090, A100, H100 Yes Yes
nvfp4-cast RTX 30xx/40xx (BF16-capable) No No
nvfp4-prequant RTX 30xx/40xx (BF16-capable) No Yes

Using LTX-2 Quantization Policies Programmatically

Import QuantizationKind from the factory and convert to a policy with optional checkpoint path:

from ltx_pipelines.utils.quantization_factory import QuantizationKind
from ltx_pipelines.ti2vid_two_stages import ti2vid_two_stages

# FP8-cast with checkpoint (requires RTX 4090/A100/H100)

fp8_policy = QuantizationKind.FP8_CAST.to_policy(
    checkpoint_path="checkpoints/model.ckpt"
)

# NVFP4-cast without checkpoint (works on most modern GPUs)

nvfp4_policy = QuantizationKind.NVFP4_CAST.to_policy()

# Attach to pipeline

pipeline = ti2vid_two_stages(
    model_paths={"base": "checkpoints/base.safetensors"},
    quantization=fp8_policy,  # or nvfp4_policy

    batch_size=1,
)
output = pipeline.run(input_video)

The to_policy() method returns a QuantizationPolicy object consumed by pipeline builders throughout ltx_pipelines.

Command-Line Usage

Pass --quantization with the enum value name (hyphenated lowercase):


# FP8-scaled-mm: fastest on H100/A100/RTX 4090

uv run python -m ltx_pipelines.ti2vid_two_stages \
    --checkpoint-path checkpoints/model.ckpt \
    --quantization fp8-scaled-mm \
    --output out.mp4 \
    input.mp4

# NVFP4-cast: no checkpoint, broader GPU support

uv run python -m ltx_pipelines.ti2vid_two_stages \
    --quantization nvfp4-cast \
    --output out.mp4 \
    input.mp4

# NVFP4-prequant: faster inference, requires checkpoint

uv run python -m ltx_pipelines.ti2vid_two_stages \
    --checkpoint-path checkpoints/model.ckpt \
    --quantization nvfp4-prequant \
    --output out.mp4 \
    input.mp4

The CLI parser in ltx_pipelines/utils/args.py validates checkpoint presence and raises descriptive errors for mismatched configurations.

Key Source Files for Quantization

File Purpose
ltx‑pipelines/src/ltx_pipelines/utils/quantization_factory.py QuantizationKind enum and policy dispatch
ltx‑pipelines/src/ltx_pipelines/utils/args.py CLI parsing and validation logic
ltx‑trainer/src/ltx_trainer/trainer.py Runtime policy enforcement
ltx‑trainer/src/ltx_trainer/quantization.py Low-level quantization routines and device checks
ltx‑pipelines/src/ltx_pipelines/utils/gpu_model.py GPU memory management context

Summary

  • Four policies: fp8-cast, fp8-scaled-mm, nvfp4-cast, nvfp4-prequant in quantization_factory.py
  • FP8 policies require NVIDIA GPUs with native FP8 (RTX 4090, A100, H100) and a checkpoint
  • NVFP4 policies work on any BF16-capable CUDA GPU; nvfp4-cast needs no checkpoint
  • Hardware validation occurs in args.py and trainer.py using torch.cuda APIs
  • API design: QuantizationKind.to_policy() produces objects for programmatic use; --quantization flag handles CLI

Frequently Asked Questions

Can I use LTX-2 FP8 quantization on an RTX 3080?

No. FP8 policies require GPUs with native FP8 hardware support. The RTX 3080 lacks this capability. Use nvfp4-cast or nvfp4-prequant instead—these work on any BF16-capable GPU including the RTX 3080.

Why does nvfp4-cast not require a checkpoint while nvfp4-prequant does?

nvfp4-cast performs 4-bit casting dynamically during inference, so original weights remain in higher precision. nvfp4-prequant converts weights to 4-bit once at load time, requiring the checkpoint to store and reload quantized parameters.

How do I detect which quantization policies my GPU supports at runtime?

Check torch.cuda.get_device_capability() before policy selection. FP8 requires capability ≥8.9 (Hopper/Ada). Alternatively, attempt instantiation and catch the ValueError raised by args.py or trainer.py validation.

What is the performance difference between fp8-cast and fp8-scaled-mm?

Both use FP8 hardware, but fp8-scaled-mm executes matrix multiplication natively in FP8 with scaling factors, maximizing throughput. fp8-cast converts to FP16/FP32 for compute, trading some speed for broader compatibility within FP8-capable hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →