LTX-2 Quantization Policies and GPU Support: Complete Implementation Guide
LTX-2 offers four quantization policies—fp8-cast, fp8-scaled-mm, nvfp4-cast, and nvfp4-prequant—with FP8 variants requiring NVIDIA Ada Lovelace or Hopper GPUs (RTX 4090, A100, H100), while NVFP4 policies work on any BF16-capable CUDA GPU.
LTX-2 quantization policies reduce memory bandwidth and accelerate inference through low-precision compute formats. This guide covers all available policies in ltx‑pipelines/src/ltx_pipelines/utils/quantization_factory.py, their hardware requirements, and how to configure them via Python or CLI.
Available Quantization Policies in LTX-2
The QuantizationKind enum defines four distinct policies. Each trades numerical precision for speed differently and carries specific checkpoint and GPU requirements.
FP8-Cast: On-the-Fly FP8-to-FP16/FP32 Conversion
fp8-cast casts activations from FP8 to higher precision formats during inference. This policy minimizes accuracy loss while still leveraging FP8 memory bandwidth.
- Requires checkpoint: Yes
- GPU support: NVIDIA GPUs with native FP8 hardware (Ada Lovelace RTX 4090, A100, H100)
The casting happens in the QuantizationPolicy returned by QuantizationKind.FP8_CAST.to_policy(checkpoint_path=...).
FP8-Scaled-MM: FP8 Matrix Multiplication with Scaling
fp8-scaled-mm performs matrix multiplication directly in FP8 using scaling factors for numerical stability. This provides maximum throughput for FP8-capable hardware.
- Requires checkpoint: Yes
- GPU support: Same as
fp8-cast(native FP8 hardware required)
Both FP8 policies rely on the same hardware capability detection in ltx_trainer/src/ltx_trainer/quantization.py.
NVFP4-Cast: Runtime 4-Bit Emulation
nvfp4-cast emulates 4-bit precision by casting BF16 tensors to NVFP4 format on the GPU. Unlike FP8 policies, this requires no pre-quantized weights.
- Requires checkpoint: No (operates on-the-fly)
- GPU support: Any CUDA GPU supporting BF16 (most RTX 30xx/40xx series)
This policy suits rapid experimentation without model conversion steps.
NVFP4-Prequant: Static 4-Bit Weight Quantization
nvfp4-prequant pre-quantizes model weights to NVFP4 before inference begins. This eliminates runtime quantization overhead.
- Requires checkpoint: Yes (weights quantized once at load time)
- GPU support: Same BF16-compatible GPUs as
nvfp4-cast
GPU Compatibility and Hardware Validation
How LTX-2 Validates GPU Capabilities
The quantization factory itself does not query hardware. Validation occurs in two locations:
ltx_pipelines/utils/args.py— CLI argument parsing raisesValueErrorfor incompatible devicesltx_trainer/src/ltx_trainer/trainer.py— Training runtime enforces policy-device alignment
Both use torch.cuda.get_device_capability() and torch.cuda.is_available() to check requirements. Attempting FP8 on unsupported GPUs (RTX 20xx, non-CUDA devices) fails immediately with:
ValueError: fp8-cast quantization requires a CUDA device with FP8 support.
GPU Support Matrix
| Policy | Minimum GPU | Native FP8 Required | Checkpoint Needed |
|---|---|---|---|
fp8-cast |
RTX 4090, A100, H100 | Yes | Yes |
fp8-scaled-mm |
RTX 4090, A100, H100 | Yes | Yes |
nvfp4-cast |
RTX 30xx/40xx (BF16-capable) | No | No |
nvfp4-prequant |
RTX 30xx/40xx (BF16-capable) | No | Yes |
Using LTX-2 Quantization Policies Programmatically
Import QuantizationKind from the factory and convert to a policy with optional checkpoint path:
from ltx_pipelines.utils.quantization_factory import QuantizationKind
from ltx_pipelines.ti2vid_two_stages import ti2vid_two_stages
# FP8-cast with checkpoint (requires RTX 4090/A100/H100)
fp8_policy = QuantizationKind.FP8_CAST.to_policy(
checkpoint_path="checkpoints/model.ckpt"
)
# NVFP4-cast without checkpoint (works on most modern GPUs)
nvfp4_policy = QuantizationKind.NVFP4_CAST.to_policy()
# Attach to pipeline
pipeline = ti2vid_two_stages(
model_paths={"base": "checkpoints/base.safetensors"},
quantization=fp8_policy, # or nvfp4_policy
batch_size=1,
)
output = pipeline.run(input_video)
The to_policy() method returns a QuantizationPolicy object consumed by pipeline builders throughout ltx_pipelines.
Command-Line Usage
Pass --quantization with the enum value name (hyphenated lowercase):
# FP8-scaled-mm: fastest on H100/A100/RTX 4090
uv run python -m ltx_pipelines.ti2vid_two_stages \
--checkpoint-path checkpoints/model.ckpt \
--quantization fp8-scaled-mm \
--output out.mp4 \
input.mp4
# NVFP4-cast: no checkpoint, broader GPU support
uv run python -m ltx_pipelines.ti2vid_two_stages \
--quantization nvfp4-cast \
--output out.mp4 \
input.mp4
# NVFP4-prequant: faster inference, requires checkpoint
uv run python -m ltx_pipelines.ti2vid_two_stages \
--checkpoint-path checkpoints/model.ckpt \
--quantization nvfp4-prequant \
--output out.mp4 \
input.mp4
The CLI parser in ltx_pipelines/utils/args.py validates checkpoint presence and raises descriptive errors for mismatched configurations.
Key Source Files for Quantization
| File | Purpose |
|---|---|
ltx‑pipelines/src/ltx_pipelines/utils/quantization_factory.py |
QuantizationKind enum and policy dispatch |
ltx‑pipelines/src/ltx_pipelines/utils/args.py |
CLI parsing and validation logic |
ltx‑trainer/src/ltx_trainer/trainer.py |
Runtime policy enforcement |
ltx‑trainer/src/ltx_trainer/quantization.py |
Low-level quantization routines and device checks |
ltx‑pipelines/src/ltx_pipelines/utils/gpu_model.py |
GPU memory management context |
Summary
- Four policies:
fp8-cast,fp8-scaled-mm,nvfp4-cast,nvfp4-prequantinquantization_factory.py - FP8 policies require NVIDIA GPUs with native FP8 (RTX 4090, A100, H100) and a checkpoint
- NVFP4 policies work on any BF16-capable CUDA GPU;
nvfp4-castneeds no checkpoint - Hardware validation occurs in
args.pyandtrainer.pyusingtorch.cudaAPIs - API design:
QuantizationKind.to_policy()produces objects for programmatic use;--quantizationflag handles CLI
Frequently Asked Questions
Can I use LTX-2 FP8 quantization on an RTX 3080?
No. FP8 policies require GPUs with native FP8 hardware support. The RTX 3080 lacks this capability. Use nvfp4-cast or nvfp4-prequant instead—these work on any BF16-capable GPU including the RTX 3080.
Why does nvfp4-cast not require a checkpoint while nvfp4-prequant does?
nvfp4-cast performs 4-bit casting dynamically during inference, so original weights remain in higher precision. nvfp4-prequant converts weights to 4-bit once at load time, requiring the checkpoint to store and reload quantized parameters.
How do I detect which quantization policies my GPU supports at runtime?
Check torch.cuda.get_device_capability() before policy selection. FP8 requires capability ≥8.9 (Hopper/Ada). Alternatively, attempt instantiation and catch the ValueError raised by args.py or trainer.py validation.
What is the performance difference between fp8-cast and fp8-scaled-mm?
Both use FP8 hardware, but fp8-scaled-mm executes matrix multiplication natively in FP8 with scaling factors, maximizing throughput. fp8-cast converts to FP16/FP32 for compute, trading some speed for broader compatibility within FP8-capable hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →