How to Enable and Configure FP8 Quantization in LTX-2 for Lower Memory Usage
LTX-2 supports FP8 (8-bit floating-point) quantization through two distinct backends—fp8-cast and fp8-scaled-mm—which can reduce the memory footprint of the 19-billion-parameter model by approximately 31 GiB while preserving output quality.
The Lightricks/LTX-2 repository provides native support for FP8 quantization via the ltx_trainer and ltx_pipelines packages. By converting torch.nn.Linear parameters from bfloat16 or float32 to torch.float8_e4m3fn, you can run inference or training on GPUs with significantly less VRAM. This guide covers the internal mechanics, configuration options, and practical implementation methods based on the latest source code.
Understanding FP8 Quantization Backends in LTX-2
LTX-2 implements two specialized FP8 backends that handle weight compression differently. Your choice depends on whether your checkpoint contains pre-computed scale tensors and how aggressively you need to minimize memory usage.
FP8-Cast Backend (On-the-Fly Upcasting)
The fp8-cast backend stores model weights in FP8 format but casts them back to bfloat16 during the forward pass. This approach is implemented in ltx_core.quantization.fp8_cast (lines 140-210) and is exposed through the policy factory in quantization_factory.py.
When to use: Select this mode when working with standard BF16 checkpoints that do not contain FP8-specific scale tensors. It provides a modest memory reduction with minimal code changes.
FP8-Scaled-MM Backend (Fused Kernels with Scale Tensors)
The fp8-scaled-mm backend utilizes FP8 weights alongside separate scale tensors and custom matrix-multiply kernels. This implementation lives in ltx_core.quantization.fp8_scaled_mm (lines 150-215) and requires checkpoints that contain *.weight_scale tensors.
When to use: Choose this for maximum memory efficiency (≈31 GiB savings on the 19B model) and when you have checkpoints specifically exported with FP8 scale tensors. The custom kernel provides optimized FP8 matrix multiplication without upcasting overhead.
How FP8 Quantization Works Internally
Precision Selection and Dtype Mapping
FP8 quantization begins with string-based precision selection. In the trainer, the quantize_model() function in packages/ltx-trainer/src/ltx_trainer/quantization.py (lines 52-70) accepts arguments like "fp8-quanto" or "fp8uz-quanto". The helper _get_quanto_dtype() (lines 71-88) maps these strings to specific quanto dtypes:
"fp8-quanto"→qfloat8(standard FP8)"fp8uz-quanto"→qfloat8_e4m3fnuz(FP8 with no zero, optimized for specific range distributions)
In the pipelines, the QuantizationKind enum in quantization_factory.py (lines 17-36) dispatches to either _build_fp8_cast_policy or _build_fp8_scaled_mm_policy.
Block-Wise GPU Quantization for Memory Efficiency
To prevent VRAM spikes during quantization, LTX-2 implements block-wise quantization on the GPU. The quantize_model() function detects the transformer_blocks attribute and processes each block individually via _quantize_blockwise():
- Move the transformer block to CUDA
- Apply
optimum.quanto.quantizeto convertLinearlayers to FP8 - Freeze quantized weights
- Move the block back to CPU
This technique ensures that the full model is never materialized in high precision on the GPU simultaneously, maintaining a low peak memory profile during the conversion process.
Exclusion Patterns for Numerical Stability
Certain layers remain in higher precision to prevent numerical instability. The EXCLUDE_PATTERNS list in quantization.py (lines 22-38) defines modules excluded from quantization, typically including projection layers and normalization layers. This list is passed directly to quanto.quantize to protect sensitive operations.
Device Compatibility and Safety Checks
FP8 quantization is not supported on Apple MPS devices. The code explicitly checks for MPS in quantization.py (lines 88-92) and raises a ValueError if detected, directing users to alternative quantization schemes like INT8 or INT4.
Enabling FP8 Quantization in LTX-2
Method 1: Trainer Python API
Import the quantization utility from ltx_trainer and apply it to your loaded model:
import torch
from ltx_trainer.quantization import quantize_model
# Load your LTX-2 model (e.g., via ltx_trainer.model_loader)
model = load_model()
# Configure FP8 precision
quantized_model = quantize_model(
model,
precision="fp8-quanto", # or "fp8uz-quanto" for the fnuz variant
quantize_activations=False, # Set True to quantize activations as well
device=torch.device("cuda") # Defaults to CUDA if available
)
This invokes the block-wise quantization routine in quantization.py (lines 73-99), ensuring efficient memory usage during conversion.
Method 2: Trainer CLI
Pass the --quantization flag when running trainer scripts:
uv run python packages/ltx-trainer/scripts/train.py \
--quantization fp8-quanto \
--checkpoint-path /path/to/ltx2-19b-dev.safetensors \
--output-dir ./quantized_output
The script handles the flag in serve_captioner.py (lines 104-110), forwarding it to quantize_model().
Method 3: Pipeline CLI
For inference pipelines, use the high-level API with specific backend selection:
Using FP8-Cast:
uv run python -m ltx_pipelines.ti2vid_two_stages \
--checkpoint-path /path/to/checkpoint.safetensors \
--quantization fp8-cast \
--prompt "your prompt here"
Using FP8-Scaled-MM:
uv run python -m ltx_pipelines.ti2vid_two_stages \
--checkpoint-path /path/to/checkpoint.safetensors \
--quantization fp8-scaled-mm \
--prompt "your prompt here"
The --quantization argument is parsed in ltx_pipelines/utils/args.py (lines 374-383) and mapped to a QuantizationPolicy via quantization_factory.py (lines 21-35).
Selecting the Optimal FP8 Mode
| Mode | Checkpoint Requirements | Memory Impact | Best For |
|---|---|---|---|
fp8-quanto |
Standard BF16 checkpoint | ~31 GiB reduction | Quick start, maximum compatibility |
fp8uz-quanto |
Standard BF16 checkpoint | ~31 GiB reduction | Scenarios requiring the fnuz (no zero) variant to avoid overflow |
fp8-cast |
BF16 checkpoint | Significant reduction, runtime BF16 | Minimal code changes, on-the-fly casting |
fp8-scaled-mm |
Must contain *.weight_scale tensors |
Maximum reduction (custom kernels) | Production inference with pre-exported FP8 checkpoints |
Troubleshooting FP8 Quantization Errors
Missing Scale Tensors: If you select fp8-scaled-mm but the checkpoint lacks the required *.weight_scale files, the loader raises an error at fp8_scaled_mm.py line 164. Resolve this by either converting the checkpoint with the fp8-cast pipeline first or exporting a proper FP8-scaled checkpoint.
MPS Device Errors: Running FP8 on Apple Silicon raises a ValueError (see quantization.py lines 88-92). Use INT8 or INT4 quantization instead for MPS compatibility.
Quality Degradation: If output quality drops, verify that critical layers are excluded from quantization. Check the EXCLUDE_PATTERNS constant in quantization.py (lines 22-38) to see which modules are protected by default, and adjust if necessary for your specific use case.
Summary
- LTX-2 provides two FP8 backends:
fp8-cast(upcasts to BF16 during inference) andfp8-scaled-mm(uses scale tensors with fused kernels). - Enable FP8 via the trainer API (
quantize_model()), trainer CLI (--quantization fp8-quanto), or pipeline CLI (--quantization fp8-castorfp8-scaled-mm). - Block-wise GPU quantization in
quantization.pyensures low peak memory during conversion by processing transformer blocks individually. - The
fp8-scaled-mmmode requires checkpoints with*.weight_scaletensors and provides the highest memory savings. - FP8 is not supported on MPS devices; use alternative quantization methods for Apple Silicon.
Frequently Asked Questions
What is the difference between fp8-cast and fp8-scaled-mm in LTX-2?
fp8-cast stores weights in FP8 but casts them to bfloat16 during the forward pass, making it compatible with standard BF16 checkpoints. fp8-scaled-mm keeps weights in FP8 throughout inference and uses separate scale tensors with custom fused kernels, requiring checkpoints that contain *.weight_scale files but providing higher memory efficiency.
How much memory does FP8 quantization save in LTX-2?
FP8 quantization reduces the memory footprint of the 19-billion-parameter LTX-2 model by approximately 31 GiB. The exact savings depend on whether you use fp8-cast or fp8-scaled-mm, with the latter providing the maximum reduction by avoiding upcasting overhead.
Can I use FP8 quantization on Apple Silicon (MPS)?
No, FP8 quantization is not supported on MPS devices. The quantize_model() function in quantization.py explicitly checks for MPS devices (lines 88-92) and raises a ValueError if detected. For Apple Silicon, use INT8 or INT4 quantization methods instead.
Why am I getting an error about missing weight_scale tensors?
This error occurs when you select fp8-scaled-mm but your checkpoint was not exported with FP8 scale tensors. According to fp8_scaled_mm.py line 164, this mode requires *.weight_scale files to perform scaled matrix multiplication. Convert your checkpoint using the fp8-cast pipeline first, or obtain a checkpoint specifically exported with FP8 scaling support.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →