How to Use FP8 Cast Quantization via CLI in LTX-2

Add --quantization fp8-cast to any LTX-2 pipeline command to enable on-the-fly FP8 weight storage with bfloat16 inference up-casting.

The LTX-2 video generation framework supports memory-efficient inference through native FP8 quantization. The fp8-cast backend stores transformer linear weights in float8_e4m3fn format while dynamically up-casting them to bfloat16 during the forward pass, significantly reducing VRAM usage without requiring specialized FP8 compute hardware.

Understanding FP8-Cast Quantization in LTX-2

LTX-2 implements two distinct FP8 quantization backends: fp8-cast and fp8-scaled-mm. The fp8-cast approach specifically targets compatibility and ease of use by handling the quantization entirely through the model loader and runtime hooks.

When activated, the pipeline loads the checkpoint into FP8 tensors and attaches a forward hook that up-casts weights on demand. This occurs in ltx_core.quantization.fp8_cast, where the Fp8CastLinear class wraps standard linear layers and performs the dtype conversion during inference.

CLI Usage and Syntax

Basic Command Structure

The fp8-cast quantization is controlled through the --quantization flag available in all LTX-2 pipeline entry points. The flag accepts fp8-cast as a value to trigger the specific policy.

ltx-pipelines ti2vid_one_stage \
    --checkpoint-path /path/to/LTX-2.safetensors \
    --gemma-root /path/to/gemma \
    --prompt "A sunrise over a futuristic city" \
    --output-path output.mp4 \
    --quantization fp8-cast \
    --num-inference-steps 50 \
    --seed 42

This command activates the fp8-cast policy for the TI2VidOneStagePipeline, causing the model loader to convert transformer linear weights to FP8 format before inference begins.

Required Arguments

The fp8-cast policy requires access to the raw checkpoint to fold any pre-quantization scales into the FP8 weight tensors. Therefore, --checkpoint-path is mandatory when using --quantization fp8-cast. The pipeline scripts (ti2vid_one_stage, ti2vid_two_stages, hdr_ic_lora, etc.) share the same argument parser via default_1_stage_arg_parser and default_2_stage_arg_parser, making the flag available across all video generation workflows.

How the FP8-Cast Policy Works Internally

CLI Parsing and Resolution

In packages/ltx-pipelines/src/ltx_pipelines/utils/args.py, the parser defines the --quantization flag to enumerate supported policies including fp8-cast and fp8-scaled-mm. After argument parsing, the _resolve_quantization helper converts the string value into a QuantizationPolicy object by calling QuantizationKind.to_policy.


# Conceptual flow from args.py

parser.add_argument("--quantization", choices=["fp8-cast", "fp8-scaled-mm"])

# Later resolved via:

policy = QuantizationKind.to_policy(args.quantization)

Policy Construction

The QuantizationKind.FP8_CAST enum maps to ltx_core.quantization.fp8_cast.build_policy in packages/ltx-pipelines/src/ltx_pipelines/utils/quantization_factory.py. This factory creates a down-cast map (TRANSFORMER_LINEAR_DOWNCAST_MAP) that replaces original bfloat16 weights with FP8 tensors, and registers ModuleOps (UPCAST_DURING_INFERENCE) that rewrite the forward pass of each linear layer.

Runtime Execution

When the DiffusionStage receives the QuantizationPolicy, it applies the policy's module_ops and sd_ops via the LTX-Core loader. During inference, Fp8CastLinear.forward in packages/ltx-core/src/ltx_core/quantization/fp8_cast.py up-casts the FP8 weight to the input dtype (typically bfloat16) and optionally performs stochastic rounding via Triton kernels.

Advanced Configuration

Enabling Stochastic Rounding

For higher fidelity during the up-cast operation, enable stochastic rounding by setting the Triton environment variable. The build_policy function automatically detects this setting and adjusts the forward behavior accordingly.

export TORCH_TRITON_ENABLED=1

ltx-pipelines ti2vid_one_stage \
    --checkpoint-path /path/to/LTX-2.safetensors \
    --quantization fp8-cast \
    --prompt "A cinematic scene" \
    --output-path output.mp4

No additional CLI flags are required; the policy constructor handles the stochastic rounding configuration internally.

Combining with torch.compile

You can stack fp8-cast quantization with PyTorch compilation for additional speed gains. The quantization policy operates before compilation, ensuring the converted model graphs are optimized together.

ltx-pipelines ti2vid_one_stage \
    --checkpoint-path /path/to/LTX-2.safetensors \
    --quantization fp8-cast \
    --compile mode=reduce-overhead fullgraph=true \
    --prompt "Abstract digital art" \
    --output-path output.mp4

Summary

  • FP8-cast quantization in LTX-2 stores transformer weights in float8_e4m3fn and up-casts to bfloat16 during inference, reducing memory footprint.
  • Activate the feature via --quantization fp8-cast in any pipeline script using the shared argument parser from args.py.
  • The policy requires a valid --checkpoint-path to fold pre-quantization scales into the weight tensors during initialization.
  • Internally, QuantizationKind.FP8_CAST triggers build_policy in quantization_factory.py, which configures Fp8CastLinear layers to handle runtime up-casting.
  • Enable stochastic rounding by setting TORCH_TRITON_ENABLED=1 for improved numerical precision during the cast operation.

Frequently Asked Questions

What is the difference between fp8-cast and fp8-scaled-mm in LTX-2?

Fp8-cast stores weights in native FP8 format and up-casts them to bfloat16 on-the-fly during the forward pass, making it compatible with standard inference hardware. Fp8-scaled-mm uses scaled matrix multiplication operations that require specialized FP8 compute support. The fp8-cast method is generally preferred for broader hardware compatibility and simpler deployment.

Do I need a specific checkpoint format for FP8-cast quantization?

No special checkpoint format is required, but you must provide the --checkpoint-path argument when using fp8-cast. The build_policy function in ltx_core.quantization.fp8_cast reads the standard checkpoint and folds any pre-quantization scales into the FP8 weight tensors during the loading process.

Can I use FP8-cast quantization with all LTX-2 pipeline scripts?

Yes. The --quantization flag is defined in the shared argument parsers (default_1_stage_arg_parser and default_2_stage_arg_parser) used by all entry points including ti2vid_one_stage, ti2vid_two_stages, and hdr_ic_lora. This ensures consistent quantization support across text-to-video, image-to-video, and other generation workflows.

How does stochastic rounding affect FP8-cast inference quality?

Stochastic rounding, enabled via TORCH_TRITON_ENABLED=1, adds random noise during the FP8-to-bfloat16 conversion to reduce quantization bias. According to the implementation in fp8_cast.py, this Triton-based operation improves numerical fidelity when up-casting weights, particularly for sensitive visual details in video generation, with minimal impact on inference latency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →