# How to Use FP8 Cast Quantization via CLI in LTX-2

> Learn to use FP8 Cast quantization via CLI in LTX-2. Enable on-the-fly FP8 weight storage with bfloat16 inference up-casting by adding --quantization fp8-cast to your LTX-2 pipeline command.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-06-21

---

**Add `--quantization fp8-cast` to any LTX-2 pipeline command to enable on-the-fly FP8 weight storage with bfloat16 inference up-casting.**

The **LTX-2** video generation framework supports memory-efficient inference through native FP8 quantization. The **fp8-cast** backend stores transformer linear weights in `float8_e4m3fn` format while dynamically up-casting them to bfloat16 during the forward pass, significantly reducing VRAM usage without requiring specialized FP8 compute hardware.

## Understanding FP8-Cast Quantization in LTX-2

LTX-2 implements two distinct FP8 quantization backends: **fp8-cast** and **fp8-scaled-mm**. The fp8-cast approach specifically targets compatibility and ease of use by handling the quantization entirely through the model loader and runtime hooks.

When activated, the pipeline loads the checkpoint into FP8 tensors and attaches a forward hook that up-casts weights on demand. This occurs in `ltx_core.quantization.fp8_cast`, where the `Fp8CastLinear` class wraps standard linear layers and performs the dtype conversion during inference.

## CLI Usage and Syntax

### Basic Command Structure

The fp8-cast quantization is controlled through the `--quantization` flag available in all LTX-2 pipeline entry points. The flag accepts `fp8-cast` as a value to trigger the specific policy.

```bash
ltx-pipelines ti2vid_one_stage \
    --checkpoint-path /path/to/LTX-2.safetensors \
    --gemma-root /path/to/gemma \
    --prompt "A sunrise over a futuristic city" \
    --output-path output.mp4 \
    --quantization fp8-cast \
    --num-inference-steps 50 \
    --seed 42

```

This command activates the fp8-cast policy for the `TI2VidOneStagePipeline`, causing the model loader to convert transformer linear weights to FP8 format before inference begins.

### Required Arguments

The fp8-cast policy requires access to the raw checkpoint to fold any pre-quantization scales into the FP8 weight tensors. Therefore, `--checkpoint-path` is mandatory when using `--quantization fp8-cast`. The pipeline scripts (`ti2vid_one_stage`, `ti2vid_two_stages`, `hdr_ic_lora`, etc.) share the same argument parser via `default_1_stage_arg_parser` and `default_2_stage_arg_parser`, making the flag available across all video generation workflows.

## How the FP8-Cast Policy Works Internally

### CLI Parsing and Resolution

In [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py), the parser defines the `--quantization` flag to enumerate supported policies including `fp8-cast` and `fp8-scaled-mm`. After argument parsing, the `_resolve_quantization` helper converts the string value into a `QuantizationPolicy` object by calling `QuantizationKind.to_policy`.

```python

# Conceptual flow from args.py

parser.add_argument("--quantization", choices=["fp8-cast", "fp8-scaled-mm"])

# Later resolved via:

policy = QuantizationKind.to_policy(args.quantization)

```

### Policy Construction

The `QuantizationKind.FP8_CAST` enum maps to `ltx_core.quantization.fp8_cast.build_policy` in [`packages/ltx-pipelines/src/ltx_pipelines/utils/quantization_factory.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/quantization_factory.py). This factory creates a down-cast map (`TRANSFORMER_LINEAR_DOWNCAST_MAP`) that replaces original bfloat16 weights with FP8 tensors, and registers `ModuleOps` (`UPCAST_DURING_INFERENCE`) that rewrite the forward pass of each linear layer.

### Runtime Execution

When the `DiffusionStage` receives the `QuantizationPolicy`, it applies the policy's `module_ops` and `sd_ops` via the LTX-Core loader. During inference, `Fp8CastLinear.forward` in [`packages/ltx-core/src/ltx_core/quantization/fp8_cast.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/quantization/fp8_cast.py) up-casts the FP8 weight to the input dtype (typically bfloat16) and optionally performs stochastic rounding via Triton kernels.

## Advanced Configuration

### Enabling Stochastic Rounding

For higher fidelity during the up-cast operation, enable stochastic rounding by setting the Triton environment variable. The `build_policy` function automatically detects this setting and adjusts the forward behavior accordingly.

```bash
export TORCH_TRITON_ENABLED=1

ltx-pipelines ti2vid_one_stage \
    --checkpoint-path /path/to/LTX-2.safetensors \
    --quantization fp8-cast \
    --prompt "A cinematic scene" \
    --output-path output.mp4

```

No additional CLI flags are required; the policy constructor handles the stochastic rounding configuration internally.

### Combining with torch.compile

You can stack fp8-cast quantization with PyTorch compilation for additional speed gains. The quantization policy operates before compilation, ensuring the converted model graphs are optimized together.

```bash
ltx-pipelines ti2vid_one_stage \
    --checkpoint-path /path/to/LTX-2.safetensors \
    --quantization fp8-cast \
    --compile mode=reduce-overhead fullgraph=true \
    --prompt "Abstract digital art" \
    --output-path output.mp4

```

## Summary

- **FP8-cast quantization** in LTX-2 stores transformer weights in `float8_e4m3fn` and up-casts to bfloat16 during inference, reducing memory footprint.
- Activate the feature via `--quantization fp8-cast` in any pipeline script using the shared argument parser from [`args.py`](https://github.com/Lightricks/LTX-2/blob/main/args.py).
- The policy requires a valid `--checkpoint-path` to fold pre-quantization scales into the weight tensors during initialization.
- Internally, `QuantizationKind.FP8_CAST` triggers `build_policy` in [`quantization_factory.py`](https://github.com/Lightricks/LTX-2/blob/main/quantization_factory.py), which configures `Fp8CastLinear` layers to handle runtime up-casting.
- Enable **stochastic rounding** by setting `TORCH_TRITON_ENABLED=1` for improved numerical precision during the cast operation.

## Frequently Asked Questions

### What is the difference between fp8-cast and fp8-scaled-mm in LTX-2?

**Fp8-cast** stores weights in native FP8 format and up-casts them to bfloat16 on-the-fly during the forward pass, making it compatible with standard inference hardware. **Fp8-scaled-mm** uses scaled matrix multiplication operations that require specialized FP8 compute support. The fp8-cast method is generally preferred for broader hardware compatibility and simpler deployment.

### Do I need a specific checkpoint format for FP8-cast quantization?

No special checkpoint format is required, but you must provide the `--checkpoint-path` argument when using fp8-cast. The `build_policy` function in `ltx_core.quantization.fp8_cast` reads the standard checkpoint and folds any pre-quantization scales into the FP8 weight tensors during the loading process.

### Can I use FP8-cast quantization with all LTX-2 pipeline scripts?

Yes. The `--quantization` flag is defined in the shared argument parsers (`default_1_stage_arg_parser` and `default_2_stage_arg_parser`) used by all entry points including `ti2vid_one_stage`, `ti2vid_two_stages`, and `hdr_ic_lora`. This ensures consistent quantization support across text-to-video, image-to-video, and other generation workflows.

### How does stochastic rounding affect FP8-cast inference quality?

Stochastic rounding, enabled via `TORCH_TRITON_ENABLED=1`, adds random noise during the FP8-to-bfloat16 conversion to reduce quantization bias. According to the implementation in [`fp8_cast.py`](https://github.com/Lightricks/LTX-2/blob/main/fp8_cast.py), this Triton-based operation improves numerical fidelity when up-casting weights, particularly for sensitive visual details in video generation, with minimal impact on inference latency.