# How to Switch Between TransformerEngine and FourOverSix NVFP4 Backends in LongLive

> Easily switch between TransformerEngine and FourOverSix NVFP4 backends in LongLive via configuration or command line. Optimize your model performance now.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: how-to-guide
- Published: 2026-05-24

---

**Set `model_quant_use_transformer_engine` to `true` for TransformerEngine or `false` for FourOverSix in your configuration file, or use the `--backend` flag when running the checkpoint conversion script.**

LongLive supports two distinct NVFP4 (4-bit Floating Point) inference backends for efficient quantization: TransformerEngine and FourOverSix. Switching between these backends requires modifying a single boolean configuration flag that determines how quantized weights are loaded and processed. This article explains the exact mechanisms used in the NVlabs/LongLive repository to control backend selection.

## Understanding the Two NVFP4 Backends

LongLive implements two different strategies for NVFP4 quantization, each requiring specific checkpoint formats.

### TransformerEngine Backend

The **TransformerEngine** backend operates on BF16 generator checkpoints (`model_te.pt`). At runtime, the model remains in BF16 precision while a lightweight wrapper (`TransformerEngineLinear` in [`utils/quant.py`](https://github.com/NVlabs/LongLive/blob/main/utils/quant.py) lines 190-210) routes linear layers through the TransformerEngine quantizer. This approach quantizes weights dynamically during inference rather than storing them pre-quantized.

### FourOverSix Backend

The **FourOverSix** backend requires pre-materialized NVFP4 state dictionaries (`model_4o6.pt`) where weights are already quantized to NVFP4 format. When loading these checkpoints, the system uses `quantize_model_for_fouroversix_nvfp4` (defined in [`utils/quant.py`](https://github.com/NVlabs/LongLive/blob/main/utils/quant.py) lines 401-437) to instantiate layers that contain pre-quantized parameters. No runtime quantizer is necessary since the quantization occurred during the checkpoint creation phase.

## The Configuration Flag That Controls Everything

Selection between these backends is entirely driven by the **`model_quant_use_transformer_engine`** boolean flag. This single configuration value acts as the routing mechanism throughout the codebase:

- When set to `true`, the system expects BF16 checkpoints and applies TransformerEngine wrappers
- When set to `false`, the system expects pre-quantized FourOverSix checkpoints and loads them directly

All configuration files, CLI tools, and programmatic APIs ultimately modify this flag to dictate backend behavior.

## Switching Backends via YAML Configuration

Edit your inference configuration file to explicitly select the desired backend. The example configuration at [`configs/nvfp4/inference_nvfp4.yaml`](https://github.com/NVlabs/LongLive/blob/main/configs/nvfp4/inference_nvfp4.yaml) (line 47) contains this flag:

```yaml
model_quant: true                               # Enable NVFP4 quantization

model_quant_use_transformer_engine: true        # Use TransformerEngine backend

# Set to false for FourOverSix backend

```

According to the repository's [`README.md`](https://github.com/NVlabs/LongLive/blob/main/README.md) (lines 113-117), you must ensure the checkpoint type matches this flag: use `model_te.pt` when the flag is `true` and `model_4o6.pt` when the flag is `false`.

## Switching Backends via Command Line

The helper script [`scripts/save_merged_nvfp4_generator.py`](https://github.com/NVlabs/LongLive/blob/main/scripts/save_merged_nvfp4_generator.py) provides a `--backend` argument that automatically sets the configuration flag. Lines 174-176 in this script map the CLI argument to the boolean value:

```bash

# Create a FourOverSix checkpoint (pre-quantized)

python scripts/save_merged_nvfp4_generator.py \
    --config_path configs/nvfp4/inference_nvfp4.yaml \
    --backend fouroversix \
    --output_path checkpoints/model_4o6.pt

# Create a TransformerEngine checkpoint (BF16 with runtime wrapping)

python scripts/save_merged_nvfp4_generator.py \
    --config_path configs/nvfp4/inference_nvfp4.yaml \
    --backend transformer_engine \
    --output_path checkpoints/model_te.pt

```

The `--backend` parameter accepts `transformer_engine` or `fouroversix`, setting `config.model_quant_use_transformer_engine` accordingly before processing.

## Switching Backends Programmatically

You can toggle backends dynamically in Python by modifying the configuration object before initializing the pipeline:

```python
from omegaconf import OmegaConf
from utils.inference_utils import setup_nvfp4_pipeline
from pipeline import CausalDiffusionInferencePipeline
import torch

# Load and modify configuration

cfg = OmegaConf.load("configs/nvfp4/inference_nvfp4.yaml")
cfg.model_quant_use_transformer_engine = False  # Set True for TransformerEngine

# Initialize pipeline with selected backend

device = torch.device("cuda")
pipe = CausalDiffusionInferencePipeline(cfg, device=device)
setup_nvfp4_pipeline(pipe, cfg, device)

```

The `setup_nvfp4_pipeline` function (in [`utils/inference_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/inference_utils.py) lines 84-92) automatically routes to the appropriate loading mechanism based on this flag.

## How the Code Routes Between Backends

The `setup_nvfp4_pipeline` function in [`utils/inference_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/inference_utils.py) serves as the central dispatcher for backend selection. When `model_quant_use_transformer_engine` is `true`, it loads the BF16 checkpoint and wraps the model with `TransformerEngineLinear` modules. When `false`, it detects pre-quantized state dictionaries using `is_nvfp4_state_dict` and loads them directly into FourOverSix-quantized modules.

The implementation includes validation checks (lines 103-110 in [`utils/inference_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/inference_utils.py)) that raise `ValueError` if you attempt to load a mismatched checkpoint type. For example, loading a FourOverSix checkpoint while `model_quant_use_transformer_engine` is `true` triggers an explicit error message to prevent configuration mismatches.

## Checkpoint Compatibility and Validation

Each backend requires a specific checkpoint format:

- **TransformerEngine**: Expects `model_te.pt` containing BF16 weights
- **FourOverSix**: Expects `model_4o6.pt` containing pre-quantized NVFP4 weights

The function `is_te_nvfp4_checkpoint` in [`utils/nvfp4_checkpoint.py`](https://github.com/NVlabs/LongLive/blob/main/utils/nvfp4_checkpoint.py) (lines 31-33) helps detect TE-style checkpoints automatically. However, the `model_quant_use_transformer_engine` flag remains the authoritative source for determining routing behavior, ensuring the correct quantization path is followed regardless of filename conventions.

## Summary

- **Single flag control**: The `model_quant_use_transformer_engine` boolean determines whether LongLive uses TransformerEngine or FourOverSix backends.
- **TransformerEngine** (`true`): Loads BF16 checkpoints (`model_te.pt`) and applies runtime quantization via `TransformerEngineLinear` wrappers.
- **FourOverSix** (`false`): Loads pre-quantized NVFP4 checkpoints (`model_4o6.pt`) directly without runtime quantization overhead.
- **CLI convenience**: Use [`scripts/save_merged_nvfp4_generator.py`](https://github.com/NVlabs/LongLive/blob/main/scripts/save_merged_nvfp4_generator.py) with `--backend transformer_engine` or `--backend fouroversix` to automatically configure the flag.
- **Safety checks**: The code validates checkpoint types against the flag setting in `setup_nvfp4_pipeline` to prevent runtime errors.

## Frequently Asked Questions

### What is the difference between TransformerEngine and FourOverSix backends?

**TransformerEngine** performs quantization at runtime by wrapping BF16 linear layers with quantization modules, offering flexibility but requiring the TransformerEngine library during inference. **FourOverSix** uses pre-quantized weights stored directly in the checkpoint, eliminating runtime quantization overhead and reducing memory bandwidth requirements, but requires preprocessing the model weights before deployment.

### Can I convert a TransformerEngine checkpoint to FourOverSix format?

Yes. Use the conversion script [`scripts/save_merged_nvfp4_generator.py`](https://github.com/NVlabs/LongLive/blob/main/scripts/save_merged_nvfp4_generator.py) with `--backend fouroversix` to materialize a pre-quantized NVFP4 checkpoint from a standard model. This script quantizes the weights offline and saves them in the `model_4o6.pt` format compatible with the FourOverSix backend.

### What happens if I set the wrong backend flag for my checkpoint?

The code raises a `ValueError` with a descriptive error message. Specifically, `setup_nvfp4_pipeline` in [`utils/inference_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/inference_utils.py) checks for mismatches between the `model_quant_use_transformer_engine` setting and the actual checkpoint format, preventing silent failures or incorrect quantization behavior.

### Which backend offers better inference performance?

**FourOverSix** typically provides lower latency and higher throughput because weights are already quantized to NVFP4 format, eliminating the overhead of runtime quantization and reducing memory bandwidth. **TransformerEngine** offers more flexibility for experimentation with different quantization strategies but incurs minor overhead during the forward pass for on-the-fly quantization.