How to Switch Between TransformerEngine and FourOverSix NVFP4 Backends in LongLive

Set model_quant_use_transformer_engine to true for TransformerEngine or false for FourOverSix in your configuration file, or use the --backend flag when running the checkpoint conversion script.

LongLive supports two distinct NVFP4 (4-bit Floating Point) inference backends for efficient quantization: TransformerEngine and FourOverSix. Switching between these backends requires modifying a single boolean configuration flag that determines how quantized weights are loaded and processed. This article explains the exact mechanisms used in the NVlabs/LongLive repository to control backend selection.

Understanding the Two NVFP4 Backends

LongLive implements two different strategies for NVFP4 quantization, each requiring specific checkpoint formats.

TransformerEngine Backend

The TransformerEngine backend operates on BF16 generator checkpoints (model_te.pt). At runtime, the model remains in BF16 precision while a lightweight wrapper (TransformerEngineLinear in utils/quant.py lines 190-210) routes linear layers through the TransformerEngine quantizer. This approach quantizes weights dynamically during inference rather than storing them pre-quantized.

FourOverSix Backend

The FourOverSix backend requires pre-materialized NVFP4 state dictionaries (model_4o6.pt) where weights are already quantized to NVFP4 format. When loading these checkpoints, the system uses quantize_model_for_fouroversix_nvfp4 (defined in utils/quant.py lines 401-437) to instantiate layers that contain pre-quantized parameters. No runtime quantizer is necessary since the quantization occurred during the checkpoint creation phase.

The Configuration Flag That Controls Everything

Selection between these backends is entirely driven by the model_quant_use_transformer_engine boolean flag. This single configuration value acts as the routing mechanism throughout the codebase:

  • When set to true, the system expects BF16 checkpoints and applies TransformerEngine wrappers
  • When set to false, the system expects pre-quantized FourOverSix checkpoints and loads them directly

All configuration files, CLI tools, and programmatic APIs ultimately modify this flag to dictate backend behavior.

Switching Backends via YAML Configuration

Edit your inference configuration file to explicitly select the desired backend. The example configuration at configs/nvfp4/inference_nvfp4.yaml (line 47) contains this flag:

model_quant: true                               # Enable NVFP4 quantization

model_quant_use_transformer_engine: true        # Use TransformerEngine backend

# Set to false for FourOverSix backend

According to the repository's README.md (lines 113-117), you must ensure the checkpoint type matches this flag: use model_te.pt when the flag is true and model_4o6.pt when the flag is false.

Switching Backends via Command Line

The helper script scripts/save_merged_nvfp4_generator.py provides a --backend argument that automatically sets the configuration flag. Lines 174-176 in this script map the CLI argument to the boolean value:


# Create a FourOverSix checkpoint (pre-quantized)

python scripts/save_merged_nvfp4_generator.py \
    --config_path configs/nvfp4/inference_nvfp4.yaml \
    --backend fouroversix \
    --output_path checkpoints/model_4o6.pt

# Create a TransformerEngine checkpoint (BF16 with runtime wrapping)

python scripts/save_merged_nvfp4_generator.py \
    --config_path configs/nvfp4/inference_nvfp4.yaml \
    --backend transformer_engine \
    --output_path checkpoints/model_te.pt

The --backend parameter accepts transformer_engine or fouroversix, setting config.model_quant_use_transformer_engine accordingly before processing.

Switching Backends Programmatically

You can toggle backends dynamically in Python by modifying the configuration object before initializing the pipeline:

from omegaconf import OmegaConf
from utils.inference_utils import setup_nvfp4_pipeline
from pipeline import CausalDiffusionInferencePipeline
import torch

# Load and modify configuration

cfg = OmegaConf.load("configs/nvfp4/inference_nvfp4.yaml")
cfg.model_quant_use_transformer_engine = False  # Set True for TransformerEngine

# Initialize pipeline with selected backend

device = torch.device("cuda")
pipe = CausalDiffusionInferencePipeline(cfg, device=device)
setup_nvfp4_pipeline(pipe, cfg, device)

The setup_nvfp4_pipeline function (in utils/inference_utils.py lines 84-92) automatically routes to the appropriate loading mechanism based on this flag.

How the Code Routes Between Backends

The setup_nvfp4_pipeline function in utils/inference_utils.py serves as the central dispatcher for backend selection. When model_quant_use_transformer_engine is true, it loads the BF16 checkpoint and wraps the model with TransformerEngineLinear modules. When false, it detects pre-quantized state dictionaries using is_nvfp4_state_dict and loads them directly into FourOverSix-quantized modules.

The implementation includes validation checks (lines 103-110 in utils/inference_utils.py) that raise ValueError if you attempt to load a mismatched checkpoint type. For example, loading a FourOverSix checkpoint while model_quant_use_transformer_engine is true triggers an explicit error message to prevent configuration mismatches.

Checkpoint Compatibility and Validation

Each backend requires a specific checkpoint format:

  • TransformerEngine: Expects model_te.pt containing BF16 weights
  • FourOverSix: Expects model_4o6.pt containing pre-quantized NVFP4 weights

The function is_te_nvfp4_checkpoint in utils/nvfp4_checkpoint.py (lines 31-33) helps detect TE-style checkpoints automatically. However, the model_quant_use_transformer_engine flag remains the authoritative source for determining routing behavior, ensuring the correct quantization path is followed regardless of filename conventions.

Summary

  • Single flag control: The model_quant_use_transformer_engine boolean determines whether LongLive uses TransformerEngine or FourOverSix backends.
  • TransformerEngine (true): Loads BF16 checkpoints (model_te.pt) and applies runtime quantization via TransformerEngineLinear wrappers.
  • FourOverSix (false): Loads pre-quantized NVFP4 checkpoints (model_4o6.pt) directly without runtime quantization overhead.
  • CLI convenience: Use scripts/save_merged_nvfp4_generator.py with --backend transformer_engine or --backend fouroversix to automatically configure the flag.
  • Safety checks: The code validates checkpoint types against the flag setting in setup_nvfp4_pipeline to prevent runtime errors.

Frequently Asked Questions

What is the difference between TransformerEngine and FourOverSix backends?

TransformerEngine performs quantization at runtime by wrapping BF16 linear layers with quantization modules, offering flexibility but requiring the TransformerEngine library during inference. FourOverSix uses pre-quantized weights stored directly in the checkpoint, eliminating runtime quantization overhead and reducing memory bandwidth requirements, but requires preprocessing the model weights before deployment.

Can I convert a TransformerEngine checkpoint to FourOverSix format?

Yes. Use the conversion script scripts/save_merged_nvfp4_generator.py with --backend fouroversix to materialize a pre-quantized NVFP4 checkpoint from a standard model. This script quantizes the weights offline and saves them in the model_4o6.pt format compatible with the FourOverSix backend.

What happens if I set the wrong backend flag for my checkpoint?

The code raises a ValueError with a descriptive error message. Specifically, setup_nvfp4_pipeline in utils/inference_utils.py checks for mismatches between the model_quant_use_transformer_engine setting and the actual checkpoint format, preventing silent failures or incorrect quantization behavior.

Which backend offers better inference performance?

FourOverSix typically provides lower latency and higher throughput because weights are already quantized to NVFP4 format, eliminating the overhead of runtime quantization and reducing memory bandwidth. TransformerEngine offers more flexibility for experimentation with different quantization strategies but incurs minor overhead during the forward pass for on-the-fly quantization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →