# How to Enable FP8 Per-Channel Precision and Native BF16 Inference in KTransformers

> Learn to boost KTransformers performance with FP8 per-channel precision and native BF16 inference. Optimize your models today!

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Enable FP8 per-channel precision and native BF16 inference by compiling KTransformers with AVX-512-VBMI support, converting model weights to FP8 format with per-channel scaling, and configuring `MoeConfig.quant_config.per_channel = True` to utilize the `AMXFP8PerChannel_MOE` kernel.**

KTransformers is an open-source inference engine by kvcache-ai that accelerates Mixture-of-Experts (MoE) models through advanced CPU optimizations. Activating FP8 per-channel precision reduces memory bandwidth while maintaining accuracy via individual scaling factors for each output channel, and when combined with native BF16 compute paths, delivers maximum throughput on modern AMX-enabled hardware.

## Hardware Prerequisites for FP8 Per-Channel Inference

FP8 per-channel precision requires specific CPU instruction sets to function. According to the source code in [`kt-kernel/python/utils/amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/amx.py), the `AMXFP8PerChannel_MOE` kernel is only available on processors supporting **AVX-512-VBMI** (Vector Byte Manipulation Instructions) and ideally **AMX-Tile** with **AMX-BF16** support.

Verify your CPU capabilities using the provided diagnostic script:

```bash
python kt-kernel/scripts/check_cpu_features.py

```

If your hardware supports the required extensions, the script outputs: "✅ Your CPU supports full AVX512 (F/BF16/VNNI/VBMI) – FP8 MoE will work!" Otherwise, the script identifies missing flags and suggests recompilation targets.

## Compiling KTransformers with AVX-512-VBMI Support

The native FP8 backend must be compiled with specific flags to expose the per-channel kernel. In [`kt-kernel/setup.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/setup.py), enable the following options before building:

```bash
CPUINFER_ENABLE_AVX512_VBMI=ON
CPUINFER_ENABLE_AVX512_BF16=ON  # Required for AMX acceleration

```

Execute the build:

```bash
cd kt-kernel
python setup.py build

# Or use the provided install script

```

Successful compilation exposes `AMXFP8PerChannel_MOE` in the `kt_kernel_ext.moe` module, as defined in [`kt-kernel/python/utils/amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/amx.py). This kernel implements the BFI6 (native FP8) backend that processes FP8 weights without dequantization overhead while utilizing BF16 compute instructions where applicable.

## Converting Model Weights to FP8 Per-Channel Format

Standard FP16 or BF16 checkpoints must be converted to the FP8 per-channel format. The conversion utility located at [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py) handles quantization and scale generation:

```bash
python kt-kernel/scripts/convert_cpu_weights.py \
  --input-type bf16 \
  --input-bf16-hf-path ./original_bf16_checkpoint \
  --output-fp8-hf-path ./fp8_perchannel_checkpoint \
  --quant-config '{"method":"FP8","per_channel":true}'

```

This script produces two critical tensor types:
- **weight**: Stored as FP8 (E4M3) format for expert weights
- **weight_scale**: Individual `float32` scaling factors for each output channel

The `FP8SafeTensorLoader` in [`kt-kernel/python/utils/loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/loader.py) automatically detects the per-channel format by identifying the `weight_scale` tensor (as opposed to `weight_scale_inv` used for block-wise quantization), ensuring proper kernel selection at load time.

## Configuring MoeConfig for Native FP8 Inference

Instantiate the MoE layer with per-channel FP8 precision by configuring the quantization settings in Python:

```python
from ktransformers import MoeConfig
import kt_kernel_ext.moe as kt_moe

# 1. Configure quantization

cfg = MoeConfig()
cfg.quant_config.method = "FP8"          # Select FP8 backend

cfg.quant_config.per_channel = True     # Enable per-channel scales

cfg.quant_config.group_size = 0         # Unused for per-channel mode

# 2. Initialize the native FP8 MoE backend

moe = kt_moe.AMXFP8PerChannel_MOE(cfg)

# 3. Run inference

import torch
x = torch.randn(1, cfg.hidden_size, device="cuda")
output = moe(x)

```

When `per_channel` is set to `True`, the runtime dispatches to `AMXFP8PerChannel_MOE` rather than fallback AVX2 implementations. This path maintains separate scales per output channel in `float32` precision, preserving numeric fidelity while benefiting from 8-bit weight compression.

## Runtime Verification and Backend Selection

Confirm the native backend is active by inspecting the compiled extension capabilities:

```python
from kt_kernel_ext import moe as kt_moe

available_backends = [k for k in kt_moe.__dict__.keys() if 'AMX' in k]
print(f"Available FP8 backends: {available_backends}")

if hasattr(kt_moe, 'AMXFP8PerChannel_MOE'):
    print("✅ Native FP8 per-channel backend (BFI6) is available")
else:
    print("⚠️  Falling back to AVX2 FP8 (block-wise only)")

```

The `_HAS_FP8_PERCHANNEL_SUPPORT` flag in [`kt-kernel/python/utils/amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/amx.py) indicates whether the compiled binary includes per-channel support. If this backend is missing despite proper compilation flags, verify that your CPU reports AVX-512-VBMI in `/proc/cpuinfo` or via the [`check_cpu_features.py`](https://github.com/kvcache-ai/ktransformers/blob/main/check_cpu_features.py) script.

## Performance Characteristics

The FP8 per-channel implementation delivers optimal performance through:
- **Memory bandwidth reduction**: 2x compression compared to BF16/FP16 weights
- **Per-channel accuracy**: Individual scales prevent quantization error accumulation across channels
- **Native execution**: The BFI6 backend avoids FP32 dequantization overhead, operating directly on compressed tensors using AMX-BF16 instructions for compute-heavy operations

For a complete working example, refer to [`kt-kernel/examples/test_fp8_perchannel_moe.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/examples/test_fp8_perchannel_moe.py), which demonstrates end-to-end configuration and forward pass execution with per-channel FP8 tensors.

## Summary

- **Compile** KTransformers with `CPUINFER_ENABLE_AVX512_VBMI=ON` and `CPUINFER_ENABLE_AVX512_BF16=ON` in [`kt-kernel/setup.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/setup.py) to expose the `AMXFP8PerChannel_MOE` kernel.
- **Convert** existing checkpoints using [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py) with `"per_channel":true` to generate FP8 weights and per-channel `float32` scales.
- **Configure** `MoeConfig.quant_config.per_channel = True` to enable the native FP8 backend with per-channel precision.
- **Verify** CPU capabilities with [`kt-kernel/scripts/check_cpu_features.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/check_cpu_features.py) before deployment to ensure AMX/AVX-512-VBMI support.
- **Runtime** automatically selects `AMXFP8PerChannel_MOE` when available, falling back to AVX2 block-wise FP8 if hardware requirements are not met.

## Frequently Asked Questions

### What CPU features are required for FP8 per-channel inference in KTransformers?

FP8 per-channel precision requires **AVX-512-VBMI** (Vector Byte Manipulation Instructions) at minimum. For optimal performance with the native FP8 (BFI6) backend, **AMX-Tile** and **AMX-BF16** instructions are also recommended. You can verify support by running `python kt-kernel/scripts/check_cpu_features.py` from the repository root.

### How does per-channel FP8 differ from block-wise FP8 quantization?

Per-channel FP8 stores a separate `float32` scale factor for every output channel in the `weight_scale` tensor, providing higher numeric fidelity than block-wise quantization which shares scales across groups of weights. According to [`kt-kernel/python/utils/loader.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/loader.py), the loader detects per-channel format by the presence of `weight_scale` (as opposed to `weight_scale_inv`), and routes to `AMXFP8PerChannel_MOE` instead of the block-wise `AMXFP8_MOE` kernel.

### Can I use BF16 compute with FP8 weights in KTransformers?

Yes. When FP8 per-channel mode is enabled and the `AMXFP8PerChannel_MOE` kernel is active, the backend utilizes **AMX-BF16** instructions for the computational forward pass while keeping weights compressed in FP8 (E4M3) format. This hybrid approach, sometimes referred to as BFI6 in the documentation, maximizes memory efficiency without sacrificing compute precision.

### What happens if my CPU doesn't support AVX-512-VBMI?

If the CPU lacks AVX-512-VBMI, the `kt_kernel_ext.moe` module will not expose `AMXFP8PerChannel_MOE`, and KTransformers will fall back to AVX2-based FP8 implementations that support only block-wise quantization. You can still run FP8 inference, but you must use block-wise scales (`weight_scale_inv`) rather than per-channel scales, and you will not benefit from the native BF16 compute acceleration provided by the AMX backend.