How to Enable FP8 Per-Channel Precision and Native BF16 Inference in KTransformers
Enable FP8 per-channel precision and native BF16 inference by compiling KTransformers with AVX-512-VBMI support, converting model weights to FP8 format with per-channel scaling, and configuring MoeConfig.quant_config.per_channel = True to utilize the AMXFP8PerChannel_MOE kernel.
KTransformers is an open-source inference engine by kvcache-ai that accelerates Mixture-of-Experts (MoE) models through advanced CPU optimizations. Activating FP8 per-channel precision reduces memory bandwidth while maintaining accuracy via individual scaling factors for each output channel, and when combined with native BF16 compute paths, delivers maximum throughput on modern AMX-enabled hardware.
Hardware Prerequisites for FP8 Per-Channel Inference
FP8 per-channel precision requires specific CPU instruction sets to function. According to the source code in kt-kernel/python/utils/amx.py, the AMXFP8PerChannel_MOE kernel is only available on processors supporting AVX-512-VBMI (Vector Byte Manipulation Instructions) and ideally AMX-Tile with AMX-BF16 support.
Verify your CPU capabilities using the provided diagnostic script:
python kt-kernel/scripts/check_cpu_features.py
If your hardware supports the required extensions, the script outputs: "✅ Your CPU supports full AVX512 (F/BF16/VNNI/VBMI) – FP8 MoE will work!" Otherwise, the script identifies missing flags and suggests recompilation targets.
Compiling KTransformers with AVX-512-VBMI Support
The native FP8 backend must be compiled with specific flags to expose the per-channel kernel. In kt-kernel/setup.py, enable the following options before building:
CPUINFER_ENABLE_AVX512_VBMI=ON
CPUINFER_ENABLE_AVX512_BF16=ON # Required for AMX acceleration
Execute the build:
cd kt-kernel
python setup.py build
# Or use the provided install script
Successful compilation exposes AMXFP8PerChannel_MOE in the kt_kernel_ext.moe module, as defined in kt-kernel/python/utils/amx.py. This kernel implements the BFI6 (native FP8) backend that processes FP8 weights without dequantization overhead while utilizing BF16 compute instructions where applicable.
Converting Model Weights to FP8 Per-Channel Format
Standard FP16 or BF16 checkpoints must be converted to the FP8 per-channel format. The conversion utility located at kt-kernel/scripts/convert_cpu_weights.py handles quantization and scale generation:
python kt-kernel/scripts/convert_cpu_weights.py \
--input-type bf16 \
--input-bf16-hf-path ./original_bf16_checkpoint \
--output-fp8-hf-path ./fp8_perchannel_checkpoint \
--quant-config '{"method":"FP8","per_channel":true}'
This script produces two critical tensor types:
- weight: Stored as FP8 (E4M3) format for expert weights
- weight_scale: Individual
float32scaling factors for each output channel
The FP8SafeTensorLoader in kt-kernel/python/utils/loader.py automatically detects the per-channel format by identifying the weight_scale tensor (as opposed to weight_scale_inv used for block-wise quantization), ensuring proper kernel selection at load time.
Configuring MoeConfig for Native FP8 Inference
Instantiate the MoE layer with per-channel FP8 precision by configuring the quantization settings in Python:
from ktransformers import MoeConfig
import kt_kernel_ext.moe as kt_moe
# 1. Configure quantization
cfg = MoeConfig()
cfg.quant_config.method = "FP8" # Select FP8 backend
cfg.quant_config.per_channel = True # Enable per-channel scales
cfg.quant_config.group_size = 0 # Unused for per-channel mode
# 2. Initialize the native FP8 MoE backend
moe = kt_moe.AMXFP8PerChannel_MOE(cfg)
# 3. Run inference
import torch
x = torch.randn(1, cfg.hidden_size, device="cuda")
output = moe(x)
When per_channel is set to True, the runtime dispatches to AMXFP8PerChannel_MOE rather than fallback AVX2 implementations. This path maintains separate scales per output channel in float32 precision, preserving numeric fidelity while benefiting from 8-bit weight compression.
Runtime Verification and Backend Selection
Confirm the native backend is active by inspecting the compiled extension capabilities:
from kt_kernel_ext import moe as kt_moe
available_backends = [k for k in kt_moe.__dict__.keys() if 'AMX' in k]
print(f"Available FP8 backends: {available_backends}")
if hasattr(kt_moe, 'AMXFP8PerChannel_MOE'):
print("✅ Native FP8 per-channel backend (BFI6) is available")
else:
print("⚠️ Falling back to AVX2 FP8 (block-wise only)")
The _HAS_FP8_PERCHANNEL_SUPPORT flag in kt-kernel/python/utils/amx.py indicates whether the compiled binary includes per-channel support. If this backend is missing despite proper compilation flags, verify that your CPU reports AVX-512-VBMI in /proc/cpuinfo or via the check_cpu_features.py script.
Performance Characteristics
The FP8 per-channel implementation delivers optimal performance through:
- Memory bandwidth reduction: 2x compression compared to BF16/FP16 weights
- Per-channel accuracy: Individual scales prevent quantization error accumulation across channels
- Native execution: The BFI6 backend avoids FP32 dequantization overhead, operating directly on compressed tensors using AMX-BF16 instructions for compute-heavy operations
For a complete working example, refer to kt-kernel/examples/test_fp8_perchannel_moe.py, which demonstrates end-to-end configuration and forward pass execution with per-channel FP8 tensors.
Summary
- Compile KTransformers with
CPUINFER_ENABLE_AVX512_VBMI=ONandCPUINFER_ENABLE_AVX512_BF16=ONinkt-kernel/setup.pyto expose theAMXFP8PerChannel_MOEkernel. - Convert existing checkpoints using
kt-kernel/scripts/convert_cpu_weights.pywith"per_channel":trueto generate FP8 weights and per-channelfloat32scales. - Configure
MoeConfig.quant_config.per_channel = Trueto enable the native FP8 backend with per-channel precision. - Verify CPU capabilities with
kt-kernel/scripts/check_cpu_features.pybefore deployment to ensure AMX/AVX-512-VBMI support. - Runtime automatically selects
AMXFP8PerChannel_MOEwhen available, falling back to AVX2 block-wise FP8 if hardware requirements are not met.
Frequently Asked Questions
What CPU features are required for FP8 per-channel inference in KTransformers?
FP8 per-channel precision requires AVX-512-VBMI (Vector Byte Manipulation Instructions) at minimum. For optimal performance with the native FP8 (BFI6) backend, AMX-Tile and AMX-BF16 instructions are also recommended. You can verify support by running python kt-kernel/scripts/check_cpu_features.py from the repository root.
How does per-channel FP8 differ from block-wise FP8 quantization?
Per-channel FP8 stores a separate float32 scale factor for every output channel in the weight_scale tensor, providing higher numeric fidelity than block-wise quantization which shares scales across groups of weights. According to kt-kernel/python/utils/loader.py, the loader detects per-channel format by the presence of weight_scale (as opposed to weight_scale_inv), and routes to AMXFP8PerChannel_MOE instead of the block-wise AMXFP8_MOE kernel.
Can I use BF16 compute with FP8 weights in KTransformers?
Yes. When FP8 per-channel mode is enabled and the AMXFP8PerChannel_MOE kernel is active, the backend utilizes AMX-BF16 instructions for the computational forward pass while keeping weights compressed in FP8 (E4M3) format. This hybrid approach, sometimes referred to as BFI6 in the documentation, maximizes memory efficiency without sacrificing compute precision.
What happens if my CPU doesn't support AVX-512-VBMI?
If the CPU lacks AVX-512-VBMI, the kt_kernel_ext.moe module will not expose AMXFP8PerChannel_MOE, and KTransformers will fall back to AVX2-based FP8 implementations that support only block-wise quantization. You can still run FP8 inference, but you must use block-wise scales (weight_scale_inv) rather than per-channel scales, and you will not benefit from the native BF16 compute acceleration provided by the AMX backend.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →