How to Configure Intel AMX Acceleration (AMX-INT8/AMX-BF16) in KTransformers
To configure Intel AMX acceleration in KTransformers, verify your CPU supports the amx-tile, amx-int8, and amx-bf16 flags, ensure the kernel was compiled with AMX support (auto-detected during ./install.sh), and set the backend field to "AMXINT8" or "AMXBF16" in your injection configuration YAML file.
KTransformers is an open-source framework designed to accelerate large language model inference on CPUs using advanced instruction sets. Configuring Intel AMX acceleration in KTransformers unlocks near-GPU throughput on Sapphire Rapids and newer Xeon processors by offloading quantized Mixture-of-Experts (MoE) layers to the Intel Advanced Matrix Extensions (AMX) tile-based engine.
Prerequisites: Verify CPU AMX Support
Before configuration, confirm your hardware supports the required instructions. The system must report the AMX flags to enable the optimized kernels.
Run the following command to check for AMX capabilities:
lscpu | grep -i amx
The output must contain amx-tile, amx-int8, and amx-bf16. According to the documentation in doc/en/AMX.md (lines 37-49), KTransformers relies on these CPUID flags to trigger the automatic selection of AMX-optimized code paths at runtime.
Build Configuration: Compile with AMX Support
The KTransformers kernel must be compiled with AMX instructions enabled. The build system auto-detects CPU capabilities during the installation process.
When building from source, the install.sh script checks for AMX support and automatically adds the -DKTRANSFORMERS_CPU_USE_AMX CMake flag. As documented in kt-kernel/README.md (lines 74-78), this ensures the compiler generates AMX-tiled instructions for the MoE kernels. If you need to verify or manually control this during a custom build, inspect kt-kernel/setup.py (lines 18-23), where the CPUINFER_ENABLE_AMX environment variable configures the underlying CMake build flags.
Runtime Configuration: Select the AMX Backend
KTransformers uses an injection configuration YAML file to map specific model layers to optimized CPU operators. To enable AMX acceleration, you must specify the backend explicitly in this configuration.
Set the backend field within the kwargs of the expert injection rule to either "AMXINT8" for 8-bit integer quantization or "AMXBF16" for bfloat16 quantization. This configuration is documented in doc/en/AMX.md (lines 55-71) and implemented via the KTransformersExperts operator.
Example configuration from kt-kernel/optimize_rules/Qwen3MoE-serve-amx.yaml:
- match:
name: "^model\\.layers\\..*\\.mlp\\.experts$"
replace:
class: ktransformers.operators.experts.KTransformersExperts
kwargs:
prefill_device: "cuda"
prefill_op: "KExpertsTorch"
generate_device: "cpu"
generate_op: "KExpertsCPU"
out_device: "cuda"
backend: "AMXINT8" # Use "AMXBF16" for BF16 quantization
The generate_device: "cpu" setting is required, as AMX acceleration is implemented only in the CPU inference path. When the server starts, KTransformers loads the AMX kernels and quantizes weights on-the-fly (or uses pre-converted weights) to run the MoE layers with AMX tiles, as detailed in doc/en/AMX.md (lines 74-82).
Optional: Convert Weights for AMX Optimization
For maximum performance, convert your model weights to an AMX-friendly tiled format before inference. This step is optional if you are using pre-quantized GGUF models, but required for raw PyTorch/SafeTensor checkpoints.
Use the convert_cpu_weights.py script located in kt-kernel/scripts/ to rewrite FP16/BF16 weights into INT4 or INT8 SafeTensor files optimized for AMX tiling:
python kt-kernel/scripts/convert_cpu_weights.py \
--input path/to/model.safetensors \
--output path/to/amx_weights \
--method AMXINT8
This process, described in kt-kernel/scripts/README.md (lines 5-9), organizes the weight matrices into tiles that align with the AMX 2D register layout, minimizing data movement during matrix multiplication.
Environment Variable Control
KTransformers provides runtime toggles to override AMX behavior via environment variables. Set CPUINFER_ENABLE_AMX to control the backend selection:
export CPUINFER_ENABLE_AMX=ON # Default: enables AMX if CPU supports it
export CPUINFER_ENABLE_AMX=OFF # Forces fallback to AVX512 implementation
As defined in kt-kernel/setup.py (lines 18-23), the default value is ON when the CPU reports AMX support. Disabling this variable is useful for benchmarking against AVX512 baselines or troubleshooting compatibility issues.
Implementation Details: Python API and Kernel Selection
Under the hood, the kt-kernel/python/utils/amx.py module handles the dynamic loading of AMX kernels. It implements the AMXMoEWrapper class, which detects the compiled AMX kernel variants and selects the appropriate backend based on your YAML configuration and CPU capabilities.
You can also programmatically inject the AMX backend using the Python API:
from ktransformers import KTransformer
model = KTransformer.from_pretrained("Qwen3-Moe")
model.inject_expert_backend(
layer_regex=r"^model\.layers\..*\.mlp\.experts$",
backend="AMXBF16" # or "AMXINT8"
)
model.generate("Hello, world!")
This approach bypasses YAML configuration for specific layers, useful for debugging or A/B testing quantization methods.
Summary
- Verify hardware: Ensure
lscpureportsamx-tile,amx-int8, andamx-bf16flags before starting. - Build correctly: Use
./install.shto auto-detect and compile with-DKTRANSFORMERS_CPU_USE_AMX. - Configure YAML: Set
backend: "AMXINT8"or"AMXBF16"in the injection configuration for MoE expert layers. - Prepare weights: Run
convert_cpu_weights.pyfor SafeTensor models; skip for GGUF models. - Control runtime: Use
CPUINFER_ENABLE_AMX=ON/OFFto toggle between AMX and AVX512 kernels.
Frequently Asked Questions
What is the difference between AMXINT8 and AMXBF16 backends?
AMXINT8 uses 8-bit integer quantization, offering the highest throughput and smallest memory footprint with minimal accuracy loss on most models. AMXBF16 uses bfloat16 format, providing higher numerical precision than INT8 while still leveraging AMX acceleration, ideal for models sensitive to quantization errors. Both utilize the amx-tile instructions but operate on different data types as implemented in the AMXMoEWrapper class.
Do I need to convert weights if using GGUF models?
No. If you supply a pre-quantized GGUF model, KTransformers handles the weight layout internally and the conversion step using convert_cpu_weights.py is not required. The conversion script is only necessary for FP16/BF16 SafeTensor or PyTorch checkpoint formats that need to be tiled specifically for the AMX kernel layout.
How can I verify that AMX acceleration is actually active?
Check your server logs for kernel selection messages, or monitor the CPU flags during execution. You can also explicitly disable AMX by setting CPUINFER_ENABLE_AMX=OFF and comparing token-per-second throughput—if the performance drops significantly, AMX was active. The kt-kernel/python/utils/amx.py wrapper logs the selected backend during model loading when verbose logging is enabled.
What happens if I configure AMX on a CPU that does not support it?
KTransformers performs runtime CPU capability detection. If the amx-tile flag is not detected in CPUID, the framework automatically falls back to AVX512 or AVX2 implementations, even if the YAML configuration specifies "AMXINT8" or "AMXBF16". However, attempting to force execution via unsupported compiler flags during build time will result in illegal instruction errors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →