# How to Configure Intel AMX Acceleration (AMX-INT8/AMX-BF16) in KTransformers

> Learn how to configure Intel AMX acceleration like AMX-INT8 and AMX-BF16 in KTransformers. Unlock faster AI model performance with this quick guide.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-26

---

**To configure Intel AMX acceleration in KTransformers, verify your CPU supports the `amx-tile`, `amx-int8`, and `amx-bf16` flags, ensure the kernel was compiled with AMX support (auto-detected during [`./install.sh`](https://github.com/kvcache-ai/ktransformers/blob/main/./install.sh)), and set the `backend` field to `"AMXINT8"` or `"AMXBF16"` in your injection configuration YAML file.**

KTransformers is an open-source framework designed to accelerate large language model inference on CPUs using advanced instruction sets. Configuring Intel AMX acceleration in KTransformers unlocks near-GPU throughput on Sapphire Rapids and newer Xeon processors by offloading quantized Mixture-of-Experts (MoE) layers to the **Intel Advanced Matrix Extensions (AMX)** tile-based engine.

## Prerequisites: Verify CPU AMX Support

Before configuration, confirm your hardware supports the required instructions. The system must report the AMX flags to enable the optimized kernels.

Run the following command to check for AMX capabilities:

```bash
lscpu | grep -i amx

```

The output must contain `amx-tile`, `amx-int8`, and `amx-bf16`. According to the documentation in [`doc/en/AMX.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/AMX.md) (lines 37-49), KTransformers relies on these CPUID flags to trigger the automatic selection of AMX-optimized code paths at runtime.

## Build Configuration: Compile with AMX Support

The KTransformers kernel must be compiled with AMX instructions enabled. The build system auto-detects CPU capabilities during the installation process.

When building from source, the [`install.sh`](https://github.com/kvcache-ai/ktransformers/blob/main/install.sh) script checks for AMX support and automatically adds the `-DKTRANSFORMERS_CPU_USE_AMX` CMake flag. As documented in [`kt-kernel/README.md`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md) (lines 74-78), this ensures the compiler generates AMX-tiled instructions for the MoE kernels. If you need to verify or manually control this during a custom build, inspect [`kt-kernel/setup.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/setup.py) (lines 18-23), where the `CPUINFER_ENABLE_AMX` environment variable configures the underlying CMake build flags.

## Runtime Configuration: Select the AMX Backend

KTransformers uses an **injection configuration** YAML file to map specific model layers to optimized CPU operators. To enable AMX acceleration, you must specify the backend explicitly in this configuration.

Set the `backend` field within the `kwargs` of the expert injection rule to either `"AMXINT8"` for 8-bit integer quantization or `"AMXBF16"` for bfloat16 quantization. This configuration is documented in [`doc/en/AMX.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/AMX.md) (lines 55-71) and implemented via the `KTransformersExperts` operator.

Example configuration from [`kt-kernel/optimize_rules/Qwen3MoE-serve-amx.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/optimize_rules/Qwen3MoE-serve-amx.yaml):

```yaml
- match:
    name: "^model\\.layers\\..*\\.mlp\\.experts$"
  replace:
    class: ktransformers.operators.experts.KTransformersExperts
    kwargs:
      prefill_device: "cuda"
      prefill_op: "KExpertsTorch"
      generate_device: "cpu"
      generate_op: "KExpertsCPU"
      out_device: "cuda"
      backend: "AMXINT8"  # Use "AMXBF16" for BF16 quantization

```

The `generate_device: "cpu"` setting is required, as AMX acceleration is implemented only in the CPU inference path. When the server starts, KTransformers loads the AMX kernels and quantizes weights on-the-fly (or uses pre-converted weights) to run the MoE layers with AMX tiles, as detailed in [`doc/en/AMX.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/AMX.md) (lines 74-82).

## Optional: Convert Weights for AMX Optimization

For maximum performance, convert your model weights to an AMX-friendly tiled format before inference. This step is optional if you are using pre-quantized GGUF models, but required for raw PyTorch/SafeTensor checkpoints.

Use the [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py) script located in `kt-kernel/scripts/` to rewrite FP16/BF16 weights into INT4 or INT8 SafeTensor files optimized for AMX tiling:

```bash
python kt-kernel/scripts/convert_cpu_weights.py \
  --input path/to/model.safetensors \
  --output path/to/amx_weights \
  --method AMXINT8

```

This process, described in [`kt-kernel/scripts/README.md`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/README.md) (lines 5-9), organizes the weight matrices into tiles that align with the AMX 2D register layout, minimizing data movement during matrix multiplication.

## Environment Variable Control

KTransformers provides runtime toggles to override AMX behavior via environment variables. Set `CPUINFER_ENABLE_AMX` to control the backend selection:

```bash
export CPUINFER_ENABLE_AMX=ON  # Default: enables AMX if CPU supports it

export CPUINFER_ENABLE_AMX=OFF # Forces fallback to AVX512 implementation

```

As defined in [`kt-kernel/setup.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/setup.py) (lines 18-23), the default value is `ON` when the CPU reports AMX support. Disabling this variable is useful for benchmarking against AVX512 baselines or troubleshooting compatibility issues.

## Implementation Details: Python API and Kernel Selection

Under the hood, the [`kt-kernel/python/utils/amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/amx.py) module handles the dynamic loading of AMX kernels. It implements the `AMXMoEWrapper` class, which detects the compiled AMX kernel variants and selects the appropriate backend based on your YAML configuration and CPU capabilities.

You can also programmatically inject the AMX backend using the Python API:

```python
from ktransformers import KTransformer

model = KTransformer.from_pretrained("Qwen3-Moe")
model.inject_expert_backend(
    layer_regex=r"^model\.layers\..*\.mlp\.experts$",
    backend="AMXBF16"  # or "AMXINT8"

)
model.generate("Hello, world!")

```

This approach bypasses YAML configuration for specific layers, useful for debugging or A/B testing quantization methods.

## Summary

- **Verify hardware**: Ensure `lscpu` reports `amx-tile`, `amx-int8`, and `amx-bf16` flags before starting.
- **Build correctly**: Use [`./install.sh`](https://github.com/kvcache-ai/ktransformers/blob/main/./install.sh) to auto-detect and compile with `-DKTRANSFORMERS_CPU_USE_AMX`.
- **Configure YAML**: Set `backend: "AMXINT8"` or `"AMXBF16"` in the injection configuration for MoE expert layers.
- **Prepare weights**: Run [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py) for SafeTensor models; skip for GGUF models.
- **Control runtime**: Use `CPUINFER_ENABLE_AMX=ON/OFF` to toggle between AMX and AVX512 kernels.

## Frequently Asked Questions

### What is the difference between AMXINT8 and AMXBF16 backends?

**AMXINT8** uses 8-bit integer quantization, offering the highest throughput and smallest memory footprint with minimal accuracy loss on most models. **AMXBF16** uses bfloat16 format, providing higher numerical precision than INT8 while still leveraging AMX acceleration, ideal for models sensitive to quantization errors. Both utilize the `amx-tile` instructions but operate on different data types as implemented in the `AMXMoEWrapper` class.

### Do I need to convert weights if using GGUF models?

No. If you supply a pre-quantized GGUF model, KTransformers handles the weight layout internally and the conversion step using [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py) is not required. The conversion script is only necessary for FP16/BF16 SafeTensor or PyTorch checkpoint formats that need to be tiled specifically for the AMX kernel layout.

### How can I verify that AMX acceleration is actually active?

Check your server logs for kernel selection messages, or monitor the CPU flags during execution. You can also explicitly disable AMX by setting `CPUINFER_ENABLE_AMX=OFF` and comparing token-per-second throughput—if the performance drops significantly, AMX was active. The [`kt-kernel/python/utils/amx.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/utils/amx.py) wrapper logs the selected backend during model loading when verbose logging is enabled.

### What happens if I configure AMX on a CPU that does not support it?

KTransformers performs runtime CPU capability detection. If the `amx-tile` flag is not detected in CPUID, the framework automatically falls back to AVX512 or AVX2 implementations, even if the YAML configuration specifies `"AMXINT8"` or `"AMXBF16"`. However, attempting to force execution via unsupported compiler flags during build time will result in illegal instruction errors.