How to Enable AVX2-Only CPU Backend for KTransformers Inference

Set the CPUINFER_CPU_INSTRUCT environment variable to AVX2 before building or running KTransformers to force the AVX2-only CPU backend, which supports BF16, FP8, GPTQ_INT4, and RAWINT4 quantization on Intel Haswell+ or AMD Zen+ processors without requiring AVX512.

KTransformers (kvcache-ai/ktransformers) provides optimized CPU inference for large language models through its kt-kernel C++ layer. For machines lacking AVX512 or AMX instructions, the framework offers a dedicated AVX2-only backend that automatically handles modern quantization formats while maximizing compatibility with older hardware.

Verify CPU Capability

Before building, confirm your processor supports the required AVX2 instructions using the provided diagnostic script.

Check CPU Features with check_cpu_features.py

Run the helper script located at kt-kernel/scripts/check_cpu_features.py to detect available instruction sets:

python kt-kernel/scripts/check_cpu_features.py

If your CPU supports AVX2 but lacks AVX512/AMX, the script outputs "AVX2 Support (fallback)", indicating you should proceed with the AVX2-only configuration.

Configure the Build Environment

The AVX2 backend is controlled at compile time through environment variables that modify the CMake configuration in kt-kernel/setup.py (lines 139 and 556).

Set CPUINFER_CPU_INSTRUCT=AVX2

Export the instruction-set flag before invoking any build commands:

export CPUINFER_CPU_INSTRUCT=AVX2

When this variable is set, the build system in setup.py disables AVX512/AMX flags and selects AVX2-specific compilation options (e.g., -DLLAMA_AVX2=ON), ensuring the kernel compiles only AVX2 code paths.

Build the Kernel with AVX2 Flags

With the environment variable exported, compile the kernel using the provided install script or pip:

cd kt-kernel
bash install.sh

# Alternative: pip install -e .

The install.sh wrapper respects the CPUINFER_CPU_INSTRUCT value and passes the appropriate flags to CMake, producing binaries that target the AVX2 instruction set exclusively.

Runtime Backend Selection

Once compiled, the Python utilities in kt-kernel/python/utils/amx.py expose the AVX2-specific backend objects: AVX2BF16_MOE, AVX2FP8_MOE, AVX2GPTQInt4_MOE, and AVX2RawInt4_MOE. These classes are automatically selected at runtime when higher-level instruction sets are unavailable.

Run inference using any supported quantization method:

python -m ktransformers.infer \
    --model facebook/opt-13b \
    --kt-method BF16

Valid --kt-method values for AVX2 include BF16, FP8, GPTQ_INT4, and RAWINT4, all of which route through the optimized AVX2 implementations.

Force AVX2 on AVX512-Capable Systems

To compile the AVX2-only backend even on machines that support AVX512 or AMX, maintain the environment variable setting:

export CPUINFER_CPU_INSTRUCT=AVX2
bash install.sh  # Re-compile to overwrite previous builds

This configuration ignores richer instruction sets and ensures the kernel uses only AVX2 code paths, which can be useful for testing compatibility or deploying to heterogeneous hardware environments.

Summary

Frequently Asked Questions

What CPUs support the KTransformers AVX2 backend?

Any processor supporting the AVX2 instruction set works, including Intel Haswell (4th Gen Core, 2013+) and AMD Zen-based chips. Run kt-kernel/scripts/check_cpu_features.py to verify your specific CPU flags before building.

Can I run quantized models with the AVX2-only backend?

Yes. The AVX2 backend supports GPTQ_INT4, RAWINT4, FP8, and BF16 quantization formats. These map to the AVX2GPTQInt4_MOE, AVX2RawInt4_MOE, AVX2FP8_MOE, and AVX2BF16_MOE runtime classes respectively.

How do I verify that AVX2 is actually being used during inference?

Check that CPUINFER_CPU_INSTRUCT was set to AVX2 during the build phase. The build system in setup.py prints the selected instruction set during compilation. At runtime, the presence of AVX2-specific backend objects in kt-kernel/python/utils/amx.py confirms the compiled kernel is active.

Do I need to rebuild if I want to switch between AVX2 and AVX512?

Yes. The instruction set selection occurs at compile time via CMake flags in setup.py. To switch backends, set or unset CPUINFER_CPU_INSTRUCT=AVX2 and re-run bash install.sh to generate new binaries targeting the desired instruction set.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →