How KTransformers Compares to vLLM and Other Inference Frameworks: CPU-First MoE vs GPU-Centric Design

KTransformers is a CPU-first, heterogeneous inference stack that couples high-performance C++ kernels (AVX2/AVX512/AMX) with optional GPU offload for hot experts, providing specialized Mixture-of-Experts (MoE) scheduling that vLLM lacks.

If you are evaluating inference frameworks for large language models, understanding how KTransformers compares to vLLM and alternatives is critical for selecting the right hardware and software stack. The kvcache-ai/ktransformers repository delivers a fundamentally different architecture optimized for commodity CPU servers rather than GPU clusters, with specific advantages for MoE workloads.

Architecture and Execution Model

The primary distinction lies in where computation happens. KTransformers optimizes for CPU-optimized MoE kernels that reside in the kt-kernel library, dynamically selected at runtime based on host CPU capabilities (AVX2, AVX512, or AMX). In contrast, vLLM is GPU-centric, relying on CUDA kernels and FlashAttention for dense models with limited native MoE support.

Other frameworks like HuggingFace Transformers or DeepSpeed typically require external plugins for MoE scheduling and lack specialized CPU paths for sparse expert computation. KTransformers treats CPU execution as a first-class citizen, making it viable for edge deployments or on-premise servers without GPU accelerators.

Hardware Detection and Kernel Selection

KTransformers automates hardware capability detection through the install.sh script, which examines CPU instruction sets and installs pre-built many-linux wheels without requiring compilation. As documented in kt-kernel/README.md, the system loads the optimal kernel variant automatically, though users can override detection via environment variables.

vLLM relies on CUDA device queries and offers no CPU fallback mechanism, while other frameworks generally require manual device selection and library management. This makes KTransformers uniquely suited for AVX2-only machines such as older Xeon or AMD Zen+ processors where GPU inference is unavailable or cost-prohibitive.

MoE Expert Scheduling Strategies

Where KTransformers truly differentiates itself is in fine-grained expert placement. The framework implements a GPU-CPU expert mask with four placement strategies—uniform, frequency, front-loading, and random—detailed in doc/en/kt-kernel/experts-sched-Tutorial.md. A dynamic expert-update mechanism can reshuffle experts during long prefills, optimizing throughput based on activation statistics.

vLLM's roadmap mentions future multi-token MoE scheduling but currently lacks a mature expert-mask system. Other frameworks either omit MoE scheduling entirely or defer to external plugins like DeepSpeed MoE. This capability allows KTransformers to achieve up to 100 tokens/s on an 80B MoE model using a 30% GPU-expert ratio on commodity hardware (4×RTX 4090 + 2×Xeon Gold 6454S).

Quantization Backend Support

KTransformers supports multiple weight formats through dedicated CPU paths: FP8, BF16, INT4/INT8 (AMX), and GGUF (Llamafile). The scripts/convert_cpu_weights.py utility quantizes GPU-side weights into CPU-optimized formats for AMX backends.

vLLM primarily supports FP16/BF16 on GPU with CPU quantization as a non-focus area, while other frameworks rely on GPU-oriented quantization libraries like bitsandbytes. For CPU inference, KTransformers provides production-ready quantization without requiring CUDA toolkits.

Integration and Deployment Patterns

The framework offers native SGLang integration via the sglang-kt plugin, enabling heterogeneous inference where GPU handles "hot" experts and CPU manages "cold" experts. The kt CLI abstracts this complexity with commands like kt run, kt chat, and kt doctor for hardware compatibility checks.


# Launch SGLang server with heterogeneous MoE

python -m sglang.launch_server \
    --model /mnt/models/Qwen3-30B-A3B \
    --kt-method AMXINT8 \
    --kt-weight-path /mnt/models/Qwen3-30B-A3B-INT8 \
    --kt-cpuinfer 64 \
    --kt-threadpool-count 2 \
    --kt-num-gpu-experts 32 \
    --kt-expert-placement-strategy frequency \
    --init-expert-location /mnt/stats/activation_stats.pt \
    --kt-enable-dynamic-expert-update \
    --kt-gpu-prefill-token-threshold 512

vLLM offers no direct SGLang integration, and other frameworks typically require custom wrappers to achieve similar heterogeneous execution.

Installation Footprint and Requirements

KTransformers distributes pre-built wheels for Python 3.10-3.12 with no CUDA toolkit required for CPU paths. GPU support uses a static CUDA runtime bundled in the wheel. This contrasts with vLLM, which requires compatible CUDA toolkits and PyTorch/FlashAttention builds, and other frameworks that often demand compiled extensions.


# CPU-only inference with INT8 AMX backend

from kt_kernel import KTMoEWrapper

wrapper = KTMoEWrapper(
    layer_idx=0,
    num_experts=8,
    num_experts_per_tok=2,
    hidden_size=4096,
    moe_intermediate_size=14336,
    num_gpu_experts=0,          # CPU-only mode

    cpuinfer_threads=32,
    threadpool_count=2,
    weight_path="/path/to/int8_weights",
    chunked_prefill_size=512,
    method="AMXINT8",           # AMX-optimized INT8 kernels

)

wrapper.load_weights(physical_to_logical_map)
output = wrapper.forward(hidden_states, topk_ids, topk_weights, cuda_stream=None)

Summary

  • KTransformers excels at heterogeneous CPU+GPU MoE inference with specialized AVX2/AVX512/AMX kernels, while vLLM optimizes for GPU-only dense model throughput.
  • The framework provides automatic CPU instruction set detection via install.sh and pre-built wheels, eliminating compilation requirements present in other stacks.
  • Four expert placement strategies (uniform, frequency, front-loading, random) with dynamic updates during prefills offer scheduling flexibility unavailable in vLLM or HuggingFace Transformers.
  • Native support for INT4/INT8 AMX and GGUF formats through convert_cpu_weights.py enables efficient CPU quantization without GPU dependencies.
  • SGLang integration and the kt CLI provide production-ready deployment tools for heterogeneous inference scenarios.

Frequently Asked Questions

Can KTransformers run without any GPU?

Yes. KTransformers supports pure CPU inference using optimized kernels for AVX2, AVX512, and AMX instruction sets. Set num_gpu_experts=0 in the KTMoEWrapper configuration or omit GPU flags in the SGLang launcher. The framework automatically detects available CPU capabilities through install.sh and loads the appropriate kernel variant from kt-kernel.

How does expert scheduling in KTransformers improve performance over standard MoE inference?

KTransformers implements a GPU-CPU expert mask that places frequently activated experts on GPU while keeping others on CPU, reducing memory bandwidth bottlenecks. The frequency placement strategy uses activation statistics to optimize placement, while dynamic expert updates during long prefills allow real-time optimization. This heterogeneous approach can achieve 100 tokens/s on 80B MoE models with only 30% GPU expert coverage, outperforming GPU-only configurations when expert placement is optimized according to the benchmarks in doc/en/kt-kernel/experts-sched-Tutorial.md.

What quantization formats does KTransformers support compared to vLLM?

KTransformers supports FP8, BF16, INT4/INT8 (optimized for Intel AMX), and GGUF (Llamafile) formats with dedicated CPU inference paths. The scripts/convert_cpu_weights.py utility converts standard weights to these formats. vLLM primarily focuses on FP16/BF16 GPU inference with limited CPU quantization support, making KTransformers the better choice for compressed CPU inference on commodity hardware.

Is KTransformers compatible with existing model serving infrastructure?

Yes. KTransformers provides a native SGLang plugin (sglang-kt) that integrates with existing serving stacks, and the kt CLI offers commands for server deployment (kt run), interactive chat (kt chat), and system diagnostics (kt doctor). While vLLM offers its own serving infrastructure, KTransformers' SGLang integration enables heterogeneous GPU+CPU serving within compatible ecosystem tools.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →