# How KTransformers Compares to vLLM and Other Inference Frameworks: CPU-First MoE vs GPU-Centric Design

> Compare KTransformers CPU-first MoE inference with vLLM's GPU-centric design. Discover KTransformers' specialized MoE scheduling and high-performance C++ kernels for efficient AI model execution.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: comparison
- Published: 2026-07-26

---

**KTransformers is a CPU-first, heterogeneous inference stack that couples high-performance C++ kernels (AVX2/AVX512/AMX) with optional GPU offload for hot experts, providing specialized Mixture-of-Experts (MoE) scheduling that vLLM lacks.**

If you are evaluating inference frameworks for large language models, understanding how KTransformers compares to vLLM and alternatives is critical for selecting the right hardware and software stack. The `kvcache-ai/ktransformers` repository delivers a fundamentally different architecture optimized for commodity CPU servers rather than GPU clusters, with specific advantages for MoE workloads.

## Architecture and Execution Model

The primary distinction lies in where computation happens. **KTransformers** optimizes for CPU-optimized MoE kernels that reside in the `kt-kernel` library, dynamically selected at runtime based on host CPU capabilities (AVX2, AVX512, or AMX). In contrast, **vLLM** is GPU-centric, relying on CUDA kernels and FlashAttention for dense models with limited native MoE support.

Other frameworks like HuggingFace Transformers or DeepSpeed typically require external plugins for MoE scheduling and lack specialized CPU paths for sparse expert computation. KTransformers treats CPU execution as a first-class citizen, making it viable for edge deployments or on-premise servers without GPU accelerators.

## Hardware Detection and Kernel Selection

KTransformers automates hardware capability detection through the [`install.sh`](https://github.com/kvcache-ai/ktransformers/blob/main/install.sh) script, which examines CPU instruction sets and installs pre-built many-linux wheels without requiring compilation. As documented in [`kt-kernel/README.md`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md), the system loads the optimal kernel variant automatically, though users can override detection via environment variables.

vLLM relies on CUDA device queries and offers no CPU fallback mechanism, while other frameworks generally require manual device selection and library management. This makes KTransformers uniquely suited for **AVX2-only machines** such as older Xeon or AMD Zen+ processors where GPU inference is unavailable or cost-prohibitive.

## MoE Expert Scheduling Strategies

Where KTransformers truly differentiates itself is in fine-grained expert placement. The framework implements a **GPU-CPU expert mask** with four placement strategies—`uniform`, `frequency`, `front-loading`, and `random`—detailed in [`doc/en/kt-kernel/experts-sched-Tutorial.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/experts-sched-Tutorial.md). A dynamic expert-update mechanism can reshuffle experts during long prefills, optimizing throughput based on activation statistics.

vLLM's roadmap mentions future multi-token MoE scheduling but currently lacks a mature expert-mask system. Other frameworks either omit MoE scheduling entirely or defer to external plugins like DeepSpeed MoE. This capability allows KTransformers to achieve up to **100 tokens/s** on an 80B MoE model using a 30% GPU-expert ratio on commodity hardware (4×RTX 4090 + 2×Xeon Gold 6454S).

## Quantization Backend Support

KTransformers supports multiple weight formats through dedicated CPU paths: **FP8, BF16, INT4/INT8 (AMX), and GGUF (Llamafile)**. The [`scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/scripts/convert_cpu_weights.py) utility quantizes GPU-side weights into CPU-optimized formats for AMX backends.

vLLM primarily supports FP16/BF16 on GPU with CPU quantization as a non-focus area, while other frameworks rely on GPU-oriented quantization libraries like bitsandbytes. For CPU inference, KTransformers provides production-ready quantization without requiring CUDA toolkits.

## Integration and Deployment Patterns

The framework offers native **SGLang integration** via the `sglang-kt` plugin, enabling heterogeneous inference where GPU handles "hot" experts and CPU manages "cold" experts. The `kt` CLI abstracts this complexity with commands like `kt run`, `kt chat`, and `kt doctor` for hardware compatibility checks.

```bash

# Launch SGLang server with heterogeneous MoE

python -m sglang.launch_server \
    --model /mnt/models/Qwen3-30B-A3B \
    --kt-method AMXINT8 \
    --kt-weight-path /mnt/models/Qwen3-30B-A3B-INT8 \
    --kt-cpuinfer 64 \
    --kt-threadpool-count 2 \
    --kt-num-gpu-experts 32 \
    --kt-expert-placement-strategy frequency \
    --init-expert-location /mnt/stats/activation_stats.pt \
    --kt-enable-dynamic-expert-update \
    --kt-gpu-prefill-token-threshold 512

```

vLLM offers no direct SGLang integration, and other frameworks typically require custom wrappers to achieve similar heterogeneous execution.

## Installation Footprint and Requirements

KTransformers distributes pre-built wheels for Python 3.10-3.12 with no CUDA toolkit required for CPU paths. GPU support uses a **static CUDA runtime** bundled in the wheel. This contrasts with vLLM, which requires compatible CUDA toolkits and PyTorch/FlashAttention builds, and other frameworks that often demand compiled extensions.

```python

# CPU-only inference with INT8 AMX backend

from kt_kernel import KTMoEWrapper

wrapper = KTMoEWrapper(
    layer_idx=0,
    num_experts=8,
    num_experts_per_tok=2,
    hidden_size=4096,
    moe_intermediate_size=14336,
    num_gpu_experts=0,          # CPU-only mode

    cpuinfer_threads=32,
    threadpool_count=2,
    weight_path="/path/to/int8_weights",
    chunked_prefill_size=512,
    method="AMXINT8",           # AMX-optimized INT8 kernels

)

wrapper.load_weights(physical_to_logical_map)
output = wrapper.forward(hidden_states, topk_ids, topk_weights, cuda_stream=None)

```

## Summary

- **KTransformers** excels at heterogeneous CPU+GPU MoE inference with specialized AVX2/AVX512/AMX kernels, while **vLLM** optimizes for GPU-only dense model throughput.
- The framework provides automatic CPU instruction set detection via [`install.sh`](https://github.com/kvcache-ai/ktransformers/blob/main/install.sh) and pre-built wheels, eliminating compilation requirements present in other stacks.
- Four expert placement strategies (`uniform`, `frequency`, `front-loading`, `random`) with dynamic updates during prefills offer scheduling flexibility unavailable in vLLM or HuggingFace Transformers.
- Native support for INT4/INT8 AMX and GGUF formats through [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py) enables efficient CPU quantization without GPU dependencies.
- SGLang integration and the `kt` CLI provide production-ready deployment tools for heterogeneous inference scenarios.

## Frequently Asked Questions

### Can KTransformers run without any GPU?

Yes. KTransformers supports pure CPU inference using optimized kernels for AVX2, AVX512, and AMX instruction sets. Set `num_gpu_experts=0` in the `KTMoEWrapper` configuration or omit GPU flags in the SGLang launcher. The framework automatically detects available CPU capabilities through [`install.sh`](https://github.com/kvcache-ai/ktransformers/blob/main/install.sh) and loads the appropriate kernel variant from `kt-kernel`.

### How does expert scheduling in KTransformers improve performance over standard MoE inference?

KTransformers implements a GPU-CPU expert mask that places frequently activated experts on GPU while keeping others on CPU, reducing memory bandwidth bottlenecks. The `frequency` placement strategy uses activation statistics to optimize placement, while dynamic expert updates during long prefills allow real-time optimization. This heterogeneous approach can achieve 100 tokens/s on 80B MoE models with only 30% GPU expert coverage, outperforming GPU-only configurations when expert placement is optimized according to the benchmarks in [`doc/en/kt-kernel/experts-sched-Tutorial.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/experts-sched-Tutorial.md).

### What quantization formats does KTransformers support compared to vLLM?

KTransformers supports FP8, BF16, INT4/INT8 (optimized for Intel AMX), and GGUF (Llamafile) formats with dedicated CPU inference paths. The [`scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/scripts/convert_cpu_weights.py) utility converts standard weights to these formats. vLLM primarily focuses on FP16/BF16 GPU inference with limited CPU quantization support, making KTransformers the better choice for compressed CPU inference on commodity hardware.

### Is KTransformers compatible with existing model serving infrastructure?

Yes. KTransformers provides a native SGLang plugin (`sglang-kt`) that integrates with existing serving stacks, and the `kt` CLI offers commands for server deployment (`kt run`), interactive chat (`kt chat`), and system diagnostics (`kt doctor`). While vLLM offers its own serving infrastructure, KTransformers' SGLang integration enables heterogeneous GPU+CPU serving within compatible ecosystem tools.