# KT-Kernel CPU-GPU Heterogeneous Computing Architecture Explained

> Explore the KT-Kernel architecture for CPU-GPU heterogeneous computing. Learn how it optimizes inference by running MoE layers across processors for peak performance. Discover the ktransformers innovation at kvcache-ai.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: architecture
- Published: 2026-07-26

---

**KT-Kernel is a hybrid inference engine that executes Mixture-of-Experts layers across CPU and GPU by partitioning cold experts onto CPU-optimized SIMD kernels while reserving hot experts for GPU acceleration.**

The kvcache-ai/ktransformers repository implements KT-Kernel to enable high-throughput inference for sparse large language models. This architecture automatically partitions computational workloads between host CPUs and discrete GPUs, leveraging AVX-512/AMX instructions for CPU-side execution while embedding a static CUDA runtime to eliminate external toolkit dependencies.

## High-Level Architecture Components

KT-Kernel separates the **CPU-only MoE backend** from the **GPU-accelerated hot-expert backend**, orchestrating both through a NUMA-aware thread-pool and lightweight task-queue.

| Component | Role | Implementation |
|---|---|---|
| **Automatic CPU variant detection** | Selects optimal kernel variant (AMX, AVX512+BF16, AVX512+VNNI, AVX512 Base, AVX2) at import time | `kt_kernel.__cpu_variant__` runtime detection |
| **CPU-optimized MoE kernels** | Native-precision (FP8/BF16/INT4) and quantized (INT8) kernels for MoE experts | Operators in `kt-kernel/operators/` (e.g., [`softmax.hpp`](https://github.com/kvcache-ai/ktransformers/blob/main/softmax.hpp), [`rope.hpp`](https://github.com/kvcache-ai/ktransformers/blob/main/rope.hpp), [`tp.hpp`](https://github.com/kvcache-ai/ktransformers/blob/main/tp.hpp)) |
| **NUMA-aware thread-pool** | Creates one pool per NUMA node, pins threads to local sockets, allocates per-socket shared buffers | [`cpu_backend/worker_pool.h`](https://github.com/kvcache-ai/ktransformers/blob/main/cpu_backend/worker_pool.h) & [`worker_pool.cpp`](https://github.com/kvcache-ai/ktransformers/blob/main/worker_pool.cpp) |
| **Task-queue & shared-memory** | Lock-free queue pushing inference tasks to appropriate worker pools; pre-allocated reusable buffers | [`cpu_backend/task_queue.h`](https://github.com/kvcache-ai/ktransformers/blob/main/cpu_backend/task_queue.h) & [`task_queue.cpp`](https://github.com/kvcache-ai/ktransformers/blob/main/task_queue.cpp) |
| **Static CUDA runtime** | Ships compiled CUDA runtime inside the wheel, requiring no external CUDA toolkit | `kt_kernel_ext` (C++ CUDA kernels) |
| **Python API** | Thin wrapper exposing C++ MoE kernels to Python and SGLang | [`kt_kernel/__init__.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt_kernel/__init__.py) exposing `KTMoEWrapper` |

## Heterogeneous MoE Forward Pass Data Flow

The execution pipeline enables pipeline parallelism and dynamic expert placement through six distinct stages:

1. **Model loading** — GPU weights are loaded by SGLang while CPU-only expert weights are quantized (INT4/INT8) and stored under `--kt-weight-path`.

2. **Kernel selection** — On import, `kt_kernel` detects host CPU capabilities and sets `__cpu_variant__` to the optimal backend (e.g., `AMX`, `AVX512+BF16`).

3. **Task creation** — The Python `KTMoEWrapper` builds a forward-task containing hidden states, top-k expert ids, and weights, then pushes it onto the `TaskQueue`.

4. **Thread-pool dispatch** — The NUMA-local worker thread picks the task and maps expert ids to either the CPU backend (via compiled SIMD kernels) or the GPU backend.

5. **GPU execution** — For experts marked *hot* (`--kt-num-gpu-experts`), the wrapper calls `CPUInfer.submit_with_cuda_stream` to stream data to the GPU kernel; otherwise, the CPU kernel runs directly.

6. **Result collection** — The worker writes output into a pre-allocated shared buffer defined in [`cpu_backend/shared_mem_buffer.h`](https://github.com/kvcache-ai/ktransformers/blob/main/cpu_backend/shared_mem_buffer.h), allowing the Python side to block (`sync_forward`) or continue asynchronously.

## Key Source Files and Implementation

KT-Kernel's implementation spans C++ SIMD kernels, CUDA operations, and Python bindings:

- **`kt-kernel/operators/*.hpp`** — Contains SIMD-optimized kernels for MoE operations including softmax, rotary-positional embeddings, and tensor-parallel logic.

- **[`kt-kernel/cpu_backend/worker_pool.h`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/cpu_backend/worker_pool.h)** — Implements the NUMA-aware thread-pool that pins workers to specific sockets and manages per-socket buffer allocation.

- **[`kt-kernel/cpu_backend/task_queue.h`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/cpu_backend/task_queue.h)** — Defines the lock-free queue dispatching MoE forward tasks to worker pools with minimal latency.

- **[`kt-kernel/cpu_backend/cpuinfer.h`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/cpu_backend/cpuinfer.h)** — The C++ entry point exposing the CPU MoE inference API to Python via `pybind11`, referenced from Python as `kt_kernel_ext.CPUInfer`.

- **[`kt-kernel/python/kt_kernel/__init__.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/kt_kernel/__init__.py)** — Python package entry defining `KTMoEWrapper` and re-exporting version constants.

- **[`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py)** — Helper script quantizing GPU-side weights into AMX-friendly INT4/INT8 formats required by the CPU backend.

## Usage Examples

### Direct Python API

Create a heterogeneous MoE wrapper and execute forward passes:

```python
from kt_kernel import KTMoEWrapper

# Initialize wrapper for heterogeneous execution

wrapper = KTMoEWrapper(
    layer_idx=0,
    num_experts=8,
    num_experts_per_tok=2,
    hidden_size=4096,
    moe_intermediate_size=14336,
    num_gpu_experts=2,            # hot experts on GPU

    cpuinfer_threads=32,          # physical cores

    threadpool_count=2,           # NUMA nodes

    weight_path="/path/to/cpu-weights",
    method="AMXINT8",             # Intel AMX backend

)

wrapper.load_weights()

# Execute heterogeneous forward pass

output = wrapper.forward(
    hidden_states=torch.randn(1, 4096),
    topk_ids=torch.tensor([[1, 3]]),
    topk_weights=torch.tensor([[0.6, 0.4]]),
    cuda_stream=None              # Set stream for GPU, None for CPU-only

)

```

### SGLang Server Integration

Launch a server with hot experts on GPU and cold experts on CPU:

```bash
python -m sglang.launch_server \
  --model /mnt/models/Qwen3-30B-A3B \
  --kt-method AMXINT8 \
  --kt-weight-path /mnt/models/Qwen3-30B-A3B-INT8 \
  --kt-cpuinfer 64 \
  --kt-threadpool-count 2 \
  --kt-num-gpu-experts 32 \
  --kt-max-deferred-experts-per-token 2 \
  --kt-enable-dynamic-expert-update

```

### Inspecting CPU Capabilities

Verify the selected CPU variant at runtime:

```python
import kt_kernel
print(f"CPU variant: {kt_kernel.__cpu_variant__}")
print(f"KT-Kernel version: {kt_kernel.__version__}")

```

## Summary

- **Hybrid execution model** — KT-Kernel partitions MoE workloads between CPU-optimized kernels and GPU-accelerated hot experts, maximizing hardware utilization.
- **Automatic optimization** — The runtime automatically selects the best CPU instruction set (AMX, AVX-512 variants) via `kt_kernel.__cpu_variant__` without manual tuning.
- **NUMA-aware scaling** — Per-socket thread pools and lock-free task queues in [`worker_pool.h`](https://github.com/kvcache-ai/ktransformers/blob/main/worker_pool.h) and [`task_queue.h`](https://github.com/kvcache-ai/ktransformers/blob/main/task_queue.h) minimize cross-socket memory traffic.
- **Zero-dependency GPU support** — Static CUDA runtime eliminates installation requirements beyond driver ≥ 11.8.
- **Flexible deployment** — The `KTMoEWrapper` Python API and SGLang integration support both standalone scripts and production server deployments.

## Frequently Asked Questions

### What makes KT-Kernel different from standard GPU inference engines?

KT-Kernel specifically optimizes for sparse MoE architectures by executing the majority of experts on CPU while reserving GPU resources for frequently-accessed "hot" experts. According to the ktransformers source code, this approach leverages modern CPU SIMD extensions (AVX-512, AMX) to handle bulk expert computation while maintaining low latency for the GPU-critical path.

### How does automatic CPU variant detection work?

On package import, KT-Kernel probes the host CPU capabilities and sets `kt_kernel.__cpu_variant__` to the optimal backend—prioritizing AMX on Intel Sapphire Rapids or newer, then falling back through AVX512+BF16, AVX512+VNNI, AVX512 Base, and AVX2. This selection happens automatically in [`kt_kernel/__init__.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt_kernel/__init__.py), ensuring the fastest available instructions are used without user configuration.

### What is the difference between hot and cold experts in KT-Kernel?

**Hot experts** are pinned to GPU memory via `--kt-num-gpu-experts` and accessed through the static CUDA runtime for minimal latency. **Cold experts** reside in system RAM, execute on CPU-optimized kernels in `kt-kernel/operators/`, and are quantized to INT4/INT8 to reduce memory bandwidth pressure. The `KTMoEWrapper` routes each token to the appropriate backend based on expert ID.

### Why is NUMA awareness critical for KT-Kernel performance?

The [`cpu_backend/worker_pool.h`](https://github.com/kvcache-ai/ktransformers/blob/main/cpu_backend/worker_pool.h) implementation creates discrete thread pools per NUMA node and pins threads to local sockets. This design maximizes memory bandwidth for large-scale CPUs (e.g., dual-socket Intel Xeon systems) by ensuring expert weights and intermediate activations reside in local socket memory, eliminating cross-socket traffic that would otherwise bottleneck CPU inference throughput.