KT-Kernel CPU-GPU Heterogeneous Computing Architecture Explained
KT-Kernel is a hybrid inference engine that executes Mixture-of-Experts layers across CPU and GPU by partitioning cold experts onto CPU-optimized SIMD kernels while reserving hot experts for GPU acceleration.
The kvcache-ai/ktransformers repository implements KT-Kernel to enable high-throughput inference for sparse large language models. This architecture automatically partitions computational workloads between host CPUs and discrete GPUs, leveraging AVX-512/AMX instructions for CPU-side execution while embedding a static CUDA runtime to eliminate external toolkit dependencies.
High-Level Architecture Components
KT-Kernel separates the CPU-only MoE backend from the GPU-accelerated hot-expert backend, orchestrating both through a NUMA-aware thread-pool and lightweight task-queue.
| Component | Role | Implementation |
|---|---|---|
| Automatic CPU variant detection | Selects optimal kernel variant (AMX, AVX512+BF16, AVX512+VNNI, AVX512 Base, AVX2) at import time | kt_kernel.__cpu_variant__ runtime detection |
| CPU-optimized MoE kernels | Native-precision (FP8/BF16/INT4) and quantized (INT8) kernels for MoE experts | Operators in kt-kernel/operators/ (e.g., softmax.hpp, rope.hpp, tp.hpp) |
| NUMA-aware thread-pool | Creates one pool per NUMA node, pins threads to local sockets, allocates per-socket shared buffers | cpu_backend/worker_pool.h & worker_pool.cpp |
| Task-queue & shared-memory | Lock-free queue pushing inference tasks to appropriate worker pools; pre-allocated reusable buffers | cpu_backend/task_queue.h & task_queue.cpp |
| Static CUDA runtime | Ships compiled CUDA runtime inside the wheel, requiring no external CUDA toolkit | kt_kernel_ext (C++ CUDA kernels) |
| Python API | Thin wrapper exposing C++ MoE kernels to Python and SGLang | kt_kernel/__init__.py exposing KTMoEWrapper |
Heterogeneous MoE Forward Pass Data Flow
The execution pipeline enables pipeline parallelism and dynamic expert placement through six distinct stages:
-
Model loading — GPU weights are loaded by SGLang while CPU-only expert weights are quantized (INT4/INT8) and stored under
--kt-weight-path. -
Kernel selection — On import,
kt_kerneldetects host CPU capabilities and sets__cpu_variant__to the optimal backend (e.g.,AMX,AVX512+BF16). -
Task creation — The Python
KTMoEWrapperbuilds a forward-task containing hidden states, top-k expert ids, and weights, then pushes it onto theTaskQueue. -
Thread-pool dispatch — The NUMA-local worker thread picks the task and maps expert ids to either the CPU backend (via compiled SIMD kernels) or the GPU backend.
-
GPU execution — For experts marked hot (
--kt-num-gpu-experts), the wrapper callsCPUInfer.submit_with_cuda_streamto stream data to the GPU kernel; otherwise, the CPU kernel runs directly. -
Result collection — The worker writes output into a pre-allocated shared buffer defined in
cpu_backend/shared_mem_buffer.h, allowing the Python side to block (sync_forward) or continue asynchronously.
Key Source Files and Implementation
KT-Kernel's implementation spans C++ SIMD kernels, CUDA operations, and Python bindings:
-
kt-kernel/operators/*.hpp— Contains SIMD-optimized kernels for MoE operations including softmax, rotary-positional embeddings, and tensor-parallel logic. -
kt-kernel/cpu_backend/worker_pool.h— Implements the NUMA-aware thread-pool that pins workers to specific sockets and manages per-socket buffer allocation. -
kt-kernel/cpu_backend/task_queue.h— Defines the lock-free queue dispatching MoE forward tasks to worker pools with minimal latency. -
kt-kernel/cpu_backend/cpuinfer.h— The C++ entry point exposing the CPU MoE inference API to Python viapybind11, referenced from Python askt_kernel_ext.CPUInfer. -
kt-kernel/python/kt_kernel/__init__.py— Python package entry definingKTMoEWrapperand re-exporting version constants. -
kt-kernel/scripts/convert_cpu_weights.py— Helper script quantizing GPU-side weights into AMX-friendly INT4/INT8 formats required by the CPU backend.
Usage Examples
Direct Python API
Create a heterogeneous MoE wrapper and execute forward passes:
from kt_kernel import KTMoEWrapper
# Initialize wrapper for heterogeneous execution
wrapper = KTMoEWrapper(
layer_idx=0,
num_experts=8,
num_experts_per_tok=2,
hidden_size=4096,
moe_intermediate_size=14336,
num_gpu_experts=2, # hot experts on GPU
cpuinfer_threads=32, # physical cores
threadpool_count=2, # NUMA nodes
weight_path="/path/to/cpu-weights",
method="AMXINT8", # Intel AMX backend
)
wrapper.load_weights()
# Execute heterogeneous forward pass
output = wrapper.forward(
hidden_states=torch.randn(1, 4096),
topk_ids=torch.tensor([[1, 3]]),
topk_weights=torch.tensor([[0.6, 0.4]]),
cuda_stream=None # Set stream for GPU, None for CPU-only
)
SGLang Server Integration
Launch a server with hot experts on GPU and cold experts on CPU:
python -m sglang.launch_server \
--model /mnt/models/Qwen3-30B-A3B \
--kt-method AMXINT8 \
--kt-weight-path /mnt/models/Qwen3-30B-A3B-INT8 \
--kt-cpuinfer 64 \
--kt-threadpool-count 2 \
--kt-num-gpu-experts 32 \
--kt-max-deferred-experts-per-token 2 \
--kt-enable-dynamic-expert-update
Inspecting CPU Capabilities
Verify the selected CPU variant at runtime:
import kt_kernel
print(f"CPU variant: {kt_kernel.__cpu_variant__}")
print(f"KT-Kernel version: {kt_kernel.__version__}")
Summary
- Hybrid execution model — KT-Kernel partitions MoE workloads between CPU-optimized kernels and GPU-accelerated hot experts, maximizing hardware utilization.
- Automatic optimization — The runtime automatically selects the best CPU instruction set (AMX, AVX-512 variants) via
kt_kernel.__cpu_variant__without manual tuning. - NUMA-aware scaling — Per-socket thread pools and lock-free task queues in
worker_pool.handtask_queue.hminimize cross-socket memory traffic. - Zero-dependency GPU support — Static CUDA runtime eliminates installation requirements beyond driver ≥ 11.8.
- Flexible deployment — The
KTMoEWrapperPython API and SGLang integration support both standalone scripts and production server deployments.
Frequently Asked Questions
What makes KT-Kernel different from standard GPU inference engines?
KT-Kernel specifically optimizes for sparse MoE architectures by executing the majority of experts on CPU while reserving GPU resources for frequently-accessed "hot" experts. According to the ktransformers source code, this approach leverages modern CPU SIMD extensions (AVX-512, AMX) to handle bulk expert computation while maintaining low latency for the GPU-critical path.
How does automatic CPU variant detection work?
On package import, KT-Kernel probes the host CPU capabilities and sets kt_kernel.__cpu_variant__ to the optimal backend—prioritizing AMX on Intel Sapphire Rapids or newer, then falling back through AVX512+BF16, AVX512+VNNI, AVX512 Base, and AVX2. This selection happens automatically in kt_kernel/__init__.py, ensuring the fastest available instructions are used without user configuration.
What is the difference between hot and cold experts in KT-Kernel?
Hot experts are pinned to GPU memory via --kt-num-gpu-experts and accessed through the static CUDA runtime for minimal latency. Cold experts reside in system RAM, execute on CPU-optimized kernels in kt-kernel/operators/, and are quantized to INT4/INT8 to reduce memory bandwidth pressure. The KTMoEWrapper routes each token to the appropriate backend based on expert ID.
Why is NUMA awareness critical for KT-Kernel performance?
The cpu_backend/worker_pool.h implementation creates discrete thread pools per NUMA node and pins threads to local sockets. This design maximizes memory bandwidth for large-scale CPUs (e.g., dual-socket Intel Xeon systems) by ensuring expert weights and intermediate activations reside in local socket memory, eliminating cross-socket traffic that would otherwise bottleneck CPU inference throughput.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →