KT-Kernel CPU-GPU Heterogeneous Computing Architecture Explained

KT-Kernel is a hybrid inference engine that executes Mixture-of-Experts layers across CPU and GPU by partitioning cold experts onto CPU-optimized SIMD kernels while reserving hot experts for GPU acceleration.

The kvcache-ai/ktransformers repository implements KT-Kernel to enable high-throughput inference for sparse large language models. This architecture automatically partitions computational workloads between host CPUs and discrete GPUs, leveraging AVX-512/AMX instructions for CPU-side execution while embedding a static CUDA runtime to eliminate external toolkit dependencies.

High-Level Architecture Components

KT-Kernel separates the CPU-only MoE backend from the GPU-accelerated hot-expert backend, orchestrating both through a NUMA-aware thread-pool and lightweight task-queue.

Component Role Implementation
Automatic CPU variant detection Selects optimal kernel variant (AMX, AVX512+BF16, AVX512+VNNI, AVX512 Base, AVX2) at import time kt_kernel.__cpu_variant__ runtime detection
CPU-optimized MoE kernels Native-precision (FP8/BF16/INT4) and quantized (INT8) kernels for MoE experts Operators in kt-kernel/operators/ (e.g., softmax.hpp, rope.hpp, tp.hpp)
NUMA-aware thread-pool Creates one pool per NUMA node, pins threads to local sockets, allocates per-socket shared buffers cpu_backend/worker_pool.h & worker_pool.cpp
Task-queue & shared-memory Lock-free queue pushing inference tasks to appropriate worker pools; pre-allocated reusable buffers cpu_backend/task_queue.h & task_queue.cpp
Static CUDA runtime Ships compiled CUDA runtime inside the wheel, requiring no external CUDA toolkit kt_kernel_ext (C++ CUDA kernels)
Python API Thin wrapper exposing C++ MoE kernels to Python and SGLang kt_kernel/__init__.py exposing KTMoEWrapper

Heterogeneous MoE Forward Pass Data Flow

The execution pipeline enables pipeline parallelism and dynamic expert placement through six distinct stages:

  1. Model loading — GPU weights are loaded by SGLang while CPU-only expert weights are quantized (INT4/INT8) and stored under --kt-weight-path.

  2. Kernel selection — On import, kt_kernel detects host CPU capabilities and sets __cpu_variant__ to the optimal backend (e.g., AMX, AVX512+BF16).

  3. Task creation — The Python KTMoEWrapper builds a forward-task containing hidden states, top-k expert ids, and weights, then pushes it onto the TaskQueue.

  4. Thread-pool dispatch — The NUMA-local worker thread picks the task and maps expert ids to either the CPU backend (via compiled SIMD kernels) or the GPU backend.

  5. GPU execution — For experts marked hot (--kt-num-gpu-experts), the wrapper calls CPUInfer.submit_with_cuda_stream to stream data to the GPU kernel; otherwise, the CPU kernel runs directly.

  6. Result collection — The worker writes output into a pre-allocated shared buffer defined in cpu_backend/shared_mem_buffer.h, allowing the Python side to block (sync_forward) or continue asynchronously.

Key Source Files and Implementation

KT-Kernel's implementation spans C++ SIMD kernels, CUDA operations, and Python bindings:

Usage Examples

Direct Python API

Create a heterogeneous MoE wrapper and execute forward passes:

from kt_kernel import KTMoEWrapper

# Initialize wrapper for heterogeneous execution

wrapper = KTMoEWrapper(
    layer_idx=0,
    num_experts=8,
    num_experts_per_tok=2,
    hidden_size=4096,
    moe_intermediate_size=14336,
    num_gpu_experts=2,            # hot experts on GPU

    cpuinfer_threads=32,          # physical cores

    threadpool_count=2,           # NUMA nodes

    weight_path="/path/to/cpu-weights",
    method="AMXINT8",             # Intel AMX backend

)

wrapper.load_weights()

# Execute heterogeneous forward pass

output = wrapper.forward(
    hidden_states=torch.randn(1, 4096),
    topk_ids=torch.tensor([[1, 3]]),
    topk_weights=torch.tensor([[0.6, 0.4]]),
    cuda_stream=None              # Set stream for GPU, None for CPU-only

)

SGLang Server Integration

Launch a server with hot experts on GPU and cold experts on CPU:

python -m sglang.launch_server \
  --model /mnt/models/Qwen3-30B-A3B \
  --kt-method AMXINT8 \
  --kt-weight-path /mnt/models/Qwen3-30B-A3B-INT8 \
  --kt-cpuinfer 64 \
  --kt-threadpool-count 2 \
  --kt-num-gpu-experts 32 \
  --kt-max-deferred-experts-per-token 2 \
  --kt-enable-dynamic-expert-update

Inspecting CPU Capabilities

Verify the selected CPU variant at runtime:

import kt_kernel
print(f"CPU variant: {kt_kernel.__cpu_variant__}")
print(f"KT-Kernel version: {kt_kernel.__version__}")

Summary

  • Hybrid execution model — KT-Kernel partitions MoE workloads between CPU-optimized kernels and GPU-accelerated hot experts, maximizing hardware utilization.
  • Automatic optimization — The runtime automatically selects the best CPU instruction set (AMX, AVX-512 variants) via kt_kernel.__cpu_variant__ without manual tuning.
  • NUMA-aware scaling — Per-socket thread pools and lock-free task queues in worker_pool.h and task_queue.h minimize cross-socket memory traffic.
  • Zero-dependency GPU support — Static CUDA runtime eliminates installation requirements beyond driver ≥ 11.8.
  • Flexible deployment — The KTMoEWrapper Python API and SGLang integration support both standalone scripts and production server deployments.

Frequently Asked Questions

What makes KT-Kernel different from standard GPU inference engines?

KT-Kernel specifically optimizes for sparse MoE architectures by executing the majority of experts on CPU while reserving GPU resources for frequently-accessed "hot" experts. According to the ktransformers source code, this approach leverages modern CPU SIMD extensions (AVX-512, AMX) to handle bulk expert computation while maintaining low latency for the GPU-critical path.

How does automatic CPU variant detection work?

On package import, KT-Kernel probes the host CPU capabilities and sets kt_kernel.__cpu_variant__ to the optimal backend—prioritizing AMX on Intel Sapphire Rapids or newer, then falling back through AVX512+BF16, AVX512+VNNI, AVX512 Base, and AVX2. This selection happens automatically in kt_kernel/__init__.py, ensuring the fastest available instructions are used without user configuration.

What is the difference between hot and cold experts in KT-Kernel?

Hot experts are pinned to GPU memory via --kt-num-gpu-experts and accessed through the static CUDA runtime for minimal latency. Cold experts reside in system RAM, execute on CPU-optimized kernels in kt-kernel/operators/, and are quantized to INT4/INT8 to reduce memory bandwidth pressure. The KTMoEWrapper routes each token to the appropriate backend based on expert ID.

Why is NUMA awareness critical for KT-Kernel performance?

The cpu_backend/worker_pool.h implementation creates discrete thread pools per NUMA node and pins threads to local sockets. This design maximizes memory bandwidth for large-scale CPUs (e.g., dual-socket Intel Xeon systems) by ensuring expert weights and intermediate activations reside in local socket memory, eliminating cross-socket traffic that would otherwise bottleneck CPU inference throughput.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →