How to Integrate KTransformers with SGLang for Heterogeneous LLM Serving
You can integrate KTransformers with SGLang by installing the kvcache-ai fork of SGLang, preparing separate CPU and GPU weight sets, and launching the server with --kt-* flags that route "hot" experts to the GPU and "cold" experts to CPU cores via NUMA-aware thread pools.
Integrating KTransformers with SGLang enables heterogeneous inference for large language models by combining KTransformers' high-performance CPU-side Mixture-of-Experts (MoE) kernels with SGLang's GPU inference engine. This architecture allows you to serve models with frequently accessed "hot" experts running on GPU while dispatching less active "cold" experts to CPU, optimizing both latency and memory usage on multi-socket NUMA servers. The kvcache-ai/ktransformers repository provides the kt-kernel module that seamlessly connects with SGLang through specialized command-line flags and weight conversion utilities.
Architectural Overview
Runtime Components
The integration creates a unified serving stack with two primary components. The SGLang server handles model loading, tokenization, request routing, and GPU inference, launched via python -m sglang.launch_server. The kt-kernel supplies a CPU MoE backend that creates thread pools per NUMA node and executes CPU experts using AVX-512, AMX, or LLAMAFILE kernels. These components communicate through a set of --kt-* command-line flags registered by the kvcache-ai fork.
Data Flow
Model weights require dual preparation: standard GPU weights for SGLang and a CPU-optimized weight directory produced by the conversion script. At startup, the server constructs a unified routing table that directs tokens to GPU or CPU based on expert placement. The --kt-num-gpu-experts parameter determines how many experts remain on GPU, while the CPU backend handles remaining experts through parallel thread pools. During inference, GPU execution proceeds asynchronously while CPU kernels run in parallel across NUMA nodes.
Key Runtime Parameters
| Parameter | Purpose | Typical Configuration |
|---|---|---|
--kt-method |
Selects CPU backend implementation: AMXINT8, BF16, or LLAMAFILE |
AMXINT8 for Intel Sapphire Rapids; LLAMAFILE for universal CPUs |
--kt-weight-path |
Filesystem path to CPU-side quantized weights | /path/to/cpu-weights |
--kt-cpuinfer |
Number of CPU inference threads (physical cores) | 64 for 2-socket 64-core systems |
--kt-threadpool-count |
Number of thread pools mapping to NUMA nodes | 2 for dual-socket servers |
--kt-num-gpu-experts |
Count of MoE experts retained on GPU | 32 (adjust based on VRAM) |
--kt-max-deferred-experts-per-token |
Enables pipelined execution; 0 disables async | 2 for latency/quality balance |
--kt-gpu-prefill-token-threshold |
Token threshold for layer-wise GPU prefill | 2048-4096 |
Installation and Prerequisites
You must install the kvcache-ai fork of SGLang to enable kt-kernel support. The repository provides automated installation via install.sh or manual pip installation.
# Automated installation from repository root
./install.sh
# Or manual installation
pip install kt-kernel sglang-kt
The sglang-kt package registers all --kt-* flags automatically. Verify installation using the checker utility in kt-kernel/python/cli/utils/sglang_checker.py, which validates that the forked SGLang supports kt-kernel parameters.
Preparing Model Weights
GPU Weights
Download or prepare the standard GPU weights directory as you would for native SGLang inference. This directory contains the model files normally loaded by the SGLang server.
CPU Weight Conversion
Convert GPU weights to CPU-optimized formats using the conversion script located at kt-kernel/scripts/convert_cpu_weights.py. This script rewrites expert tensors into formats required by your selected --kt-method.
python scripts/convert_cpu_weights.py \
--input-path /path/to/model \
--input-type bf16 \
--output /path/to/cpu-weights \
--quant-method int8
Supported quantization methods include int8, int4, and moe_int8 for AMD backends. The output directory specified becomes the --kt-weight-path parameter during server launch.
Launching the Integrated Server
Basic Launch Command
Launch the heterogeneous server by appending kt-kernel flags to the standard SGLang launch command. The implementation in kt-kernel/python/cli/commands/run.py builds these commands internally when using kt run.
python -m sglang.launch_server \
--host 0.0.0.0 \
--port 8000 \
--model /mnt/data/models/Qwen3-30B-A3B \
--trust-remote-code \
--mem-fraction-static 0.92 \
--chunked-prefill-size 4096 \
--served-model-name Qwen3-30B-A3B \
--enable-mixed-chunk \
--kt-method AMXINT8 \
--kt-weight-path /mnt/data/models/Qwen3-30B-A3B-INT8 \
--kt-cpuinfer 64 \
--kt-threadpool-count 2 \
--kt-num-gpu-experts 32 \
--kt-max-deferred-experts-per-token 2
Dynamic Expert Placement
Enable automatic expert rebalancing by adding --kt-enable-dynamic-expert-update. This feature collects routing statistics during initial prefills and adjusts GPU expert placement automatically, optimizing for workloads where the optimal GPU-to-CPU ratio is unknown a priori.
Verification and Client Usage
Python Client Integration
After server startup, verify the integration by checking logs for the message SGLang kt-kernel support verified. Query the server using the OpenAI-compatible API:
import sglang as sg
from sglang import OpenAIClient
client = OpenAIClient(base_url="http://localhost:8000/v1")
resp = client.chat.completions.create(
model="Qwen3",
messages=[{"role": "user", "content": "Explain the difference between AMX and AVX512"}],
max_tokens=256,
)
print(resp.choices[0].message.content)
Direct CPU MoE Wrapper Usage
For CPU-only inference without SGLang, instantiate KTMoEWrapper directly from kt_kernel:
from kt_kernel import KTMoEWrapper
import torch
wrapper = KTMoEWrapper(
layer_idx=0,
num_experts=8,
num_experts_per_tok=2,
hidden_size=4096,
moe_intermediate_size=14336,
num_gpu_experts=0,
cpuinfer_threads=64,
threadpool_count=2,
weight_path="/path/to/cpu-weights",
chunked_prefill_size=512,
method="AMXINT8",
)
wrapper.load_weights()
hidden = torch.randn(1, 4096, dtype=torch.float16)
topk_ids = torch.tensor([[0, 1]], dtype=torch.int64)
topk_weights = torch.tensor([[0.6, 0.4]], dtype=torch.float16)
out = wrapper.forward(hidden, topk_ids, topk_weights, None)
Summary
- Install the kvcache-ai fork (
sglang-kt) to enable--kt-*flag support when integrating KTransformers with SGLang. - Prepare dual weight sets: standard GPU weights for SGLang and CPU-optimized weights via
kt-kernel/scripts/convert_cpu_weights.py. - Configure heterogeneous execution using
--kt-num-gpu-expertsto balance GPU VRAM and CPU compute resources. - Enable NUMA-aware parallelism by setting
--kt-threadpool-countto match your system's socket count. - Use
--kt-enable-dynamic-expert-updatefor automatic optimization of expert placement during runtime.
Frequently Asked Questions
What is the difference between --kt-method AMXINT8 and LLAMAFILE?
AMXINT8 utilizes Intel Advanced Matrix Extensions (AMX) instructions available on Sapphire Rapids and newer CPUs, providing optimized int8 matrix multiplication. LLAMAFILE offers broader CPU compatibility using portable SIMD implementations suitable for various x86 and ARM architectures. Choose AMXINT8 for Intel server platforms with AMX support; use LLAMAFILE for universal deployment across heterogeneous hardware.
How do I determine the correct value for --kt-num-gpu-experts?
Set this value based on your GPU's VRAM capacity and the model's expert size. Start with 32 experts for a 30B parameter model on an A100-80GB, then monitor GPU memory utilization. If OOM errors occur, reduce the value; if VRAM remains underutilized, increase until you reach the optimal balance for your specific workload's expert access patterns.
Can I run KTransformers without SGLang for CPU-only inference?
Yes. The kt_kernel module provides KTMoEWrapper for direct CPU inference without GPU dependencies. Set num_gpu_experts=0 in the wrapper initialization and provide only CPU weights via the weight_path parameter. This configuration runs pure CPU MoE inference using the AVX-512/AMX kernels without requiring the SGLang server components.
Where can I diagnose integration issues between KTransformers and SGLang?
Use the diagnostic command implemented in kt-kernel/python/cli/commands/doctor.py. This utility checks CPU instruction set detection (AVX-512, AMX), CUDA availability, and SGLang fork compatibility. Run this before launching the server to verify that your environment meets all requirements for heterogeneous LLM serving.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →