How to Integrate KTransformers with SGLang for Heterogeneous LLM Serving

You can integrate KTransformers with SGLang by installing the kvcache-ai fork of SGLang, preparing separate CPU and GPU weight sets, and launching the server with --kt-* flags that route "hot" experts to the GPU and "cold" experts to CPU cores via NUMA-aware thread pools.

Integrating KTransformers with SGLang enables heterogeneous inference for large language models by combining KTransformers' high-performance CPU-side Mixture-of-Experts (MoE) kernels with SGLang's GPU inference engine. This architecture allows you to serve models with frequently accessed "hot" experts running on GPU while dispatching less active "cold" experts to CPU, optimizing both latency and memory usage on multi-socket NUMA servers. The kvcache-ai/ktransformers repository provides the kt-kernel module that seamlessly connects with SGLang through specialized command-line flags and weight conversion utilities.

Architectural Overview

Runtime Components

The integration creates a unified serving stack with two primary components. The SGLang server handles model loading, tokenization, request routing, and GPU inference, launched via python -m sglang.launch_server. The kt-kernel supplies a CPU MoE backend that creates thread pools per NUMA node and executes CPU experts using AVX-512, AMX, or LLAMAFILE kernels. These components communicate through a set of --kt-* command-line flags registered by the kvcache-ai fork.

Data Flow

Model weights require dual preparation: standard GPU weights for SGLang and a CPU-optimized weight directory produced by the conversion script. At startup, the server constructs a unified routing table that directs tokens to GPU or CPU based on expert placement. The --kt-num-gpu-experts parameter determines how many experts remain on GPU, while the CPU backend handles remaining experts through parallel thread pools. During inference, GPU execution proceeds asynchronously while CPU kernels run in parallel across NUMA nodes.

Key Runtime Parameters

Parameter Purpose Typical Configuration
--kt-method Selects CPU backend implementation: AMXINT8, BF16, or LLAMAFILE AMXINT8 for Intel Sapphire Rapids; LLAMAFILE for universal CPUs
--kt-weight-path Filesystem path to CPU-side quantized weights /path/to/cpu-weights
--kt-cpuinfer Number of CPU inference threads (physical cores) 64 for 2-socket 64-core systems
--kt-threadpool-count Number of thread pools mapping to NUMA nodes 2 for dual-socket servers
--kt-num-gpu-experts Count of MoE experts retained on GPU 32 (adjust based on VRAM)
--kt-max-deferred-experts-per-token Enables pipelined execution; 0 disables async 2 for latency/quality balance
--kt-gpu-prefill-token-threshold Token threshold for layer-wise GPU prefill 2048-4096

Installation and Prerequisites

You must install the kvcache-ai fork of SGLang to enable kt-kernel support. The repository provides automated installation via install.sh or manual pip installation.


# Automated installation from repository root

./install.sh

# Or manual installation

pip install kt-kernel sglang-kt

The sglang-kt package registers all --kt-* flags automatically. Verify installation using the checker utility in kt-kernel/python/cli/utils/sglang_checker.py, which validates that the forked SGLang supports kt-kernel parameters.

Preparing Model Weights

GPU Weights

Download or prepare the standard GPU weights directory as you would for native SGLang inference. This directory contains the model files normally loaded by the SGLang server.

CPU Weight Conversion

Convert GPU weights to CPU-optimized formats using the conversion script located at kt-kernel/scripts/convert_cpu_weights.py. This script rewrites expert tensors into formats required by your selected --kt-method.

python scripts/convert_cpu_weights.py \
    --input-path /path/to/model \
    --input-type bf16 \
    --output /path/to/cpu-weights \
    --quant-method int8

Supported quantization methods include int8, int4, and moe_int8 for AMD backends. The output directory specified becomes the --kt-weight-path parameter during server launch.

Launching the Integrated Server

Basic Launch Command

Launch the heterogeneous server by appending kt-kernel flags to the standard SGLang launch command. The implementation in kt-kernel/python/cli/commands/run.py builds these commands internally when using kt run.

python -m sglang.launch_server \
    --host 0.0.0.0 \
    --port 8000 \
    --model /mnt/data/models/Qwen3-30B-A3B \
    --trust-remote-code \
    --mem-fraction-static 0.92 \
    --chunked-prefill-size 4096 \
    --served-model-name Qwen3-30B-A3B \
    --enable-mixed-chunk \
    --kt-method AMXINT8 \
    --kt-weight-path /mnt/data/models/Qwen3-30B-A3B-INT8 \
    --kt-cpuinfer 64 \
    --kt-threadpool-count 2 \
    --kt-num-gpu-experts 32 \
    --kt-max-deferred-experts-per-token 2

Dynamic Expert Placement

Enable automatic expert rebalancing by adding --kt-enable-dynamic-expert-update. This feature collects routing statistics during initial prefills and adjusts GPU expert placement automatically, optimizing for workloads where the optimal GPU-to-CPU ratio is unknown a priori.

Verification and Client Usage

Python Client Integration

After server startup, verify the integration by checking logs for the message SGLang kt-kernel support verified. Query the server using the OpenAI-compatible API:

import sglang as sg
from sglang import OpenAIClient

client = OpenAIClient(base_url="http://localhost:8000/v1")
resp = client.chat.completions.create(
    model="Qwen3",
    messages=[{"role": "user", "content": "Explain the difference between AMX and AVX512"}],
    max_tokens=256,
)
print(resp.choices[0].message.content)

Direct CPU MoE Wrapper Usage

For CPU-only inference without SGLang, instantiate KTMoEWrapper directly from kt_kernel:

from kt_kernel import KTMoEWrapper
import torch

wrapper = KTMoEWrapper(
    layer_idx=0,
    num_experts=8,
    num_experts_per_tok=2,
    hidden_size=4096,
    moe_intermediate_size=14336,
    num_gpu_experts=0,
    cpuinfer_threads=64,
    threadpool_count=2,
    weight_path="/path/to/cpu-weights",
    chunked_prefill_size=512,
    method="AMXINT8",
)

wrapper.load_weights()

hidden = torch.randn(1, 4096, dtype=torch.float16)
topk_ids = torch.tensor([[0, 1]], dtype=torch.int64)
topk_weights = torch.tensor([[0.6, 0.4]], dtype=torch.float16)

out = wrapper.forward(hidden, topk_ids, topk_weights, None)

Summary

  • Install the kvcache-ai fork (sglang-kt) to enable --kt-* flag support when integrating KTransformers with SGLang.
  • Prepare dual weight sets: standard GPU weights for SGLang and CPU-optimized weights via kt-kernel/scripts/convert_cpu_weights.py.
  • Configure heterogeneous execution using --kt-num-gpu-experts to balance GPU VRAM and CPU compute resources.
  • Enable NUMA-aware parallelism by setting --kt-threadpool-count to match your system's socket count.
  • Use --kt-enable-dynamic-expert-update for automatic optimization of expert placement during runtime.

Frequently Asked Questions

What is the difference between --kt-method AMXINT8 and LLAMAFILE?

AMXINT8 utilizes Intel Advanced Matrix Extensions (AMX) instructions available on Sapphire Rapids and newer CPUs, providing optimized int8 matrix multiplication. LLAMAFILE offers broader CPU compatibility using portable SIMD implementations suitable for various x86 and ARM architectures. Choose AMXINT8 for Intel server platforms with AMX support; use LLAMAFILE for universal deployment across heterogeneous hardware.

How do I determine the correct value for --kt-num-gpu-experts?

Set this value based on your GPU's VRAM capacity and the model's expert size. Start with 32 experts for a 30B parameter model on an A100-80GB, then monitor GPU memory utilization. If OOM errors occur, reduce the value; if VRAM remains underutilized, increase until you reach the optimal balance for your specific workload's expert access patterns.

Can I run KTransformers without SGLang for CPU-only inference?

Yes. The kt_kernel module provides KTMoEWrapper for direct CPU inference without GPU dependencies. Set num_gpu_experts=0 in the wrapper initialization and provide only CPU weights via the weight_path parameter. This configuration runs pure CPU MoE inference using the AVX-512/AMX kernels without requiring the SGLang server components.

Where can I diagnose integration issues between KTransformers and SGLang?

Use the diagnostic command implemented in kt-kernel/python/cli/commands/doctor.py. This utility checks CPU instruction set detection (AVX-512, AMX), CUDA availability, and SGLang fork compatibility. Run this before launching the server to verify that your environment meets all requirements for heterogeneous LLM serving.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →