# How to Integrate KTransformers with SGLang for Heterogeneous LLM Serving

> Learn to integrate KTransformers with SGLang for efficient LLM serving. Optimize performance by routing experts to CPU and GPU and leverage NUMA-aware thread pools.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-26

---

**You can integrate KTransformers with SGLang by installing the kvcache-ai fork of SGLang, preparing separate CPU and GPU weight sets, and launching the server with `--kt-*` flags that route "hot" experts to the GPU and "cold" experts to CPU cores via NUMA-aware thread pools.**

Integrating KTransformers with SGLang enables heterogeneous inference for large language models by combining KTransformers' high-performance CPU-side Mixture-of-Experts (MoE) kernels with SGLang's GPU inference engine. This architecture allows you to serve models with frequently accessed "hot" experts running on GPU while dispatching less active "cold" experts to CPU, optimizing both latency and memory usage on multi-socket NUMA servers. The `kvcache-ai/ktransformers` repository provides the `kt-kernel` module that seamlessly connects with SGLang through specialized command-line flags and weight conversion utilities.

## Architectural Overview

### Runtime Components

The integration creates a unified serving stack with two primary components. The **SGLang server** handles model loading, tokenization, request routing, and GPU inference, launched via `python -m sglang.launch_server`. The **kt-kernel** supplies a CPU MoE backend that creates thread pools per NUMA node and executes CPU experts using AVX-512, AMX, or LLAMAFILE kernels. These components communicate through a set of `--kt-*` command-line flags registered by the kvcache-ai fork.

### Data Flow

Model weights require dual preparation: standard GPU weights for SGLang and a CPU-optimized weight directory produced by the conversion script. At startup, the server constructs a unified routing table that directs tokens to GPU or CPU based on expert placement. The `--kt-num-gpu-experts` parameter determines how many experts remain on GPU, while the CPU backend handles remaining experts through parallel thread pools. During inference, GPU execution proceeds asynchronously while CPU kernels run in parallel across NUMA nodes.

### Key Runtime Parameters

| Parameter | Purpose | Typical Configuration |
|-----------|---------|---------------------|
| `--kt-method` | Selects CPU backend implementation: `AMXINT8`, `BF16`, or `LLAMAFILE` | `AMXINT8` for Intel Sapphire Rapids; `LLAMAFILE` for universal CPUs |
| `--kt-weight-path` | Filesystem path to CPU-side quantized weights | `/path/to/cpu-weights` |
| `--kt-cpuinfer` | Number of CPU inference threads (physical cores) | `64` for 2-socket 64-core systems |
| `--kt-threadpool-count` | Number of thread pools mapping to NUMA nodes | `2` for dual-socket servers |
| `--kt-num-gpu-experts` | Count of MoE experts retained on GPU | `32` (adjust based on VRAM) |
| `--kt-max-deferred-experts-per-token` | Enables pipelined execution; 0 disables async | `2` for latency/quality balance |
| `--kt-gpu-prefill-token-threshold` | Token threshold for layer-wise GPU prefill | `2048-4096` |

## Installation and Prerequisites

You must install the kvcache-ai fork of SGLang to enable kt-kernel support. The repository provides automated installation via [`install.sh`](https://github.com/kvcache-ai/ktransformers/blob/main/install.sh) or manual pip installation.

```bash

# Automated installation from repository root

./install.sh

# Or manual installation

pip install kt-kernel sglang-kt

```

The `sglang-kt` package registers all `--kt-*` flags automatically. Verify installation using the checker utility in [`kt-kernel/python/cli/utils/sglang_checker.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/cli/utils/sglang_checker.py), which validates that the forked SGLang supports kt-kernel parameters.

## Preparing Model Weights

### GPU Weights

Download or prepare the standard GPU weights directory as you would for native SGLang inference. This directory contains the model files normally loaded by the SGLang server.

### CPU Weight Conversion

Convert GPU weights to CPU-optimized formats using the conversion script located at [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py). This script rewrites expert tensors into formats required by your selected `--kt-method`.

```bash
python scripts/convert_cpu_weights.py \
    --input-path /path/to/model \
    --input-type bf16 \
    --output /path/to/cpu-weights \
    --quant-method int8

```

Supported quantization methods include `int8`, `int4`, and `moe_int8` for AMD backends. The output directory specified becomes the `--kt-weight-path` parameter during server launch.

## Launching the Integrated Server

### Basic Launch Command

Launch the heterogeneous server by appending kt-kernel flags to the standard SGLang launch command. The implementation in [`kt-kernel/python/cli/commands/run.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/cli/commands/run.py) builds these commands internally when using `kt run`.

```bash
python -m sglang.launch_server \
    --host 0.0.0.0 \
    --port 8000 \
    --model /mnt/data/models/Qwen3-30B-A3B \
    --trust-remote-code \
    --mem-fraction-static 0.92 \
    --chunked-prefill-size 4096 \
    --served-model-name Qwen3-30B-A3B \
    --enable-mixed-chunk \
    --kt-method AMXINT8 \
    --kt-weight-path /mnt/data/models/Qwen3-30B-A3B-INT8 \
    --kt-cpuinfer 64 \
    --kt-threadpool-count 2 \
    --kt-num-gpu-experts 32 \
    --kt-max-deferred-experts-per-token 2

```

### Dynamic Expert Placement

Enable automatic expert rebalancing by adding `--kt-enable-dynamic-expert-update`. This feature collects routing statistics during initial prefills and adjusts GPU expert placement automatically, optimizing for workloads where the optimal GPU-to-CPU ratio is unknown a priori.

## Verification and Client Usage

### Python Client Integration

After server startup, verify the integration by checking logs for the message `SGLang kt-kernel support verified`. Query the server using the OpenAI-compatible API:

```python
import sglang as sg
from sglang import OpenAIClient

client = OpenAIClient(base_url="http://localhost:8000/v1")
resp = client.chat.completions.create(
    model="Qwen3",
    messages=[{"role": "user", "content": "Explain the difference between AMX and AVX512"}],
    max_tokens=256,
)
print(resp.choices[0].message.content)

```

### Direct CPU MoE Wrapper Usage

For CPU-only inference without SGLang, instantiate `KTMoEWrapper` directly from `kt_kernel`:

```python
from kt_kernel import KTMoEWrapper
import torch

wrapper = KTMoEWrapper(
    layer_idx=0,
    num_experts=8,
    num_experts_per_tok=2,
    hidden_size=4096,
    moe_intermediate_size=14336,
    num_gpu_experts=0,
    cpuinfer_threads=64,
    threadpool_count=2,
    weight_path="/path/to/cpu-weights",
    chunked_prefill_size=512,
    method="AMXINT8",
)

wrapper.load_weights()

hidden = torch.randn(1, 4096, dtype=torch.float16)
topk_ids = torch.tensor([[0, 1]], dtype=torch.int64)
topk_weights = torch.tensor([[0.6, 0.4]], dtype=torch.float16)

out = wrapper.forward(hidden, topk_ids, topk_weights, None)

```

## Summary

- Install the kvcache-ai fork (`sglang-kt`) to enable `--kt-*` flag support when integrating KTransformers with SGLang.
- Prepare dual weight sets: standard GPU weights for SGLang and CPU-optimized weights via [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py).
- Configure heterogeneous execution using `--kt-num-gpu-experts` to balance GPU VRAM and CPU compute resources.
- Enable NUMA-aware parallelism by setting `--kt-threadpool-count` to match your system's socket count.
- Use `--kt-enable-dynamic-expert-update` for automatic optimization of expert placement during runtime.

## Frequently Asked Questions

### What is the difference between `--kt-method AMXINT8` and `LLAMAFILE`?

**AMXINT8** utilizes Intel Advanced Matrix Extensions (AMX) instructions available on Sapphire Rapids and newer CPUs, providing optimized int8 matrix multiplication. **LLAMAFILE** offers broader CPU compatibility using portable SIMD implementations suitable for various x86 and ARM architectures. Choose AMXINT8 for Intel server platforms with AMX support; use LLAMAFILE for universal deployment across heterogeneous hardware.

### How do I determine the correct value for `--kt-num-gpu-experts`?

Set this value based on your GPU's VRAM capacity and the model's expert size. Start with 32 experts for a 30B parameter model on an A100-80GB, then monitor GPU memory utilization. If OOM errors occur, reduce the value; if VRAM remains underutilized, increase until you reach the optimal balance for your specific workload's expert access patterns.

### Can I run KTransformers without SGLang for CPU-only inference?

Yes. The `kt_kernel` module provides `KTMoEWrapper` for direct CPU inference without GPU dependencies. Set `num_gpu_experts=0` in the wrapper initialization and provide only CPU weights via the `weight_path` parameter. This configuration runs pure CPU MoE inference using the AVX-512/AMX kernels without requiring the SGLang server components.

### Where can I diagnose integration issues between KTransformers and SGLang?

Use the diagnostic command implemented in [`kt-kernel/python/cli/commands/doctor.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/cli/commands/doctor.py). This utility checks CPU instruction set detection (AVX-512, AMX), CUDA availability, and SGLang fork compatibility. Run this before launching the server to verify that your environment meets all requirements for heterogeneous LLM serving.