# How to Integrate KTransformers with SGLang for Production Model Serving

> Integrate KTransformers with SGLang for efficient production model serving. Learn to enable heterogeneous inference with GPU and CPU offloading for optimized performance.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-20

---

**Integrate KTransformers with SGLang by installing the `sglang-kt` fork, preparing separate GPU and quantized CPU weights, and launching the server with `--kt-*` parameters to enable heterogeneous inference where hot experts run on GPU and cold experts run on CPU.**

The kvcache-ai/ktransformers repository provides high-performance CPU-optimized MoE kernels that enable heterogeneous inference for large language models. By integrating KTransformers with SGLang, you can deploy production-grade serving pipelines that split computation between GPU and CPU, drastically reducing memory costs while maintaining throughput. This guide walks through the complete setup using the `kt-kernel` package and the compatible `sglang-kt` fork.

## Install the SGLang Fork and KT-Kernel

To integrate KTransformers with SGLang, you must use the `sglang-kt` package rather than the official SGLang distribution. This fork contains the necessary hooks to dispatch expert kernels to the CPU backend.

Remove any existing SGLang installation to avoid conflicts:

```bash
pip uninstall sglang -y

```

Install the integrated packages using the provided install script or pip:

```bash

# Option 1: One-click install from repository root

./install.sh

# Option 2: Direct pip install

pip install kt-kernel sglang-kt

```

Verify the installation using the included checker utility:

```bash
python -m kt_kernel.python.cli.utils.sglang_checker

```

*Reference: Installation details are documented in [`kt-kernel/README.md`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md) under the Installation Steps section.*

## Prepare Model Weights for Heterogeneous Inference

The integration requires two distinct weight sets: standard weights for GPU execution and quantized weights for CPU execution via kt-kernel.

### GPU Weights

Load your original or previously quantized model files onto the GPU as usual. These weights remain unchanged and are loaded by the standard SGLang model loader.

### CPU Weights with AMX Quantization

Convert your model to a CPU-optimized format (AMXINT4, AMXINT8, or LLAMAFILE) using the conversion utility. This reduces memory footprint and enables Intel AMX/AVX-512 acceleration.

```bash
python kt-kernel/scripts/convert_cpu_weights.py \
    --input-path /path/to/original/model \
    --input-type bf16 \
    --output /path/to/cpu-weights \
    --quant-method int8

```

*The script produces a directory containing INT8 or INT4 files that kt-kernel reads during inference. Supported methods include `int4`, `int8`, and `moe_int8`.*

*Reference: See [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py) for implementation details.*

### Convert Fused-Expert LoRA Checkpoints

If you are using fused-expert LoRA adapters, convert them to the SGLang-compatible format:

```bash
python kt-kernel/scripts/convert_kt_to_sglang_adapter.py \
    --input /path/to/kt-checkpoint \
    --output /path/to/sglang-adapter

```

*Reference: Implementation is in [`kt-kernel/scripts/convert_kt_to_sglang_adapter.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_kt_to_sglang_adapter.py).*

## Launch the SGLang Server with KT-Kernel Parameters

Start the SGLang server with extended flags that specify the CPU weight path, backend method, and heterogeneous execution configuration.

```bash
python -m sglang.launch_server \
    --host 0.0.0.0 \
    --port 30000 \
    --model /mnt/data/models/Qwen3-30B-A3B \
    --kt-weight-path /mnt/data/models/Qwen3-30B-A3B/cpu-weights \
    --kt-method AMXINT8 \
    --kt-cpuinfer 64 \
    --kt-threadpool-count 2 \
    --kt-num-gpu-experts 32 \
    --kt-max-deferred-experts-per-token 2 \
    --attention-backend flashinfer \
    --trust-remote-code \
    --mem-fraction-static 0.80 \
    --chunked-prefill-size 16384 \
    --max-running-requests 4 \
    --served-model-name Qwen3

```

**Key KT-Kernel Parameters:**

- **`--kt-method`**: Selects the CPU backend. Options include `AMXINT4`, `AMXINT8`, `BF16`, or `LLAMAFILE`.
- **`--kt-weight-path`**: Directory containing the CPU-side quantized weights generated by [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py).
- **`--kt-cpuinfer`**: Number of CPU inference threads (typically set to physical core count).
- **`--kt-threadpool-count`**: Number of thread pools, usually matched to NUMA node count for optimal memory locality.
- **`--kt-num-gpu-experts`**: Determines how many MoE experts remain on GPU (hot experts); the rest execute on CPU.
- **`--kt-max-deferred-experts-per-token`**: Enables pipelined execution, allowing CPU to begin the next token while GPU processes the current one.

*Reference: Full parameter documentation is in [`kt-kernel/README.md`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md) under KT-Kernel Parameters.*

## Run Inference via the Client SDK

Once the server is running, interact with it using the standard SGLang Python SDK. The heterogeneous execution is transparent to the client.

```python
from sglang import SGLangClient

client = SGLangClient(host="localhost", port=30000)

response = client.generate(
    prompt="Explain the benefits of heterogeneous inference with KTransformers and SGLang.",
    max_new_tokens=256,
)

print(response.output_text)

```

The client requires no special configuration; the server internally routes expert kernels to GPU or CPU based on the `--kt-num-gpu-experts` and deferred execution settings.

## Summary

- **Install** the `sglang-kt` fork and `kt-kernel` packages from the kvcache-ai repository to enable CPU/GPU heterogeneous execution.
- **Prepare dual weight sets**: Standard GPU weights and CPU-quantized weights using [`convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_cpu_weights.py) (or [`convert_kt_to_sglang_adapter.py`](https://github.com/kvcache-ai/ktransformers/blob/main/convert_kt_to_sglang_adapter.py) for LoRA adapters).
- **Launch** the server with `--kt-method`, `--kt-weight-path`, and `--kt-num-gpu-experts` to control the hot/cold expert split between GPU and CPU.
- The **client interface** remains standard SGLang; the server handles internal dispatch to kt-kernel backends (AMX, AVX-512, Llamafile) automatically.

## Frequently Asked Questions

### What is the difference between `sglang-kt` and standard SGLang?

The `sglang-kt` fork is a modified version of SGLang maintained in `third_party/sglang/python/` that includes hooks for the kt-kernel backend. It allows the server to dispatch MoE expert kernels to CPU inference engines while exposing the standard SGLang API to clients. You must uninstall the official `sglang` package before installing `sglang-kt` to avoid symbol conflicts.

### Which quantization formats does KTransformers support for CPU weights?

KTransformers supports `AMXINT4`, `AMXINT8`, `BF16`, and `LLAMAFILE` formats for CPU execution, selected via the `--kt-method` parameter. Use [`kt-kernel/scripts/convert_cpu_weights.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_cpu_weights.py) to generate these quantized files from original BF16 or FP16 checkpoints, choosing the method based on your accuracy and latency requirements.

### How do I determine the optimal number of GPU experts versus CPU experts?

Configure `--kt-num-gpu-experts` based on available GPU memory and the model's expert count (e.g., 32 hot experts for a 256-expert MoE model). Set `--kt-max-deferred-experts-per-token` to 2 or higher to enable pipelining, which allows CPU computation to overlap with GPU execution. Benchmark throughput with your specific hardware to find the optimal split.

### Can I use existing LoRA adapters with this integration?

Yes, fused-expert LoRA checkpoints can be converted using [`kt-kernel/scripts/convert_kt_to_sglang_adapter.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/scripts/convert_kt_to_sglang_adapter.py). This script transforms the checkpoints into a directory structure compatible with SGLang's adapter loading mechanism, allowing you to serve fine-tuned models through the heterogeneous inference pipeline.