How to Integrate KTransformers with SGLang for Production Model Serving

Integrate KTransformers with SGLang by installing the sglang-kt fork, preparing separate GPU and quantized CPU weights, and launching the server with --kt-* parameters to enable heterogeneous inference where hot experts run on GPU and cold experts run on CPU.

The kvcache-ai/ktransformers repository provides high-performance CPU-optimized MoE kernels that enable heterogeneous inference for large language models. By integrating KTransformers with SGLang, you can deploy production-grade serving pipelines that split computation between GPU and CPU, drastically reducing memory costs while maintaining throughput. This guide walks through the complete setup using the kt-kernel package and the compatible sglang-kt fork.

Install the SGLang Fork and KT-Kernel

To integrate KTransformers with SGLang, you must use the sglang-kt package rather than the official SGLang distribution. This fork contains the necessary hooks to dispatch expert kernels to the CPU backend.

Remove any existing SGLang installation to avoid conflicts:

pip uninstall sglang -y

Install the integrated packages using the provided install script or pip:


# Option 1: One-click install from repository root

./install.sh

# Option 2: Direct pip install

pip install kt-kernel sglang-kt

Verify the installation using the included checker utility:

python -m kt_kernel.python.cli.utils.sglang_checker

Reference: Installation details are documented in kt-kernel/README.md under the Installation Steps section.

Prepare Model Weights for Heterogeneous Inference

The integration requires two distinct weight sets: standard weights for GPU execution and quantized weights for CPU execution via kt-kernel.

GPU Weights

Load your original or previously quantized model files onto the GPU as usual. These weights remain unchanged and are loaded by the standard SGLang model loader.

CPU Weights with AMX Quantization

Convert your model to a CPU-optimized format (AMXINT4, AMXINT8, or LLAMAFILE) using the conversion utility. This reduces memory footprint and enables Intel AMX/AVX-512 acceleration.

python kt-kernel/scripts/convert_cpu_weights.py \
    --input-path /path/to/original/model \
    --input-type bf16 \
    --output /path/to/cpu-weights \
    --quant-method int8

The script produces a directory containing INT8 or INT4 files that kt-kernel reads during inference. Supported methods include int4, int8, and moe_int8.

Reference: See kt-kernel/scripts/convert_cpu_weights.py for implementation details.

Convert Fused-Expert LoRA Checkpoints

If you are using fused-expert LoRA adapters, convert them to the SGLang-compatible format:

python kt-kernel/scripts/convert_kt_to_sglang_adapter.py \
    --input /path/to/kt-checkpoint \
    --output /path/to/sglang-adapter

Reference: Implementation is in kt-kernel/scripts/convert_kt_to_sglang_adapter.py.

Launch the SGLang Server with KT-Kernel Parameters

Start the SGLang server with extended flags that specify the CPU weight path, backend method, and heterogeneous execution configuration.

python -m sglang.launch_server \
    --host 0.0.0.0 \
    --port 30000 \
    --model /mnt/data/models/Qwen3-30B-A3B \
    --kt-weight-path /mnt/data/models/Qwen3-30B-A3B/cpu-weights \
    --kt-method AMXINT8 \
    --kt-cpuinfer 64 \
    --kt-threadpool-count 2 \
    --kt-num-gpu-experts 32 \
    --kt-max-deferred-experts-per-token 2 \
    --attention-backend flashinfer \
    --trust-remote-code \
    --mem-fraction-static 0.80 \
    --chunked-prefill-size 16384 \
    --max-running-requests 4 \
    --served-model-name Qwen3

Key KT-Kernel Parameters:

  • --kt-method: Selects the CPU backend. Options include AMXINT4, AMXINT8, BF16, or LLAMAFILE.
  • --kt-weight-path: Directory containing the CPU-side quantized weights generated by convert_cpu_weights.py.
  • --kt-cpuinfer: Number of CPU inference threads (typically set to physical core count).
  • --kt-threadpool-count: Number of thread pools, usually matched to NUMA node count for optimal memory locality.
  • --kt-num-gpu-experts: Determines how many MoE experts remain on GPU (hot experts); the rest execute on CPU.
  • --kt-max-deferred-experts-per-token: Enables pipelined execution, allowing CPU to begin the next token while GPU processes the current one.

Reference: Full parameter documentation is in kt-kernel/README.md under KT-Kernel Parameters.

Run Inference via the Client SDK

Once the server is running, interact with it using the standard SGLang Python SDK. The heterogeneous execution is transparent to the client.

from sglang import SGLangClient

client = SGLangClient(host="localhost", port=30000)

response = client.generate(
    prompt="Explain the benefits of heterogeneous inference with KTransformers and SGLang.",
    max_new_tokens=256,
)

print(response.output_text)

The client requires no special configuration; the server internally routes expert kernels to GPU or CPU based on the --kt-num-gpu-experts and deferred execution settings.

Summary

  • Install the sglang-kt fork and kt-kernel packages from the kvcache-ai repository to enable CPU/GPU heterogeneous execution.
  • Prepare dual weight sets: Standard GPU weights and CPU-quantized weights using convert_cpu_weights.py (or convert_kt_to_sglang_adapter.py for LoRA adapters).
  • Launch the server with --kt-method, --kt-weight-path, and --kt-num-gpu-experts to control the hot/cold expert split between GPU and CPU.
  • The client interface remains standard SGLang; the server handles internal dispatch to kt-kernel backends (AMX, AVX-512, Llamafile) automatically.

Frequently Asked Questions

What is the difference between sglang-kt and standard SGLang?

The sglang-kt fork is a modified version of SGLang maintained in third_party/sglang/python/ that includes hooks for the kt-kernel backend. It allows the server to dispatch MoE expert kernels to CPU inference engines while exposing the standard SGLang API to clients. You must uninstall the official sglang package before installing sglang-kt to avoid symbol conflicts.

Which quantization formats does KTransformers support for CPU weights?

KTransformers supports AMXINT4, AMXINT8, BF16, and LLAMAFILE formats for CPU execution, selected via the --kt-method parameter. Use kt-kernel/scripts/convert_cpu_weights.py to generate these quantized files from original BF16 or FP16 checkpoints, choosing the method based on your accuracy and latency requirements.

How do I determine the optimal number of GPU experts versus CPU experts?

Configure --kt-num-gpu-experts based on available GPU memory and the model's expert count (e.g., 32 hot experts for a 256-expert MoE model). Set --kt-max-deferred-experts-per-token to 2 or higher to enable pipelining, which allows CPU computation to overlap with GPU execution. Benchmark throughput with your specific hardware to find the optimal split.

Can I use existing LoRA adapters with this integration?

Yes, fused-expert LoRA checkpoints can be converted using kt-kernel/scripts/convert_kt_to_sglang_adapter.py. This script transforms the checkpoints into a directory structure compatible with SGLang's adapter loading mechanism, allowing you to serve fine-tuned models through the heterogeneous inference pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →