Deploying KTransformers in Production Environments: A Complete Hardware and Software Guide

Deploy KTransformers in production by matching your Python/Torch/CUDA versions to the pre-compiled wheel, running a lightweight sanity check on a few prompts, and containerizing the service with pinned kernel versions to ensure reproducible CPU-GPU heterogeneous inference.

Deploying KTransformers in production environments requires careful coordination between heterogeneous hardware resources and precise software versioning. The kvcache-ai/ktransformers library enables CPU↔GPU heterogeneous inference for Mixture-of-Experts (MoE) large language models, demanding specific Xeon CPU features and CUDA toolkit alignment. This guide covers the authoritative best practices derived from the source code documentation and kernel implementations.

Hardware and Operating System Prerequisites

Production deployments require a foundation of modern Intel hardware and recent Ubuntu distributions to support the kernel's optimized inference paths.

According to doc/en/kt-kernel/experts-sched-Tutorial.md, the minimum specifications include:

  • Operating System: Ubuntu 20.04 or newer (strongly recommended for driver compatibility)
  • CPU: Xeon processor with AVX2, AVX-512, or AMX instruction sets. The physical core count must meet or exceed the number of cold experts retained on CPU
  • GPU: NVIDIA GPU supporting CUDA 12.0 or newer (CUDA 12.8+ required for FP8/MXFP8 precision)
  • Drivers: NVIDIA driver version 525 or newer with matching CUDA toolkit

The doc/en/kt-kernel/MiniMax-M3-Tutorial.md specifically emphasizes that AMX-enabled CPUs provide significant throughput gains for INT4/INT8 quantized expert computation.

Installing the Correct KTransformers Wheel

KTransformers distributes pre-compiled wheels that bundle native kernels targeting specific architecture triples. Selecting the wrong wheel causes immediate ABI mismatches and kernel loading failures.

Install the wheel matching your Python version, PyTorch version, and CUDA version exactly:


# Example: Python 3.10, PyTorch 2.2, CUDA 12.0

pip install \
  https://github.com/kvcache-ai/ktransformers/releases/download/v0.6.1/ktransformers-0.6.1-cp310-cp310-linux_x86_64-cuda120.whl

The installation guide at doc/en/SFT_Installation_Guide_KimiK2.5.md stresses that production systems must pin the exact wheel version to prevent runtime kernel registration errors.

Pre-Deployment Sanity Checks

Before exposing the service to traffic, execute a lightweight verification to confirm the kernel loads correctly and the CPU/GPU expert routing functions properly.

Run a dry-run inference test:

kt serve \
  --model deepseek-v3.2 \
  --device cuda \
  --cpu-experts 8 \
  --gpu-experts 4 \
  --max-new-tokens 64 \
  --prompt "Hello, KTransformers!" \
  --dry-run

If the command streams token-by-token output without CUDA_ERROR or CPU kernel not found exceptions, the environment is production-ready. The doc/en/SFT_Installation_Guide_KimiK2.5.md explicitly mandates this check before any production deployment.

Deployment Architectures

KTransformers supports three primary deployment modes, each documented in separate tutorial files within the repository.

Standalone KT CLI Mode

For single-node serving with minimal dependencies, use the native kt CLI. This mode provides direct control over expert placement and memory allocation:

kt serve \
  --model deepseek-v3.2 \
  --device cuda \
  --cpuinfer $(nproc --all) \
  --gpuint 4 \
  --kv-cache-dtype fp8_e4m3 \
  --port 8000 \
  --log-level info \
  --max-batch-size 32 \
  --max-new-tokens 256

All supported flags and default values are documented in kt-kernel/README.md.

SGLang Integration for Large MoE Models

For advanced request routing and heterogeneous expert scheduling, deploy via the SGLang integration. This approach handles hot experts on GPU and cold experts on CPU transparently:


# Install the SGLang fork containing KT integration

pip install sglang==0.2.0+ktransformers

# Launch with KT routing enabled

sglang server \
  --model deepseek-v3.2 \
  --kt-method AMXINT4 \
  --kt-num-gpu-experts 4 \
  --kt-cpuinfer $(nproc --all) \
  --port 8080

The doc/en/kt-kernel/deepseek-v3.2-sglang-tutorial.md provides the complete flag reference for hybrid deployment scenarios.

Containerized Deployment

Create immutable Docker images to ensure the kernel binary, CUDA libraries, and Python environment remain synchronized across development and production.

Based on docker/Dockerfile:

FROM nvidia/cuda:12.0.1-runtime-ubuntu20.04

RUN apt-get update && apt-get install -y \
    python3-pip git curl && rm -rf /var/lib/apt/lists/*

# Install exact wheel version

RUN pip install \
    https://github.com/kvcache-ai/ktransformers/releases/download/v0.6.1/ktransformers-0.6.1-cp310-cp310-linux_x86_64-cuda120.whl

COPY kt-prod.yaml /app/kt-prod.yaml

EXPOSE 8000
ENTRYPOINT ["kt", "serve", "--config", "/app/kt-prod.yaml"]

Build and deploy:

docker build -t kt-prod .
docker run -d --gpus all -p 8000:8000 kt-prod

Production Configuration Management

Externalize settings via YAML configuration files to version-control deployment parameters separately from container images.

Example kt-prod.yaml:

model: deepseek-v3.2
device: cuda
cpuinfer: 32          # Cores for cold experts

gpuint: 4             # GPU experts count

kv_cache_dtype: fp8_e4m3
max_new_tokens: 256
max_batch_size: 32
log_level: info
port: 8000

Configuration options are fully documented in doc/en/kt-kernel/kt-cli.md and the CLI help (kt serve --help).

Monitoring and Observability

Production systems require visibility into heterogeneous resource utilization and request latency.

  • Logging: Set log_level: info (or debug for troubleshooting). Logs write to stdout; forward to aggregators like Loki or ELK
  • Metrics: Expose Prometheus metrics via --metrics-port 9090 (available in KT v0.6+)
  • Health Checks: Implement a /healthz endpoint that executes kt health for load balancer integration

Scaling and Resource Optimization

Optimize throughput by aligning expert placement with hardware capabilities:

  • Horizontal Scaling: Run multiple containers behind a load balancer, varying cpuinfer and gpuint per node to host different expert subsets
  • Memory Budgeting: Use FP8 KV caching (--kv-cache-dtype fp8_e4m3) to reduce GPU memory footprint, enabling higher concurrency
  • Expert Scheduling: Consult doc/en/kt-kernel/experts-sched-Tutorial.md to determine optimal CPU-to-GPU expert ratios based on your specific hardware topology

Security and Version Control

Maintain deployment integrity through strict version management:

  • Pin Exact Versions: Track the Git SHA of the KT wheel and base Docker image in your CI/CD pipeline
  • Kernel Updates: New optimizations (AMX, AVX-512VNNI) release with each minor version; monitor doc/en/AMX.md for performance improvements
  • Workload Isolation: Run KT in dedicated namespaces or VMs, as the kernel replaces FlashInfer in the environment and may conflict with other GPU workloads

Summary

  • Match hardware exactly: Deploy on Ubuntu 20.04+ with Xeon CPUs featuring AVX2/AVX512/AMX and CUDA 12.0+ (12.8+ for FP8)
  • Install precise wheels: Select the ktransformers wheel matching your Python, PyTorch, and CUDA versions to avoid ABI failures
  • Validate before serving: Execute a lightweight sanity check with --dry-run to verify kernel loading and expert routing
  • Choose appropriate architecture: Use standalone kt serve for simple deployments or SGLang integration for complex MoE routing
  • Containerize immutably: Build production images using docker/Dockerfile with pinned versions and externalized YAML configs
  • Monitor continuously: Enable Prometheus metrics and structured logging for observable heterogeneous inference

Frequently Asked Questions

What hardware specifications are required for deploying KTransformers in production environments?

Production deployments require a Xeon CPU with AVX2, AVX-512, or AMX instruction sets and an NVIDIA GPU supporting CUDA 12.0 or newer. The CPU must have at least as many physical cores as the number of cold experts retained on CPU, while CUDA 12.8 or newer is mandatory for FP8/MXFP8 precision modes.

How do I verify that my KTransformers installation is ready for production traffic?

Run the lightweight sanity check using kt serve --dry-run with a short prompt and limited token generation. This validates that the native kernel loads without CUDA errors and that the CPU-GPU expert routing functions correctly before exposing the service to real traffic, as recommended in doc/en/SFT_Installation_Guide_KimiK2.5.md.

Should I use the standalone KT CLI or SGLang integration for production?

Use the standalone kt CLI for simple single-node deployments where you need direct control over expert placement. Choose SGLang integration when serving large MoE models requiring sophisticated request routing and automatic management of hot GPU experts versus cold CPU experts.

How do I handle version upgrades in production KTransformers deployments?

Pin the exact wheel version, Python version, and CUDA toolkit in your Dockerfile, tracking the Git SHA in your CI/CD pipeline. When upgrading, consult doc/en/AMX.md for new kernel optimizations and test in a staging environment first, as kernel binaries are tightly coupled to specific PyTorch and CUDA ABI versions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →