Deploying KTransformers in Production Environments: A Complete Hardware and Software Guide
Deploy KTransformers in production by matching your Python/Torch/CUDA versions to the pre-compiled wheel, running a lightweight sanity check on a few prompts, and containerizing the service with pinned kernel versions to ensure reproducible CPU-GPU heterogeneous inference.
Deploying KTransformers in production environments requires careful coordination between heterogeneous hardware resources and precise software versioning. The kvcache-ai/ktransformers library enables CPU↔GPU heterogeneous inference for Mixture-of-Experts (MoE) large language models, demanding specific Xeon CPU features and CUDA toolkit alignment. This guide covers the authoritative best practices derived from the source code documentation and kernel implementations.
Hardware and Operating System Prerequisites
Production deployments require a foundation of modern Intel hardware and recent Ubuntu distributions to support the kernel's optimized inference paths.
According to doc/en/kt-kernel/experts-sched-Tutorial.md, the minimum specifications include:
- Operating System: Ubuntu 20.04 or newer (strongly recommended for driver compatibility)
- CPU: Xeon processor with AVX2, AVX-512, or AMX instruction sets. The physical core count must meet or exceed the number of cold experts retained on CPU
- GPU: NVIDIA GPU supporting CUDA 12.0 or newer (CUDA 12.8+ required for FP8/MXFP8 precision)
- Drivers: NVIDIA driver version 525 or newer with matching CUDA toolkit
The doc/en/kt-kernel/MiniMax-M3-Tutorial.md specifically emphasizes that AMX-enabled CPUs provide significant throughput gains for INT4/INT8 quantized expert computation.
Installing the Correct KTransformers Wheel
KTransformers distributes pre-compiled wheels that bundle native kernels targeting specific architecture triples. Selecting the wrong wheel causes immediate ABI mismatches and kernel loading failures.
Install the wheel matching your Python version, PyTorch version, and CUDA version exactly:
# Example: Python 3.10, PyTorch 2.2, CUDA 12.0
pip install \
https://github.com/kvcache-ai/ktransformers/releases/download/v0.6.1/ktransformers-0.6.1-cp310-cp310-linux_x86_64-cuda120.whl
The installation guide at doc/en/SFT_Installation_Guide_KimiK2.5.md stresses that production systems must pin the exact wheel version to prevent runtime kernel registration errors.
Pre-Deployment Sanity Checks
Before exposing the service to traffic, execute a lightweight verification to confirm the kernel loads correctly and the CPU/GPU expert routing functions properly.
Run a dry-run inference test:
kt serve \
--model deepseek-v3.2 \
--device cuda \
--cpu-experts 8 \
--gpu-experts 4 \
--max-new-tokens 64 \
--prompt "Hello, KTransformers!" \
--dry-run
If the command streams token-by-token output without CUDA_ERROR or CPU kernel not found exceptions, the environment is production-ready. The doc/en/SFT_Installation_Guide_KimiK2.5.md explicitly mandates this check before any production deployment.
Deployment Architectures
KTransformers supports three primary deployment modes, each documented in separate tutorial files within the repository.
Standalone KT CLI Mode
For single-node serving with minimal dependencies, use the native kt CLI. This mode provides direct control over expert placement and memory allocation:
kt serve \
--model deepseek-v3.2 \
--device cuda \
--cpuinfer $(nproc --all) \
--gpuint 4 \
--kv-cache-dtype fp8_e4m3 \
--port 8000 \
--log-level info \
--max-batch-size 32 \
--max-new-tokens 256
All supported flags and default values are documented in kt-kernel/README.md.
SGLang Integration for Large MoE Models
For advanced request routing and heterogeneous expert scheduling, deploy via the SGLang integration. This approach handles hot experts on GPU and cold experts on CPU transparently:
# Install the SGLang fork containing KT integration
pip install sglang==0.2.0+ktransformers
# Launch with KT routing enabled
sglang server \
--model deepseek-v3.2 \
--kt-method AMXINT4 \
--kt-num-gpu-experts 4 \
--kt-cpuinfer $(nproc --all) \
--port 8080
The doc/en/kt-kernel/deepseek-v3.2-sglang-tutorial.md provides the complete flag reference for hybrid deployment scenarios.
Containerized Deployment
Create immutable Docker images to ensure the kernel binary, CUDA libraries, and Python environment remain synchronized across development and production.
Based on docker/Dockerfile:
FROM nvidia/cuda:12.0.1-runtime-ubuntu20.04
RUN apt-get update && apt-get install -y \
python3-pip git curl && rm -rf /var/lib/apt/lists/*
# Install exact wheel version
RUN pip install \
https://github.com/kvcache-ai/ktransformers/releases/download/v0.6.1/ktransformers-0.6.1-cp310-cp310-linux_x86_64-cuda120.whl
COPY kt-prod.yaml /app/kt-prod.yaml
EXPOSE 8000
ENTRYPOINT ["kt", "serve", "--config", "/app/kt-prod.yaml"]
Build and deploy:
docker build -t kt-prod .
docker run -d --gpus all -p 8000:8000 kt-prod
Production Configuration Management
Externalize settings via YAML configuration files to version-control deployment parameters separately from container images.
Example kt-prod.yaml:
model: deepseek-v3.2
device: cuda
cpuinfer: 32 # Cores for cold experts
gpuint: 4 # GPU experts count
kv_cache_dtype: fp8_e4m3
max_new_tokens: 256
max_batch_size: 32
log_level: info
port: 8000
Configuration options are fully documented in doc/en/kt-kernel/kt-cli.md and the CLI help (kt serve --help).
Monitoring and Observability
Production systems require visibility into heterogeneous resource utilization and request latency.
- Logging: Set
log_level: info(ordebugfor troubleshooting). Logs write to stdout; forward to aggregators like Loki or ELK - Metrics: Expose Prometheus metrics via
--metrics-port 9090(available in KT v0.6+) - Health Checks: Implement a
/healthzendpoint that executeskt healthfor load balancer integration
Scaling and Resource Optimization
Optimize throughput by aligning expert placement with hardware capabilities:
- Horizontal Scaling: Run multiple containers behind a load balancer, varying
cpuinferandgpuintper node to host different expert subsets - Memory Budgeting: Use FP8 KV caching (
--kv-cache-dtype fp8_e4m3) to reduce GPU memory footprint, enabling higher concurrency - Expert Scheduling: Consult
doc/en/kt-kernel/experts-sched-Tutorial.mdto determine optimal CPU-to-GPU expert ratios based on your specific hardware topology
Security and Version Control
Maintain deployment integrity through strict version management:
- Pin Exact Versions: Track the Git SHA of the KT wheel and base Docker image in your CI/CD pipeline
- Kernel Updates: New optimizations (AMX, AVX-512VNNI) release with each minor version; monitor
doc/en/AMX.mdfor performance improvements - Workload Isolation: Run KT in dedicated namespaces or VMs, as the kernel replaces FlashInfer in the environment and may conflict with other GPU workloads
Summary
- Match hardware exactly: Deploy on Ubuntu 20.04+ with Xeon CPUs featuring AVX2/AVX512/AMX and CUDA 12.0+ (12.8+ for FP8)
- Install precise wheels: Select the ktransformers wheel matching your Python, PyTorch, and CUDA versions to avoid ABI failures
- Validate before serving: Execute a lightweight sanity check with
--dry-runto verify kernel loading and expert routing - Choose appropriate architecture: Use standalone
kt servefor simple deployments or SGLang integration for complex MoE routing - Containerize immutably: Build production images using
docker/Dockerfilewith pinned versions and externalized YAML configs - Monitor continuously: Enable Prometheus metrics and structured logging for observable heterogeneous inference
Frequently Asked Questions
What hardware specifications are required for deploying KTransformers in production environments?
Production deployments require a Xeon CPU with AVX2, AVX-512, or AMX instruction sets and an NVIDIA GPU supporting CUDA 12.0 or newer. The CPU must have at least as many physical cores as the number of cold experts retained on CPU, while CUDA 12.8 or newer is mandatory for FP8/MXFP8 precision modes.
How do I verify that my KTransformers installation is ready for production traffic?
Run the lightweight sanity check using kt serve --dry-run with a short prompt and limited token generation. This validates that the native kernel loads without CUDA errors and that the CPU-GPU expert routing functions correctly before exposing the service to real traffic, as recommended in doc/en/SFT_Installation_Guide_KimiK2.5.md.
Should I use the standalone KT CLI or SGLang integration for production?
Use the standalone kt CLI for simple single-node deployments where you need direct control over expert placement. Choose SGLang integration when serving large MoE models requiring sophisticated request routing and automatic management of hot GPU experts versus cold CPU experts.
How do I handle version upgrades in production KTransformers deployments?
Pin the exact wheel version, Python version, and CUDA toolkit in your Dockerfile, tracking the Git SHA in your CI/CD pipeline. When upgrading, consult doc/en/AMX.md for new kernel optimizations and test in a staging environment first, as kernel binaries are tightly coupled to specific PyTorch and CUDA ABI versions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →