# Deploying KTransformers in Production Environments: A Complete Hardware and Software Guide

> Deploy KTransformers in production environments with this hardware and software guide. Learn best practices for version matching, containerization, and reproducible inference.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: best-practices
- Published: 2026-07-20

---

**Deploy KTransformers in production by matching your Python/Torch/CUDA versions to the pre-compiled wheel, running a lightweight sanity check on a few prompts, and containerizing the service with pinned kernel versions to ensure reproducible CPU-GPU heterogeneous inference.**

Deploying KTransformers in production environments requires careful coordination between heterogeneous hardware resources and precise software versioning. The kvcache-ai/ktransformers library enables **CPU↔GPU heterogeneous inference** for Mixture-of-Experts (MoE) large language models, demanding specific Xeon CPU features and CUDA toolkit alignment. This guide covers the authoritative best practices derived from the source code documentation and kernel implementations.

## Hardware and Operating System Prerequisites

Production deployments require a foundation of modern Intel hardware and recent Ubuntu distributions to support the kernel's optimized inference paths.

According to [`doc/en/kt-kernel/experts-sched-Tutorial.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/experts-sched-Tutorial.md), the minimum specifications include:

- **Operating System**: Ubuntu 20.04 or newer (strongly recommended for driver compatibility)
- **CPU**: Xeon processor with AVX2, AVX-512, or AMX instruction sets. The physical core count must meet or exceed the number of *cold* experts retained on CPU
- **GPU**: NVIDIA GPU supporting CUDA 12.0 or newer (CUDA 12.8+ required for FP8/MXFP8 precision)
- **Drivers**: NVIDIA driver version 525 or newer with matching CUDA toolkit

The [`doc/en/kt-kernel/MiniMax-M3-Tutorial.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/MiniMax-M3-Tutorial.md) specifically emphasizes that AMX-enabled CPUs provide significant throughput gains for INT4/INT8 quantized expert computation.

## Installing the Correct KTransformers Wheel

KTransformers distributes pre-compiled wheels that bundle native kernels targeting specific architecture triples. Selecting the wrong wheel causes immediate ABI mismatches and kernel loading failures.

Install the wheel matching your **Python version, PyTorch version, and CUDA version** exactly:

```bash

# Example: Python 3.10, PyTorch 2.2, CUDA 12.0

pip install \
  https://github.com/kvcache-ai/ktransformers/releases/download/v0.6.1/ktransformers-0.6.1-cp310-cp310-linux_x86_64-cuda120.whl

```

The installation guide at [`doc/en/SFT_Installation_Guide_KimiK2.5.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/SFT_Installation_Guide_KimiK2.5.md) stresses that production systems must pin the exact wheel version to prevent runtime kernel registration errors.

## Pre-Deployment Sanity Checks

Before exposing the service to traffic, execute a lightweight verification to confirm the kernel loads correctly and the CPU/GPU expert routing functions properly.

Run a dry-run inference test:

```bash
kt serve \
  --model deepseek-v3.2 \
  --device cuda \
  --cpu-experts 8 \
  --gpu-experts 4 \
  --max-new-tokens 64 \
  --prompt "Hello, KTransformers!" \
  --dry-run

```

If the command streams token-by-token output without `CUDA_ERROR` or `CPU kernel not found` exceptions, the environment is production-ready. The [`doc/en/SFT_Installation_Guide_KimiK2.5.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/SFT_Installation_Guide_KimiK2.5.md) explicitly mandates this check before any production deployment.

## Deployment Architectures

KTransformers supports three primary deployment modes, each documented in separate tutorial files within the repository.

### Standalone KT CLI Mode

For single-node serving with minimal dependencies, use the native `kt` CLI. This mode provides direct control over expert placement and memory allocation:

```bash
kt serve \
  --model deepseek-v3.2 \
  --device cuda \
  --cpuinfer $(nproc --all) \
  --gpuint 4 \
  --kv-cache-dtype fp8_e4m3 \
  --port 8000 \
  --log-level info \
  --max-batch-size 32 \
  --max-new-tokens 256

```

All supported flags and default values are documented in [`kt-kernel/README.md`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md).

### SGLang Integration for Large MoE Models

For advanced request routing and heterogeneous expert scheduling, deploy via the SGLang integration. This approach handles hot experts on GPU and cold experts on CPU transparently:

```bash

# Install the SGLang fork containing KT integration

pip install sglang==0.2.0+ktransformers

# Launch with KT routing enabled

sglang server \
  --model deepseek-v3.2 \
  --kt-method AMXINT4 \
  --kt-num-gpu-experts 4 \
  --kt-cpuinfer $(nproc --all) \
  --port 8080

```

The [`doc/en/kt-kernel/deepseek-v3.2-sglang-tutorial.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/deepseek-v3.2-sglang-tutorial.md) provides the complete flag reference for hybrid deployment scenarios.

### Containerized Deployment

Create immutable Docker images to ensure the kernel binary, CUDA libraries, and Python environment remain synchronized across development and production.

Based on `docker/Dockerfile`:

```dockerfile
FROM nvidia/cuda:12.0.1-runtime-ubuntu20.04

RUN apt-get update && apt-get install -y \
    python3-pip git curl && rm -rf /var/lib/apt/lists/*

# Install exact wheel version

RUN pip install \
    https://github.com/kvcache-ai/ktransformers/releases/download/v0.6.1/ktransformers-0.6.1-cp310-cp310-linux_x86_64-cuda120.whl

COPY kt-prod.yaml /app/kt-prod.yaml

EXPOSE 8000
ENTRYPOINT ["kt", "serve", "--config", "/app/kt-prod.yaml"]

```

Build and deploy:

```bash
docker build -t kt-prod .
docker run -d --gpus all -p 8000:8000 kt-prod

```

## Production Configuration Management

Externalize settings via YAML configuration files to version-control deployment parameters separately from container images.

Example [`kt-prod.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-prod.yaml):

```yaml
model: deepseek-v3.2
device: cuda
cpuinfer: 32          # Cores for cold experts

gpuint: 4             # GPU experts count

kv_cache_dtype: fp8_e4m3
max_new_tokens: 256
max_batch_size: 32
log_level: info
port: 8000

```

Configuration options are fully documented in [`doc/en/kt-kernel/kt-cli.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/kt-cli.md) and the CLI help (`kt serve --help`).

## Monitoring and Observability

Production systems require visibility into heterogeneous resource utilization and request latency.

- **Logging**: Set `log_level: info` (or `debug` for troubleshooting). Logs write to stdout; forward to aggregators like Loki or ELK
- **Metrics**: Expose Prometheus metrics via `--metrics-port 9090` (available in KT v0.6+)
- **Health Checks**: Implement a `/healthz` endpoint that executes `kt health` for load balancer integration

## Scaling and Resource Optimization

Optimize throughput by aligning expert placement with hardware capabilities:

- **Horizontal Scaling**: Run multiple containers behind a load balancer, varying `cpuinfer` and `gpuint` per node to host different expert subsets
- **Memory Budgeting**: Use FP8 KV caching (`--kv-cache-dtype fp8_e4m3`) to reduce GPU memory footprint, enabling higher concurrency
- **Expert Scheduling**: Consult [`doc/en/kt-kernel/experts-sched-Tutorial.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/experts-sched-Tutorial.md) to determine optimal CPU-to-GPU expert ratios based on your specific hardware topology

## Security and Version Control

Maintain deployment integrity through strict version management:

- **Pin Exact Versions**: Track the Git SHA of the KT wheel and base Docker image in your CI/CD pipeline
- **Kernel Updates**: New optimizations (AMX, AVX-512VNNI) release with each minor version; monitor [`doc/en/AMX.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/AMX.md) for performance improvements
- **Workload Isolation**: Run KT in dedicated namespaces or VMs, as the kernel replaces FlashInfer in the environment and may conflict with other GPU workloads

## Summary

- **Match hardware exactly**: Deploy on Ubuntu 20.04+ with Xeon CPUs featuring AVX2/AVX512/AMX and CUDA 12.0+ (12.8+ for FP8)
- **Install precise wheels**: Select the ktransformers wheel matching your Python, PyTorch, and CUDA versions to avoid ABI failures
- **Validate before serving**: Execute a lightweight sanity check with `--dry-run` to verify kernel loading and expert routing
- **Choose appropriate architecture**: Use standalone `kt serve` for simple deployments or SGLang integration for complex MoE routing
- **Containerize immutably**: Build production images using `docker/Dockerfile` with pinned versions and externalized YAML configs
- **Monitor continuously**: Enable Prometheus metrics and structured logging for observable heterogeneous inference

## Frequently Asked Questions

### What hardware specifications are required for deploying KTransformers in production environments?

Production deployments require a Xeon CPU with AVX2, AVX-512, or AMX instruction sets and an NVIDIA GPU supporting CUDA 12.0 or newer. The CPU must have at least as many physical cores as the number of cold experts retained on CPU, while CUDA 12.8 or newer is mandatory for FP8/MXFP8 precision modes.

### How do I verify that my KTransformers installation is ready for production traffic?

Run the lightweight sanity check using `kt serve --dry-run` with a short prompt and limited token generation. This validates that the native kernel loads without CUDA errors and that the CPU-GPU expert routing functions correctly before exposing the service to real traffic, as recommended in [`doc/en/SFT_Installation_Guide_KimiK2.5.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/SFT_Installation_Guide_KimiK2.5.md).

### Should I use the standalone KT CLI or SGLang integration for production?

Use the **standalone `kt` CLI** for simple single-node deployments where you need direct control over expert placement. Choose **SGLang integration** when serving large MoE models requiring sophisticated request routing and automatic management of hot GPU experts versus cold CPU experts.

### How do I handle version upgrades in production KTransformers deployments?

Pin the exact wheel version, Python version, and CUDA toolkit in your Dockerfile, tracking the Git SHA in your CI/CD pipeline. When upgrading, consult [`doc/en/AMX.md`](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/AMX.md) for new kernel optimizations and test in a staging environment first, as kernel binaries are tightly coupled to specific PyTorch and CUDA ABI versions.