How to Set Up KTransformers on Windows Native Environment: Complete Installation Guide
Install CUDA 12.5, Python 3.12, and MSVC 2022 Build Tools, then deploy the pre-compiled Windows wheel or run install.bat to launch the heterogeneous LLM inference engine without WSL2.
KTransformers is a high-performance inference framework developed by the kvcache-ai/ktransformers repository that combines optimized C++/CUDA kernels (kt-kernel) with a lightweight Python wrapper. While the project maintains full Linux support, you can set up KTransformers on a native Windows environment using pre-compiled binaries that bundle the kernel, PyTorch, and CPU extensions (AVX2/AVX512) into a single wheel.
Prerequisites
Hardware and System Requirements
- Windows 10 or 11 (64-bit)
- NVIDIA GPU with recent driver support for CUDA 12.5
- CPU supporting AVX2 or AVX512 instruction sets
- At least 32 GB system RAM recommended for large model inference
Software Dependencies
| Component | Version | Installation Source |
|---|---|---|
| CUDA Toolkit | 12.5 | NVIDIA CUDA Archive |
| Python | 3.12 (64-bit) | python.org or Miniconda |
| MSVC Build Tools | 2022 v143 | Visual Studio Installer with "Desktop development with C++" workload |
| PyTorch | 2.4+cu125 | pip install torch==2.4.0+cu125 --extra-index-url https://download.pytorch.org/whl/cu125 |
Note: The official Windows wheel (ktransformers-0.2.0+cu125torch24avx2-cp312-cp312-win_amd64.whl) targets Python 3.12 (indicated by the cp312 ABI tag). Ensure your Python version matches the wheel.
Installation Methods
Method 1: Install the Pre-compiled Wheel (Recommended)
Download the Windows-specific wheel from the GitHub releases page:
curl -L -o ktransformers-0.2.0+cu125torch24avx2-cp312-cp312-win_amd64.whl ^
https://github.com/kvcache-ai/ktransformers/releases/download/v0.2.0/ktransformers-0.2.0+cu125torch24avx2-cp312-cp312-win_amd64.whl
Install the package directly:
pip install ktransformers-0.2.0+cu125torch24avx2-cp312-cp312-win_amd64.whl
Reference: Windows wheel URL documented in doc/en/install.md (line 89).
Method 2: Build from Source Using install.bat
Clone the repository and initialize submodules:
git clone https://github.com/kvcache-ai/ktransformers.git
cd ktransformers
git submodule update --init --recursive
Execute the Windows installer script located at archive/install.bat:
install.bat
This batch script performs three operations:
- Clears previous build artifacts (
rmdir /S /Q build dist) - Installs Python dependencies from
requirements-local_chat.txt - Installs KTransformers with
pip install . --no-build-isolation
Source: archive/install.bat – view source.
Verify the Installation
Confirm the package imports correctly and check the version:
python -c "import ktransformers; print(ktransformers.__version__)"
If this executes without ImportError or DLL load failures, the kt-kernel binaries are properly linked.
Running Inference
Local Chat Demo
Run the single-user command-line interface from ktransformers/local_chat.py:
python -m ktransformers.local_chat ^
--model_path deepseek-ai/DeepSeek-V2-Lite-Chat ^
--gguf_path .\DeepSeek-V2-Lite-Chat-GGUF
Source: Entry point defined in ktransformers/local_chat.py.
Server Mode (Multi-Concurrency)
Launch the HTTP server from ktransformers/server/main.py for multi-client serving:
python ktransformers/server/main.py ^
--model_path C:\models\DeepSeek-V3 ^
--gguf_path C:\models\DeepSeek-V3-GGUF ^
--cpu_infer 24 ^
--optimize_config_path ktransformers/optimize/optimize_rules/DeepSeek-V3-Chat-serve.yaml ^
--backend_type balance_serve ^
--port 10002
Source: Server implementation in ktransformers/server/main.py.
Architecture Overview
The Windows deployment packages three core components into the wheel:
kt-kernel– Native C++/CUDA binaries implementing CPU INT4/INT8 kernels (AVX2/AVX512) and GPU MoE routing. Documentation:kt-kernel/README.md.- Python Wrapper –
ktransformers/ktransformers.pyexposesKTransformerEngine, forwarding calls to the native library viactypes. - Optimization Rules – YAML configurations in
ktransformers/optimize/optimize_rules/define expert placement and quantization policies.
Because the wheel contains pre-compiled binaries, the Windows setup skips the CMake build process required on Linux.
Troubleshooting
- ImportError: DLL load failed – Ensure CUDA 12.5 is installed and PyTorch matches the
cu125build. The wheel is compiled against CUDA 12.5; other versions will cause library loading errors. - MSVC Build Tools required – Even when using the wheel, running
install.batmay trigger compilation steps. Install the "Desktop development with C++" workload via Visual Studio Installer. - Python version mismatch – The available wheel requires Python 3.12 (
cp312). Using Python 3.11 will result in a "not supported wheel on this platform" error.
Summary
- KTransformers on Windows relies on a pre-compiled wheel (
ktransformers-0.2.0+cu125torch24avx2-cp312-cp312-win_amd64.whl) that bundles the CUDA 12.5 kernel and Python 3.12 bindings. - Install via
pip install <wheel>.whlor runarchive/install.batafter cloning the repository. - Prerequisites include CUDA 12.5, Python 3.12, MSVC 2022 v143, and compatible PyTorch wheels.
- Launch inference using
ktransformers/local_chat.pyfor single-user mode orktransformers/server/main.pyfor HTTP serving. - Native Windows support is maintained via pre-compiled releases, though the project recommends WSL2 for development builds according to
doc/en/install.md.
Frequently Asked Questions
Is native Windows support deprecated for KTransformers?
According to doc/en/install.md in the kvcache-ai/ktransformers repository, native Windows support is temporarily deprecated in favor of WSL2 Ubuntu for development. However, the pre-compiled wheels released for Windows 10/11 remain functional and receive kernel updates.
Which Python version is required for the Windows wheel?
The current Windows wheel (ktransformers-0.2.0+cu125torch24avx2-cp312-cp312-win_amd64.whl) specifically requires Python 3.12, as indicated by the cp312 ABI tag in the filename. You must use Python 3.12 to install this wheel; Python 3.11 or earlier will reject the package as an incompatible wheel platform.
Can I run KTransformers on Windows without an NVIDIA GPU?
While you can install the package, the kt-kernel relies heavily on CUDA for heterogeneous compute and MoE routing. CPU-only inference is possible but not optimized for Windows native builds, as the AVX2/AVX512 kernels are designed to work alongside GPU offloading. Performance without a CUDA-capable GPU will be significantly degraded.
Where are the optimization rules stored for different model architectures?
Optimization configurations reside in ktransformers/optimize/optimize_rules/ as YAML files. When starting the server or local chat, specify the rule file using --optimize_config_path, such as DeepSeek-V3-Chat-serve.yaml for DeepSeek-V3 models or equivalent files for other architectures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →