How to Set Up KTransformers with Ascend NPU for Inference: A Complete Guide
KTransformers enables high-performance MoE inference on Huawei Ascend NPUs by combining the balance_serve backend with custom NPU operators located in ktransformers/operators/ascend, requiring CANN 8.3+, Docker deployment, and specific ARM64 build patches.
KTransformers is an open-source inference engine optimized for large-scale Mixture-of-Experts (MoE) models like DeepSeek-V3/R1. This guide explains how to set up KTransformers with Ascend NPU for inference, leveraging the balance_serve backend to schedule experts across NPU and CPU workers while minimizing latency through persistent NPU graphs.
Prerequisites and Architecture Overview
Running KTransformers on Ascend NPUs requires a specific software stack and hardware configuration. According to the DeepseekR1_V3_tutorial_zh_for_Ascend_NPU.md documentation, the architecture consists of:
- Ascend NPU Hardware: Atlas 300I A2 server or compatible Ascend NPU supported by CANN 8.3.xx
- CANN Toolkit: Ascend driver, HDK, and NNAL libraries (
/usr/local/Ascend/ascend-toolkit/set_env.sh) - Docker Environment: Pre-built Ubuntu 22.04 aarch64 image (
mindie:2.2.RC1-800I-A2-py311-openeuler24.03-lts) - PyTorch for NPU:
torch-npupackage for dispatching tensors to the Ascend device
The inference workflow splits MoE computation between the NPU (running optimized kernels in ascend_experts.py) and CPU workers, orchestrated by the balance_serve backend defined in ktransformers/server/utils/create_interface.py.
Step-by-Step Installation
1. Launch the Docker Container
Mount the Ascend drivers and toolkit into an interactive container:
docker run -it -d --net=host --shm-size=500g \
--name kt-ascend \
-w /workspace \
--device=/dev/davinci_manager \
--device=/dev/hisi_hdc \
--device=/dev/devmm_svm \
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver:ro \
-v /usr/local/dcmi:/usr/local/dcmi:ro \
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi:ro \
-v /usr/local/sbin/:/usr/local/sbin:ro \
-v <path_to_your_project>:/workspace \
mindie:2.2.RC1-800I-A2-py311-openeuler24.03-lts bash
2. Install Dependencies and torch-npu
Inside the container, install system libraries and Python packages:
# Install system dependencies
yum install -y zlib-devel libtbb-devel openssl-devel libaio-devel \
libcurl-devel
# Install PyTorch CPU wheel first, then torch-npu dependencies
pip3 install numpy==1.26.4 \
torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 \
--index-url https://download.pytorch.org/whl/cpu
pip3 install packaging ninja fire protobuf attrs \
decorator cloudpickle ml-dtypes scipy tornado absl-py psutil \
sqlalchemy transformers==4.57.1
3. Clone and Patch for ARM64
Ascend NPU servers run on ARM64 architecture, requiring a patch to disable x86-specific IQK kernels:
git clone https://github.com/kvcache-ai/ktransformers.git
cd ktransformers
git submodule update --init --recursive
# Comment out ARM-specific SIMD kernels (not needed for Ascend NPU)
sed -i 's/#define iqk_mul_mat iqk_mul_mat_arm82/\/\/#define iqk_mul_mat iqk_mul_mat_arm82/' \
./third_party/llamafile/iqk_mul_mat_arm82.cpp
4. Build with NPU Support
Compile the NPU-optimized operators using the installation script:
USE_BALANCE_SERVE=1 USE_NUMA=1 bash ./install.sh
This builds the custom operators in ktransformers/operators/ascend/ascend_experts.py, including KExpertsCPUW8A8 and KTransformersExpertsW8A8, and enables the NUMA-aware memory allocator for optimal CPU-NPU data movement.
Model Preparation
For the DeepSeek-R1/R3 models, merge the Q4 and W8A8 GGUF weights using the provided utility:
python merge_safetensor_gguf.py \
--safetensor_path /mnt/weights/DeepSeek-R1-Q4_K_M \
--gguf_path /mnt/weights/DeepSeek-R1-W8A8 \
--output_path /mnt/weights/DeepSeek-R1-q4km-w8a8
This creates a unified weight directory compatible with the NPU balance_serve backend.
Configuration and Runtime Settings
Configure the inference parameters in ktransformers/config/config.yaml or via command-line arguments:
- Backend: Set
backend_type: balance_serveto enable multi-concurrency NPU scheduling - Device: Specify
device: nputo trigger the Ascend code path - Memory: Adjust
page_size: 128andchunk_size: 16384for NPU-optimized KV cache management
The optimization rules in ktransformers/optimize/optimize_rules/npu/DeepSeek-V3-Chat-300IA2-npu-serve.yaml define kernel fusion patterns specific to the Atlas 300I A2 NPU.
Launching the Inference Server
Set the required environment variables and start the server:
export USE_MERGE=0
export INF_NAN_MODE_FORCE_DISABLE=1
export TASK_QUEUE_ENABLE=0
export RANK=0
export LOCAL_WORLD_SIZE=1
# Source Ascend toolkit
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
python ktransformers/server/main.py \
--port 10002 \
--model_path /mnt/weights/DeepSeek-R1-q4km-w8a8 \
--gguf_path /mnt/weights/DeepSeek-R1-q4km-w8a8 \
--model_name DeepSeekV3ForCausalLM \
--cpu_infer 100 \
--optimize_config_path ./ktransformers/optimize/optimize_rules/npu/DeepSeek-V3-Chat-300IA2-npu-serve.yaml \
--max_new_tokens 1024 \
--cache_lens 20480 \
--max_batch_size 4 \
--use_cuda_graph \
--tp 1 \
--backend_type balance_serve
Key flags explained:
--backend_type balance_serve: Activates the multi-concurrency engine that schedules MoE experts across NPU and CPU--use_cuda_graph: Enables the NPU graph runner (ktransformers/util/npu_graph_runner.py) to reduce kernel launch overhead by capturing and replaying NPU operations--cpu_infer 100: Allocates 100 CPU threads for non-expert operations and data preprocessing--optimize_config_path: Points to the YAML file containing NPU-specific kernel fusion rules
Sending Inference Requests
Once the server is running on port 10002, send requests using the HTTP API:
import requests
import json
url = "http://localhost:10002/generate"
payload = {
"prompt": "Explain the advantages of Ascend NPU for LLM inference.",
"max_new_tokens": 256,
"temperature": 0.7
}
response = requests.post(url, json=payload)
result = json.loads(response.text)
print(result["generated_text"])
The server executes MoE routing on the Ascend NPU via the custom operators in ascend_experts.py, while the npu_graph_runner.py module manages persistent graphs to minimize dispatch latency for repeated inference calls.
Summary
- Hardware Requirements: Atlas 300I A2 server with CANN 8.3+ toolkit and NNAL libraries installed on the host
- Docker Deployment: Use the official
mindie:2.2.RC1image to obtain a pre-configured Ubuntu 22.04 aarch64 environment - Build Process: Patch
iqk_mul_mat_arm82.cppfor ARM compatibility, then compile withUSE_BALANCE_SERVE=1 USE_NUMA=1 - Model Setup: Merge Q4 and W8A8 weights using
merge_safetensor_gguf.pybefore inference - Runtime Configuration: Use
--backend_type balance_servewith--use_cuda_graphto enable the optimized NPU execution path - Key Files:
ascend_experts.py(NPU kernels),npu_graph_runner.py(graph optimization), andmain.py(server entry point)
Frequently Asked Questions
What Ascend NPU hardware is compatible with KTransformers?
KTransformers officially supports the Atlas 300I A2 server running CANN 8.3.xx. The implementation relies on specific kernel optimizations found in ktransformers/operators/ascend/ascend_experts.py that target the Ascend 310P/910B architecture. Ensure your driver version matches the CANN toolkit requirements documented in DeepseekR1_V3_tutorial_zh_for_Ascend_NPU.md.
Why is the ARM patch necessary when building for Ascend NPU?
Ascend NPU servers typically run on ARM64 (aarch64) architecture. The IQK (integer quantization kernel) implementations in third_party/llamafile/iqk_mul_mat_arm82.cpp contain ARM-specific SIMD optimizations that conflict with the Ascend NPU's own kernel dispatch mechanism. Commenting out the #define iqk_mul_mat iqk_mul_mat_arm82 line prevents symbol collisions while allowing the NPU-specific operators to handle matrix multiplication.
How does the balance_serve backend differ from standard CUDA inference?
The balance_serve backend, implemented in the server utilities, specifically partitions MoE workloads between CPU and NPU workers. Unlike standard CUDA backends that run entire layers on GPU, balance_serve in ktransformers/server/utils/create_interface.py schedules individual experts to the Ascend NPU via ascend_experts.py while keeping attention heads and embeddings on CPU, maximizing throughput for memory-bandwidth-bound MoE models.
What is the purpose of the NPU graph runner?
The npu_graph_runner.py utility creates persistent execution graphs that cache the NPU kernel launch sequence. When --use_cuda_graph is enabled, the system captures the MoE forward pass into a static graph during the first iteration, then replays it for subsequent tokens. This eliminates Python-level overhead and CPU-NPU synchronization delays, reducing per-token latency by 30-50% compared to eager execution mode.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →