How to Set Up KTransformers with Ascend NPU for Inference: A Complete Guide

KTransformers enables high-performance MoE inference on Huawei Ascend NPUs by combining the balance_serve backend with custom NPU operators located in ktransformers/operators/ascend, requiring CANN 8.3+, Docker deployment, and specific ARM64 build patches.

KTransformers is an open-source inference engine optimized for large-scale Mixture-of-Experts (MoE) models like DeepSeek-V3/R1. This guide explains how to set up KTransformers with Ascend NPU for inference, leveraging the balance_serve backend to schedule experts across NPU and CPU workers while minimizing latency through persistent NPU graphs.

Prerequisites and Architecture Overview

Running KTransformers on Ascend NPUs requires a specific software stack and hardware configuration. According to the DeepseekR1_V3_tutorial_zh_for_Ascend_NPU.md documentation, the architecture consists of:

  • Ascend NPU Hardware: Atlas 300I A2 server or compatible Ascend NPU supported by CANN 8.3.xx
  • CANN Toolkit: Ascend driver, HDK, and NNAL libraries (/usr/local/Ascend/ascend-toolkit/set_env.sh)
  • Docker Environment: Pre-built Ubuntu 22.04 aarch64 image (mindie:2.2.RC1-800I-A2-py311-openeuler24.03-lts)
  • PyTorch for NPU: torch-npu package for dispatching tensors to the Ascend device

The inference workflow splits MoE computation between the NPU (running optimized kernels in ascend_experts.py) and CPU workers, orchestrated by the balance_serve backend defined in ktransformers/server/utils/create_interface.py.

Step-by-Step Installation

1. Launch the Docker Container

Mount the Ascend drivers and toolkit into an interactive container:

docker run -it -d --net=host --shm-size=500g \
   --name kt-ascend \
   -w /workspace \
   --device=/dev/davinci_manager \
   --device=/dev/hisi_hdc \
   --device=/dev/devmm_svm \
   -v /usr/local/Ascend/driver:/usr/local/Ascend/driver:ro \
   -v /usr/local/dcmi:/usr/local/dcmi:ro \
   -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi:ro \
   -v /usr/local/sbin/:/usr/local/sbin:ro \
   -v <path_to_your_project>:/workspace \
   mindie:2.2.RC1-800I-A2-py311-openeuler24.03-lts bash

2. Install Dependencies and torch-npu

Inside the container, install system libraries and Python packages:


# Install system dependencies

yum install -y zlib-devel libtbb-devel openssl-devel libaio-devel \
    libcurl-devel

# Install PyTorch CPU wheel first, then torch-npu dependencies

pip3 install numpy==1.26.4 \
    torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 \
    --index-url https://download.pytorch.org/whl/cpu

pip3 install packaging ninja fire protobuf attrs \
    decorator cloudpickle ml-dtypes scipy tornado absl-py psutil \
    sqlalchemy transformers==4.57.1

3. Clone and Patch for ARM64

Ascend NPU servers run on ARM64 architecture, requiring a patch to disable x86-specific IQK kernels:

git clone https://github.com/kvcache-ai/ktransformers.git
cd ktransformers
git submodule update --init --recursive

# Comment out ARM-specific SIMD kernels (not needed for Ascend NPU)

sed -i 's/#define iqk_mul_mat iqk_mul_mat_arm82/\/\/#define iqk_mul_mat iqk_mul_mat_arm82/' \
    ./third_party/llamafile/iqk_mul_mat_arm82.cpp

4. Build with NPU Support

Compile the NPU-optimized operators using the installation script:

USE_BALANCE_SERVE=1 USE_NUMA=1 bash ./install.sh

This builds the custom operators in ktransformers/operators/ascend/ascend_experts.py, including KExpertsCPUW8A8 and KTransformersExpertsW8A8, and enables the NUMA-aware memory allocator for optimal CPU-NPU data movement.

Model Preparation

For the DeepSeek-R1/R3 models, merge the Q4 and W8A8 GGUF weights using the provided utility:

python merge_safetensor_gguf.py \
    --safetensor_path /mnt/weights/DeepSeek-R1-Q4_K_M \
    --gguf_path /mnt/weights/DeepSeek-R1-W8A8 \
    --output_path /mnt/weights/DeepSeek-R1-q4km-w8a8

This creates a unified weight directory compatible with the NPU balance_serve backend.

Configuration and Runtime Settings

Configure the inference parameters in ktransformers/config/config.yaml or via command-line arguments:

  • Backend: Set backend_type: balance_serve to enable multi-concurrency NPU scheduling
  • Device: Specify device: npu to trigger the Ascend code path
  • Memory: Adjust page_size: 128 and chunk_size: 16384 for NPU-optimized KV cache management

The optimization rules in ktransformers/optimize/optimize_rules/npu/DeepSeek-V3-Chat-300IA2-npu-serve.yaml define kernel fusion patterns specific to the Atlas 300I A2 NPU.

Launching the Inference Server

Set the required environment variables and start the server:

export USE_MERGE=0
export INF_NAN_MODE_FORCE_DISABLE=1
export TASK_QUEUE_ENABLE=0
export RANK=0
export LOCAL_WORLD_SIZE=1

# Source Ascend toolkit

source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh

python ktransformers/server/main.py \
    --port 10002 \
    --model_path /mnt/weights/DeepSeek-R1-q4km-w8a8 \
    --gguf_path /mnt/weights/DeepSeek-R1-q4km-w8a8 \
    --model_name DeepSeekV3ForCausalLM \
    --cpu_infer 100 \
    --optimize_config_path ./ktransformers/optimize/optimize_rules/npu/DeepSeek-V3-Chat-300IA2-npu-serve.yaml \
    --max_new_tokens 1024 \
    --cache_lens 20480 \
    --max_batch_size 4 \
    --use_cuda_graph \
    --tp 1 \
    --backend_type balance_serve

Key flags explained:

  • --backend_type balance_serve: Activates the multi-concurrency engine that schedules MoE experts across NPU and CPU
  • --use_cuda_graph: Enables the NPU graph runner (ktransformers/util/npu_graph_runner.py) to reduce kernel launch overhead by capturing and replaying NPU operations
  • --cpu_infer 100: Allocates 100 CPU threads for non-expert operations and data preprocessing
  • --optimize_config_path: Points to the YAML file containing NPU-specific kernel fusion rules

Sending Inference Requests

Once the server is running on port 10002, send requests using the HTTP API:

import requests
import json

url = "http://localhost:10002/generate"
payload = {
    "prompt": "Explain the advantages of Ascend NPU for LLM inference.",
    "max_new_tokens": 256,
    "temperature": 0.7
}

response = requests.post(url, json=payload)
result = json.loads(response.text)
print(result["generated_text"])

The server executes MoE routing on the Ascend NPU via the custom operators in ascend_experts.py, while the npu_graph_runner.py module manages persistent graphs to minimize dispatch latency for repeated inference calls.

Summary

  • Hardware Requirements: Atlas 300I A2 server with CANN 8.3+ toolkit and NNAL libraries installed on the host
  • Docker Deployment: Use the official mindie:2.2.RC1 image to obtain a pre-configured Ubuntu 22.04 aarch64 environment
  • Build Process: Patch iqk_mul_mat_arm82.cpp for ARM compatibility, then compile with USE_BALANCE_SERVE=1 USE_NUMA=1
  • Model Setup: Merge Q4 and W8A8 weights using merge_safetensor_gguf.py before inference
  • Runtime Configuration: Use --backend_type balance_serve with --use_cuda_graph to enable the optimized NPU execution path
  • Key Files: ascend_experts.py (NPU kernels), npu_graph_runner.py (graph optimization), and main.py (server entry point)

Frequently Asked Questions

What Ascend NPU hardware is compatible with KTransformers?

KTransformers officially supports the Atlas 300I A2 server running CANN 8.3.xx. The implementation relies on specific kernel optimizations found in ktransformers/operators/ascend/ascend_experts.py that target the Ascend 310P/910B architecture. Ensure your driver version matches the CANN toolkit requirements documented in DeepseekR1_V3_tutorial_zh_for_Ascend_NPU.md.

Why is the ARM patch necessary when building for Ascend NPU?

Ascend NPU servers typically run on ARM64 (aarch64) architecture. The IQK (integer quantization kernel) implementations in third_party/llamafile/iqk_mul_mat_arm82.cpp contain ARM-specific SIMD optimizations that conflict with the Ascend NPU's own kernel dispatch mechanism. Commenting out the #define iqk_mul_mat iqk_mul_mat_arm82 line prevents symbol collisions while allowing the NPU-specific operators to handle matrix multiplication.

How does the balance_serve backend differ from standard CUDA inference?

The balance_serve backend, implemented in the server utilities, specifically partitions MoE workloads between CPU and NPU workers. Unlike standard CUDA backends that run entire layers on GPU, balance_serve in ktransformers/server/utils/create_interface.py schedules individual experts to the Ascend NPU via ascend_experts.py while keeping attention heads and embeddings on CPU, maximizing throughput for memory-bandwidth-bound MoE models.

What is the purpose of the NPU graph runner?

The npu_graph_runner.py utility creates persistent execution graphs that cache the NPU kernel launch sequence. When --use_cuda_graph is enabled, the system captures the MoE forward pass into a static graph during the first iteration, then replays it for subsequent tokens. This eliminates Python-level overhead and CPU-NPU synchronization delays, reducing per-token latency by 30-50% compared to eager execution mode.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →