# How to Set Up KTransformers with Ascend NPU for Inference: A Complete Guide

> Learn how to set up KTransformers with Ascend NPU for efficient MoE inference. Follow this guide for high-performance results using CANN and Docker.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: how-to-guide
- Published: 2026-07-26

---

**KTransformers enables high-performance MoE inference on Huawei Ascend NPUs by combining the `balance_serve` backend with custom NPU operators located in `ktransformers/operators/ascend`, requiring CANN 8.3+, Docker deployment, and specific ARM64 build patches.**

KTransformers is an open-source inference engine optimized for large-scale Mixture-of-Experts (MoE) models like DeepSeek-V3/R1. This guide explains how to set up KTransformers with Ascend NPU for inference, leveraging the `balance_serve` backend to schedule experts across NPU and CPU workers while minimizing latency through persistent NPU graphs.

## Prerequisites and Architecture Overview

Running KTransformers on Ascend NPUs requires a specific software stack and hardware configuration. According to the [`DeepseekR1_V3_tutorial_zh_for_Ascend_NPU.md`](https://github.com/kvcache-ai/ktransformers/blob/main/DeepseekR1_V3_tutorial_zh_for_Ascend_NPU.md) documentation, the architecture consists of:

- **Ascend NPU Hardware**: Atlas 300I A2 server or compatible Ascend NPU supported by CANN 8.3.xx
- **CANN Toolkit**: Ascend driver, HDK, and NNAL libraries ([`/usr/local/Ascend/ascend-toolkit/set_env.sh`](https://github.com/kvcache-ai/ktransformers/blob/main//usr/local/Ascend/ascend-toolkit/set_env.sh))
- **Docker Environment**: Pre-built Ubuntu 22.04 aarch64 image (`mindie:2.2.RC1-800I-A2-py311-openeuler24.03-lts`)
- **PyTorch for NPU**: `torch-npu` package for dispatching tensors to the Ascend device

The inference workflow splits MoE computation between the NPU (running optimized kernels in [`ascend_experts.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ascend_experts.py)) and CPU workers, orchestrated by the `balance_serve` backend defined in [`ktransformers/server/utils/create_interface.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/server/utils/create_interface.py).

## Step-by-Step Installation

### 1. Launch the Docker Container

Mount the Ascend drivers and toolkit into an interactive container:

```bash
docker run -it -d --net=host --shm-size=500g \
   --name kt-ascend \
   -w /workspace \
   --device=/dev/davinci_manager \
   --device=/dev/hisi_hdc \
   --device=/dev/devmm_svm \
   -v /usr/local/Ascend/driver:/usr/local/Ascend/driver:ro \
   -v /usr/local/dcmi:/usr/local/dcmi:ro \
   -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi:ro \
   -v /usr/local/sbin/:/usr/local/sbin:ro \
   -v <path_to_your_project>:/workspace \
   mindie:2.2.RC1-800I-A2-py311-openeuler24.03-lts bash

```

### 2. Install Dependencies and torch-npu

Inside the container, install system libraries and Python packages:

```bash

# Install system dependencies

yum install -y zlib-devel libtbb-devel openssl-devel libaio-devel \
    libcurl-devel

# Install PyTorch CPU wheel first, then torch-npu dependencies

pip3 install numpy==1.26.4 \
    torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 \
    --index-url https://download.pytorch.org/whl/cpu

pip3 install packaging ninja fire protobuf attrs \
    decorator cloudpickle ml-dtypes scipy tornado absl-py psutil \
    sqlalchemy transformers==4.57.1

```

### 3. Clone and Patch for ARM64

Ascend NPU servers run on ARM64 architecture, requiring a patch to disable x86-specific IQK kernels:

```bash
git clone https://github.com/kvcache-ai/ktransformers.git
cd ktransformers
git submodule update --init --recursive

# Comment out ARM-specific SIMD kernels (not needed for Ascend NPU)

sed -i 's/#define iqk_mul_mat iqk_mul_mat_arm82/\/\/#define iqk_mul_mat iqk_mul_mat_arm82/' \
    ./third_party/llamafile/iqk_mul_mat_arm82.cpp

```

### 4. Build with NPU Support

Compile the NPU-optimized operators using the installation script:

```bash
USE_BALANCE_SERVE=1 USE_NUMA=1 bash ./install.sh

```

This builds the custom operators in [`ktransformers/operators/ascend/ascend_experts.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/operators/ascend/ascend_experts.py), including `KExpertsCPUW8A8` and `KTransformersExpertsW8A8`, and enables the NUMA-aware memory allocator for optimal CPU-NPU data movement.

## Model Preparation

For the DeepSeek-R1/R3 models, merge the Q4 and W8A8 GGUF weights using the provided utility:

```bash
python merge_safetensor_gguf.py \
    --safetensor_path /mnt/weights/DeepSeek-R1-Q4_K_M \
    --gguf_path /mnt/weights/DeepSeek-R1-W8A8 \
    --output_path /mnt/weights/DeepSeek-R1-q4km-w8a8

```

This creates a unified weight directory compatible with the NPU `balance_serve` backend.

## Configuration and Runtime Settings

Configure the inference parameters in [`ktransformers/config/config.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/config/config.yaml) or via command-line arguments:

- **Backend**: Set `backend_type: balance_serve` to enable multi-concurrency NPU scheduling
- **Device**: Specify `device: npu` to trigger the Ascend code path
- **Memory**: Adjust `page_size: 128` and `chunk_size: 16384` for NPU-optimized KV cache management

The optimization rules in [`ktransformers/optimize/optimize_rules/npu/DeepSeek-V3-Chat-300IA2-npu-serve.yaml`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/optimize/optimize_rules/npu/DeepSeek-V3-Chat-300IA2-npu-serve.yaml) define kernel fusion patterns specific to the Atlas 300I A2 NPU.

## Launching the Inference Server

Set the required environment variables and start the server:

```bash
export USE_MERGE=0
export INF_NAN_MODE_FORCE_DISABLE=1
export TASK_QUEUE_ENABLE=0
export RANK=0
export LOCAL_WORLD_SIZE=1

# Source Ascend toolkit

source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh

python ktransformers/server/main.py \
    --port 10002 \
    --model_path /mnt/weights/DeepSeek-R1-q4km-w8a8 \
    --gguf_path /mnt/weights/DeepSeek-R1-q4km-w8a8 \
    --model_name DeepSeekV3ForCausalLM \
    --cpu_infer 100 \
    --optimize_config_path ./ktransformers/optimize/optimize_rules/npu/DeepSeek-V3-Chat-300IA2-npu-serve.yaml \
    --max_new_tokens 1024 \
    --cache_lens 20480 \
    --max_batch_size 4 \
    --use_cuda_graph \
    --tp 1 \
    --backend_type balance_serve

```

**Key flags explained:**

- **`--backend_type balance_serve`**: Activates the multi-concurrency engine that schedules MoE experts across NPU and CPU
- **`--use_cuda_graph`**: Enables the NPU graph runner ([`ktransformers/util/npu_graph_runner.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/util/npu_graph_runner.py)) to reduce kernel launch overhead by capturing and replaying NPU operations
- **`--cpu_infer 100`**: Allocates 100 CPU threads for non-expert operations and data preprocessing
- **`--optimize_config_path`**: Points to the YAML file containing NPU-specific kernel fusion rules

## Sending Inference Requests

Once the server is running on port 10002, send requests using the HTTP API:

```python
import requests
import json

url = "http://localhost:10002/generate"
payload = {
    "prompt": "Explain the advantages of Ascend NPU for LLM inference.",
    "max_new_tokens": 256,
    "temperature": 0.7
}

response = requests.post(url, json=payload)
result = json.loads(response.text)
print(result["generated_text"])

```

The server executes MoE routing on the Ascend NPU via the custom operators in [`ascend_experts.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ascend_experts.py), while the [`npu_graph_runner.py`](https://github.com/kvcache-ai/ktransformers/blob/main/npu_graph_runner.py) module manages persistent graphs to minimize dispatch latency for repeated inference calls.

## Summary

- **Hardware Requirements**: Atlas 300I A2 server with CANN 8.3+ toolkit and NNAL libraries installed on the host
- **Docker Deployment**: Use the official `mindie:2.2.RC1` image to obtain a pre-configured Ubuntu 22.04 aarch64 environment
- **Build Process**: Patch [`iqk_mul_mat_arm82.cpp`](https://github.com/kvcache-ai/ktransformers/blob/main/iqk_mul_mat_arm82.cpp) for ARM compatibility, then compile with `USE_BALANCE_SERVE=1 USE_NUMA=1`
- **Model Setup**: Merge Q4 and W8A8 weights using [`merge_safetensor_gguf.py`](https://github.com/kvcache-ai/ktransformers/blob/main/merge_safetensor_gguf.py) before inference
- **Runtime Configuration**: Use `--backend_type balance_serve` with `--use_cuda_graph` to enable the optimized NPU execution path
- **Key Files**: [`ascend_experts.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ascend_experts.py) (NPU kernels), [`npu_graph_runner.py`](https://github.com/kvcache-ai/ktransformers/blob/main/npu_graph_runner.py) (graph optimization), and [`main.py`](https://github.com/kvcache-ai/ktransformers/blob/main/main.py) (server entry point)

## Frequently Asked Questions

### What Ascend NPU hardware is compatible with KTransformers?

KTransformers officially supports the Atlas 300I A2 server running CANN 8.3.xx. The implementation relies on specific kernel optimizations found in [`ktransformers/operators/ascend/ascend_experts.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/operators/ascend/ascend_experts.py) that target the Ascend 310P/910B architecture. Ensure your driver version matches the CANN toolkit requirements documented in [`DeepseekR1_V3_tutorial_zh_for_Ascend_NPU.md`](https://github.com/kvcache-ai/ktransformers/blob/main/DeepseekR1_V3_tutorial_zh_for_Ascend_NPU.md).

### Why is the ARM patch necessary when building for Ascend NPU?

Ascend NPU servers typically run on ARM64 (aarch64) architecture. The IQK (integer quantization kernel) implementations in [`third_party/llamafile/iqk_mul_mat_arm82.cpp`](https://github.com/kvcache-ai/ktransformers/blob/main/third_party/llamafile/iqk_mul_mat_arm82.cpp) contain ARM-specific SIMD optimizations that conflict with the Ascend NPU's own kernel dispatch mechanism. Commenting out the `#define iqk_mul_mat iqk_mul_mat_arm82` line prevents symbol collisions while allowing the NPU-specific operators to handle matrix multiplication.

### How does the balance_serve backend differ from standard CUDA inference?

The `balance_serve` backend, implemented in the server utilities, specifically partitions MoE workloads between CPU and NPU workers. Unlike standard CUDA backends that run entire layers on GPU, `balance_serve` in [`ktransformers/server/utils/create_interface.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ktransformers/server/utils/create_interface.py) schedules individual experts to the Ascend NPU via [`ascend_experts.py`](https://github.com/kvcache-ai/ktransformers/blob/main/ascend_experts.py) while keeping attention heads and embeddings on CPU, maximizing throughput for memory-bandwidth-bound MoE models.

### What is the purpose of the NPU graph runner?

The [`npu_graph_runner.py`](https://github.com/kvcache-ai/ktransformers/blob/main/npu_graph_runner.py) utility creates persistent execution graphs that cache the NPU kernel launch sequence. When `--use_cuda_graph` is enabled, the system captures the MoE forward pass into a static graph during the first iteration, then replays it for subsequent tokens. This eliminates Python-level overhead and CPU-NPU synchronization delays, reducing per-token latency by 30-50% compared to eager execution mode.