# How to Deploy GLM-5.2 on Ascend NPU with xLLM: Complete Setup Guide

> Deploy GLM-5.2 on Ascend NPU using xLLM. Follow our guide to install drivers, xLLM with Ascend support, and launch the server for efficient W8A8 quantized MoE inference.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-21

---

**Deploy GLM-5.2 on Ascend NPU by installing the Ascend driver and AI Toolkit, installing xLLM with Ascend support via `pip install "xllm[ascend]"`, and launching the server with `xllm serve --model ./GLM-5.2 --device ascend` to enable W8A8 quantized MoE inference.**

The GLM-5.2 mixture-of-experts (MoE) large language model achieves optimal inference performance on Huawei Ascend NPUs through the xLLM framework’s hardware-specific optimizations. Deploying GLM-5.2 on Ascend NPU with xLLM requires configuring the Ascend runtime, applying the W8A8 quantization pipeline, and utilizing MoE Mega-Fusion kernels. This guide references the official `zai-org/GLM-5` repository, specifically the deployment instructions in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md), to provide a production-ready setup workflow.

## Prerequisites and Environment Setup

Deployment requires the **Ascend driver** and **Ascend-AI-Toolkit** installed on your Linux host to provide the kernel-level matrix-multiply and communication primitives (ReduceScatter/AllGather). The hardware layer consists of Ascend-AI-CPU and Ascend-NPU silicon, which xLLM targets for high-throughput inference.

Install the driver and toolkit via the official Huawei repositories:

```bash

# Install Ascend driver (replace <version> with your hardware-specific version)

sudo apt-get install ascend-driver-<version>

# Download and install the Ascend AI Toolkit

wget https://ascend-repo/hardware/Ascend-toolkit-<version>.run
bash Ascend-toolkit-<version>.run
source /usr/local/Ascend/ascend-toolkit/set_env.sh

```

Clone the GLM-5 repository and prepare the Python environment:

```bash
git clone https://github.com/zai-org/GLM-5.git
cd GLM-5
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt

```

## Installing xLLM with Ascend Support

The xLLM runtime translates PyTorch-style model graphs into Ascend kernels, handling MoE-fusion, quantization, and prefill/decode disaggregation. Install xLLM with the Ascend extras to pull the core library plus Ascend-specific kernels:

```bash
pip install "xllm[ascend]"

```

This installation provides the `xllm` CLI and Python package necessary to serve GLM-5.2 on NPU hardware.

## Model Conversion and Optimization

When you launch the server, xLLM automatically loads the HuggingFace checkpoint and applies the **W8A8 hybrid quantization** pipeline (QuaRot → Flex SmoothQuant → SSZ weight quantization). This process reduces compute cost while preserving accuracy by matching the Ascend NPU’s 8-bit arithmetic units.

Simultaneously, xLLM constructs the **MoE Mega-Fusion** operator, which fuses expert routing, weighted computation, and reduction into a single kernel. This eliminates intermediate tensor traffic, which is critical on NPUs where memory bandwidth is limited. The runtime also applies **Communication-Computation Fusion**, interleaving ReduceScatter/AllGather operations with GEMM kernels to hide AllReduce latency.

## Launching the xLLM Server

Start the inference server using the `xllm serve` command with the `--device ascend` flag to target NPU hardware. The server automatically splits the workload into a **prefill** phase (large-batch token generation) and a **decode** phase (single-token streaming), using the **prefill-delay** scheduler to smooth throughput peaks.

```bash

# Download the model checkpoint (Git LFS required for large files)

git lfs install
git clone https://huggingface.co/zai-org/GLM-5.2

# Launch the server

xllm serve \
    --model ./GLM-5.2 \
    --device ascend \
    --port 8080 \
    --tensor-parallel-size 1

```

Upon startup, verify the logs indicate the MoE Mega-Fusion operator has compiled and the quantized weights loaded successfully. The server exposes an OpenAI-compatible API endpoint at `http://localhost:8080/v1`.

## Client Integration and API Usage

Send requests using any OpenAI-compatible client. The `reasoning_effort` argument controls the **MTP (Multi-Token Prediction)** speculative decoding level, with options `max` (default) or `high` for increased generation speed.

```python
import openai

openai.api_base = "http://localhost:8080/v1"
openai.api_key = "placeholder"  # xLLM does not enforce authentication

response = openai.ChatCompletion.create(
    model="glm-5.2",
    messages=[
        {"role": "user", "content": "Explain the benefits of MoE fusion on Ascend NPU."}
    ],
    temperature=0.7,
    max_tokens=256,
    reasoning_effort="high"  # Enables MTP speculative decoding

)

print(response.choices[0].message.content)

```

This configuration typically achieves **2-3× higher throughput** compared to CPU-only inference, leveraging the Ascend-optimized pipeline.

## Multi-Node Deployment (Optional)

For large-scale deployments, distribute the model across multiple nodes using `xllm launch`. xLLM automatically shards the MoE experts across the cluster, applying **IndexCache** and **Prefix Caching** mechanisms to maintain low latency.

```bash
xllm launch \
    --model ./GLM-5.2 \
    --device ascend \
    --nodes 4 \
    --pp-size 2 \
    --dp-size 2 \
    --port-base 8000

```

The `--pp-size` (pipeline parallelism) and `--dp-size` (data parallelism) flags control the distribution strategy. This setup utilizes **IndexShare** to efficiently manage expert routing across the NPU cluster.

## Summary

- **Install Ascend drivers** and the Ascend-AI-Toolkit before proceeding with Python dependencies.
- **Install xLLM** with Ascend support via `pip install "xllm[ascend]"` to obtain NPU-specific kernels.
- **Deploy** using `xllm serve --device ascend`, which automatically applies W8A8 quantization and MoE Mega-Fusion.
- **Optimize throughput** using the `reasoning_effort` parameter for MTP speculative decoding and the prefill-delay scheduler for mixed workloads.
- **Scale horizontally** with `xllm launch` for multi-node NPU clusters, utilizing pipeline and data parallelism.

## Frequently Asked Questions

### What hardware and software prerequisites are required for deploying GLM-5.2 on Ascend NPU?

You need Huawei Ascend NPUs with the Ascend driver and Ascend-AI-Toolkit installed, as detailed in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md). The host system must run Linux and have Python 3.8+ with the dependencies listed in [`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt) installed.

### How does xLLM optimize MoE inference specifically for Ascend NPUs?

xLLM implements **MoE Mega-Fusion**, which combines routing, expert computation, and reduction into a single kernel to minimize memory bandwidth usage. It also fuses communication primitives (ReduceScatter/AllGather) with computation kernels to hide latency, matching the Ascend architecture’s strengths.

### What is the purpose of the `reasoning_effort` parameter in the xLLM API?

The `reasoning_effort` parameter controls **MTP (Multi-Token Prediction)** speculative decoding. Setting it to `"high"` enables aggressive speculative decoding for faster token generation, while `"max"` (default) provides balanced performance. This is implemented in the xLLM runtime’s decode phase scheduler.

### How does the prefill-decode disaggregation improve throughput?

xLLM separates the **prefill** phase (processing input prompts in large batches) from the **decode** phase (generating single tokens sequentially). The **prefill-delay** scheduler smooths peaks between these phases, ensuring stable throughput for concurrent chat sessions and preventing NPU memory bottlenecks during mixed workloads.