How to Deploy GLM-5.2 on Ascend NPU with xLLM: Complete Setup Guide

Deploy GLM-5.2 on Ascend NPU by installing the Ascend driver and AI Toolkit, installing xLLM with Ascend support via pip install "xllm[ascend]", and launching the server with xllm serve --model ./GLM-5.2 --device ascend to enable W8A8 quantized MoE inference.

The GLM-5.2 mixture-of-experts (MoE) large language model achieves optimal inference performance on Huawei Ascend NPUs through the xLLM framework’s hardware-specific optimizations. Deploying GLM-5.2 on Ascend NPU with xLLM requires configuring the Ascend runtime, applying the W8A8 quantization pipeline, and utilizing MoE Mega-Fusion kernels. This guide references the official zai-org/GLM-5 repository, specifically the deployment instructions in example/ascend.md, to provide a production-ready setup workflow.

Prerequisites and Environment Setup

Deployment requires the Ascend driver and Ascend-AI-Toolkit installed on your Linux host to provide the kernel-level matrix-multiply and communication primitives (ReduceScatter/AllGather). The hardware layer consists of Ascend-AI-CPU and Ascend-NPU silicon, which xLLM targets for high-throughput inference.

Install the driver and toolkit via the official Huawei repositories:


# Install Ascend driver (replace <version> with your hardware-specific version)

sudo apt-get install ascend-driver-<version>

# Download and install the Ascend AI Toolkit

wget https://ascend-repo/hardware/Ascend-toolkit-<version>.run
bash Ascend-toolkit-<version>.run
source /usr/local/Ascend/ascend-toolkit/set_env.sh

Clone the GLM-5 repository and prepare the Python environment:

git clone https://github.com/zai-org/GLM-5.git
cd GLM-5
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt

Installing xLLM with Ascend Support

The xLLM runtime translates PyTorch-style model graphs into Ascend kernels, handling MoE-fusion, quantization, and prefill/decode disaggregation. Install xLLM with the Ascend extras to pull the core library plus Ascend-specific kernels:

pip install "xllm[ascend]"

This installation provides the xllm CLI and Python package necessary to serve GLM-5.2 on NPU hardware.

Model Conversion and Optimization

When you launch the server, xLLM automatically loads the HuggingFace checkpoint and applies the W8A8 hybrid quantization pipeline (QuaRot → Flex SmoothQuant → SSZ weight quantization). This process reduces compute cost while preserving accuracy by matching the Ascend NPU’s 8-bit arithmetic units.

Simultaneously, xLLM constructs the MoE Mega-Fusion operator, which fuses expert routing, weighted computation, and reduction into a single kernel. This eliminates intermediate tensor traffic, which is critical on NPUs where memory bandwidth is limited. The runtime also applies Communication-Computation Fusion, interleaving ReduceScatter/AllGather operations with GEMM kernels to hide AllReduce latency.

Launching the xLLM Server

Start the inference server using the xllm serve command with the --device ascend flag to target NPU hardware. The server automatically splits the workload into a prefill phase (large-batch token generation) and a decode phase (single-token streaming), using the prefill-delay scheduler to smooth throughput peaks.


# Download the model checkpoint (Git LFS required for large files)

git lfs install
git clone https://huggingface.co/zai-org/GLM-5.2

# Launch the server

xllm serve \
    --model ./GLM-5.2 \
    --device ascend \
    --port 8080 \
    --tensor-parallel-size 1

Upon startup, verify the logs indicate the MoE Mega-Fusion operator has compiled and the quantized weights loaded successfully. The server exposes an OpenAI-compatible API endpoint at http://localhost:8080/v1.

Client Integration and API Usage

Send requests using any OpenAI-compatible client. The reasoning_effort argument controls the MTP (Multi-Token Prediction) speculative decoding level, with options max (default) or high for increased generation speed.

import openai

openai.api_base = "http://localhost:8080/v1"
openai.api_key = "placeholder"  # xLLM does not enforce authentication

response = openai.ChatCompletion.create(
    model="glm-5.2",
    messages=[
        {"role": "user", "content": "Explain the benefits of MoE fusion on Ascend NPU."}
    ],
    temperature=0.7,
    max_tokens=256,
    reasoning_effort="high"  # Enables MTP speculative decoding

)

print(response.choices[0].message.content)

This configuration typically achieves 2-3× higher throughput compared to CPU-only inference, leveraging the Ascend-optimized pipeline.

Multi-Node Deployment (Optional)

For large-scale deployments, distribute the model across multiple nodes using xllm launch. xLLM automatically shards the MoE experts across the cluster, applying IndexCache and Prefix Caching mechanisms to maintain low latency.

xllm launch \
    --model ./GLM-5.2 \
    --device ascend \
    --nodes 4 \
    --pp-size 2 \
    --dp-size 2 \
    --port-base 8000

The --pp-size (pipeline parallelism) and --dp-size (data parallelism) flags control the distribution strategy. This setup utilizes IndexShare to efficiently manage expert routing across the NPU cluster.

Summary

  • Install Ascend drivers and the Ascend-AI-Toolkit before proceeding with Python dependencies.
  • Install xLLM with Ascend support via pip install "xllm[ascend]" to obtain NPU-specific kernels.
  • Deploy using xllm serve --device ascend, which automatically applies W8A8 quantization and MoE Mega-Fusion.
  • Optimize throughput using the reasoning_effort parameter for MTP speculative decoding and the prefill-delay scheduler for mixed workloads.
  • Scale horizontally with xllm launch for multi-node NPU clusters, utilizing pipeline and data parallelism.

Frequently Asked Questions

What hardware and software prerequisites are required for deploying GLM-5.2 on Ascend NPU?

You need Huawei Ascend NPUs with the Ascend driver and Ascend-AI-Toolkit installed, as detailed in example/ascend.md. The host system must run Linux and have Python 3.8+ with the dependencies listed in requirements.txt installed.

How does xLLM optimize MoE inference specifically for Ascend NPUs?

xLLM implements MoE Mega-Fusion, which combines routing, expert computation, and reduction into a single kernel to minimize memory bandwidth usage. It also fuses communication primitives (ReduceScatter/AllGather) with computation kernels to hide latency, matching the Ascend architecture’s strengths.

What is the purpose of the reasoning_effort parameter in the xLLM API?

The reasoning_effort parameter controls MTP (Multi-Token Prediction) speculative decoding. Setting it to "high" enables aggressive speculative decoding for faster token generation, while "max" (default) provides balanced performance. This is implemented in the xLLM runtime’s decode phase scheduler.

How does the prefill-decode disaggregation improve throughput?

xLLM separates the prefill phase (processing input prompts in large batches) from the decode phase (generating single tokens sequentially). The prefill-delay scheduler smooths peaks between these phases, ensuring stable throughput for concurrent chat sessions and preventing NPU memory bottlenecks during mixed workloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →