Deploying GLM-5 on Ascend NPU Platform: Optimization Guide for High-Throughput Inference
Deploy GLM-5 on Ascend NPUs using vLLM-Ascend, SGLang, or xLLM with optimized sparse attention kernels and fused MoE operators to achieve high-throughput inference on the 744B parameter model.
The zai-org/GLM-5 repository provides a production-ready implementation of the GLM-5 model series optimized for Ascend NPU hardware. Deploying GLM-5 on Ascend NPU platforms leverages specialized sparse attention mechanisms and fused operators to efficiently serve the 744-billion-parameter model with its 1-million-token context window. The official deployment documentation in example/ascend.md and the repository README detail how to integrate these hardware-specific optimizations into your inference pipeline.
Architecture Overview: Why GLM-5 Excels on Ascend NPU
GLM-5 (including GLM-5.1 and GLM-5.2) is built on a DeepSeek Sparse Attention (DSA) backbone with IndexShare-based sparse-attention layers, which dramatically reduces per-token FLOPs while preserving the 1M-token context window.
The following architectural components are specifically optimized for Ascend NPU execution:
-
DeepSeek Sparse Attention (DSA) – Performs attention on a sparse set of tokens, cutting memory and compute cost. This matches the Ascend NPU's heterogeneous compute units, letting the hardware accelerate sparse matrix kernels.
-
IndexShare – Reuses the same indexer across every four sparse-attention layers, saving routing overhead. This reduces the number of index look-ups that would otherwise dominate memory traffic on the NPU.
-
Mixture-of-Experts (MoE) Mega-Fusion Operator – Fuses expert routing, weighted computation, and result reduction into a single kernel. This eliminates intermediate tensor reads/writes, a key bottleneck on Ascend's DMA-limited memory hierarchy.
-
Communication-Computation Fusion – Splits AllReduce into ReduceScatter + AllGather and pipelines them with matrix multiplications. This hides inter-device communication latency on multi-NPU clusters.
-
Prefill-Delay Scheduling & Prefix Caching – Separates prefill and decode phases, caches common prefixes, and smooths load spikes. This improves throughput stability when many concurrent requests share the same context.
-
Hybrid W8A8 Quantization (QuaRot + Flex SmoothQuant + SSZ) – Compresses expert weights while keeping accuracy on critical paths. This reduces memory footprint enough to fit the model inside Ascend's on-chip SRAM, enabling faster inference.
Prerequisites and Environment Setup
Before deploying GLM-5 on Ascend NPU platforms, ensure your environment meets the following requirements:
- Install the Ascend driver and toolkit from Huawei's official repositories.
- Install Python dependencies listed in
requirements.txtfrom the zai-org/GLM-5 repository. - Configure environment variables for NPU device access:
export ASCEND_DEVICE_ID=0
export LD_LIBRARY_PATH=/usr/local/Ascend/driver/lib64:${LD_LIBRARY_PATH}
The example/ascend.md file in the repository contains detailed environment setup instructions and validation steps.
Deployment Methods
Method 1: Single-Node Deployment with vLLM-Ascend
The vLLM-Ascend plugin adds Ascend-specific kernels and the MoE-fusion operator to the standard vLLM framework. This is the recommended approach for single-node deployments.
Install the plugin and launch the server:
# Install the Ascend plugin (requires Ascend driver & toolkit)
pip install "vllm[ascend]"
# Launch the server
python -m vllm.entrypoints.openai.api_server \
--model zai-org/GLM-5.2 \
--dtype bfloat16 \
--tensor-parallel-size 1 \
--port 8080 \
--reasoning_effort max
Query the deployed model using the OpenAI-compatible API:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"zai-org/GLM-5.2","messages":[{"role":"user","content":"Explain the benefits of sparse attention on Ascend NPU."}]}'
Method 2: SGLang Ascend Backend
SGLang offers an Ascend-optimized backend that implements the same fused operators with a flexible configuration system.
Install SGLang with Ascend support:
git clone https://github.com/sgl-project/sglang
cd sglang
pip install -e ".[ascend]"
Create a configuration file sgl_config.yaml:
model: "zai-org/GLM-5.2"
device: "ascend"
dtype: "bfloat16"
reasoning_effort: "max"
Start the server and query via the Python SDK:
sgl serve -c sgl_config.yaml
from sglang import SGLangClient
client = SGLangClient("http://localhost:9000")
resp = client.chat(messages=[{"role":"user","content":"What is IndexShare?"}])
print(resp.choices[0].message.content)
Method 3: Multi-Node Cluster with xLLM
For distributed deployments across multiple Ascend nodes, xLLM provides a quick-start script illustrating environment setup and multi-node configuration.
Install the xLLM Ascend plugin:
pip install "xllm[ascend]"
Create a launch script run_glm5.sh:
#!/bin/bash
export ASCEND_DEVICE_ID=$1 # device id per node
export MASTER_ADDR=$2 # IP of rank 0
export MASTER_PORT=29500
export WORLD_SIZE=$3 # total # of GPUs
export RANK=$4 # local rank
python -m xllm.entrypoints.launch \
--model zai-org/GLM-5.2 \
--dtype bfloat16 \
--reasoning_effort high \
--port 8080
Launch on each node of a 4-node cluster:
# Node 0
./run_glm5.sh 0 10.0.0.1 4 0 &
# Node 1
./run_glm5.sh 1 10.0.0.1 4 1 &
# Node 2
./run_glm5.sh 2 10.0.0.1 4 2 &
# Node 3
./run_glm5.sh 3 10.0.0.1 4 3 &
Key Optimizations for Ascend Hardware
When deploying GLM-5 on Ascend NPU platforms, the inference frameworks automatically offload heavy operations to the NPU while the host CPU orchestrates request scheduling and cache management.
The MoE Mega-Fusion Operator is particularly critical for Ascend performance. According to the source code in example/ascend.md, this operator fuses expert routing, weighted computation, and result reduction into a single kernel, eliminating intermediate tensor reads/writes that would otherwise bottleneck the DMA-limited memory hierarchy.
Similarly, Communication-Computation Fusion splits AllReduce operations into ReduceScatter and AllGather phases, pipelining them with matrix multiplications to hide inter-device communication latency on multi-NPU clusters.
Configuration Parameters
All three frameworks expose a reasoning_effort knob and an enable_thinking flag that control the model's internal speculative decoding pipeline.
reasoning_effort: Set to"max"for highest quality reasoning or"high"for balanced performance. This parameter adjusts the depth of the speculative decoding pipeline.enable_thinking: Boolean flag to enable or disable the model's chain-of-thought generation capabilities.
These parameters are documented in the README.md under the "Serve GLM-5 Series Locally" section and can be passed via command-line arguments or configuration files.
Summary
- GLM-5 leverages DeepSeek Sparse Attention (DSA) and IndexShare optimizations specifically designed for Ascend NPU's sparse compute capabilities.
- Three primary frameworks support Ascend deployment: vLLM-Ascend for single-node setups, SGLang for flexible backend configurations, and xLLM for multi-node clusters.
- The MoE Mega-Fusion Operator and Communication-Computation Fusion eliminate memory bottlenecks and hide latency on Ascend's DMA-limited architecture.
- Use
reasoning_effort(max/high) andenable_thinkingflags to control the speculative decoding pipeline and reasoning depth. - Reference
example/ascend.mdandREADME.mdin the zai-org/GLM-5 repository for the latest deployment instructions and optimization settings.
Frequently Asked Questions
What makes GLM-5 compatible with Ascend NPUs?
GLM-5 implements DeepSeek Sparse Attention (DSA) with IndexShare-based sparse-attention layers that map efficiently to Ascend NPU's heterogeneous compute units. The model's sparse matrix kernels and fused MoE operators are specifically optimized for Ascend's memory hierarchy, as detailed in example/ascend.md.
How do I enable sparse attention optimizations on Ascend?
Sparse attention optimizations are automatically enabled when using supported frameworks like vLLM-Ascend, SGLang, or xLLM. These frameworks implement the IndexShare mechanism that reuses indexers across every four sparse-attention layers, reducing memory traffic. No manual configuration is required beyond installing the Ascend-specific plugin versions.
What is the difference between reasoning_effort and enable_thinking?
The reasoning_effort parameter (max or high) controls the computational depth of the speculative decoding pipeline, affecting inference quality and latency. The enable_thinking flag is a boolean toggle that activates or deactivates the model's internal chain-of-thought generation. According to the README.md, these parameters allow fine-grained control over the trade-off between reasoning quality and throughput.
Can I deploy GLM-5 on a multi-node Ascend cluster?
Yes. The xLLM framework provides native support for multi-node deployment on Ascend clusters, as demonstrated in the example/ascend.md documentation. You can also use vLLM-Ascend with tensor parallelism across multiple NPUs by adjusting the --tensor-parallel-size parameter and setting appropriate environment variables for distributed communication.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →