# Deploying GLM-5 on Ascend NPU Platform: Optimization Guide for High-Throughput Inference

> Deploy GLM-5 on Ascend NPUs for high-throughput inference. Optimize large language models using vLLM-Ascend, SGLang, or xLLM with fused MoE and sparse attention.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-19

---

**Deploy GLM-5 on Ascend NPUs using vLLM-Ascend, SGLang, or xLLM with optimized sparse attention kernels and fused MoE operators to achieve high-throughput inference on the 744B parameter model.**

The zai-org/GLM-5 repository provides a production-ready implementation of the GLM-5 model series optimized for Ascend NPU hardware. Deploying GLM-5 on Ascend NPU platforms leverages specialized sparse attention mechanisms and fused operators to efficiently serve the 744-billion-parameter model with its 1-million-token context window. The official deployment documentation in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) and the repository README detail how to integrate these hardware-specific optimizations into your inference pipeline.

## Architecture Overview: Why GLM-5 Excels on Ascend NPU

GLM-5 (including GLM-5.1 and GLM-5.2) is built on a **DeepSeek Sparse Attention (DSA)** backbone with **IndexShare**-based sparse-attention layers, which dramatically reduces per-token FLOPs while preserving the 1M-token context window.

The following architectural components are specifically optimized for Ascend NPU execution:

- **DeepSeek Sparse Attention (DSA)** – Performs attention on a sparse set of tokens, cutting memory and compute cost. This matches the Ascend NPU's heterogeneous compute units, letting the hardware accelerate sparse matrix kernels.

- **IndexShare** – Reuses the same indexer across every four sparse-attention layers, saving routing overhead. This reduces the number of index look-ups that would otherwise dominate memory traffic on the NPU.

- **Mixture-of-Experts (MoE) Mega-Fusion Operator** – Fuses expert routing, weighted computation, and result reduction into a single kernel. This eliminates intermediate tensor reads/writes, a key bottleneck on Ascend's DMA-limited memory hierarchy.

- **Communication-Computation Fusion** – Splits AllReduce into ReduceScatter + AllGather and pipelines them with matrix multiplications. This hides inter-device communication latency on multi-NPU clusters.

- **Prefill-Delay Scheduling & Prefix Caching** – Separates prefill and decode phases, caches common prefixes, and smooths load spikes. This improves throughput stability when many concurrent requests share the same context.

- **Hybrid W8A8 Quantization (QuaRot + Flex SmoothQuant + SSZ)** – Compresses expert weights while keeping accuracy on critical paths. This reduces memory footprint enough to fit the model inside Ascend's on-chip SRAM, enabling faster inference.

## Prerequisites and Environment Setup

Before deploying GLM-5 on Ascend NPU platforms, ensure your environment meets the following requirements:

1. Install the Ascend driver and toolkit from Huawei's official repositories.
2. Install Python dependencies listed in [`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt) from the zai-org/GLM-5 repository.
3. Configure environment variables for NPU device access:

```bash
export ASCEND_DEVICE_ID=0
export LD_LIBRARY_PATH=/usr/local/Ascend/driver/lib64:${LD_LIBRARY_PATH}

```

The [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) file in the repository contains detailed environment setup instructions and validation steps.

## Deployment Methods

### Method 1: Single-Node Deployment with vLLM-Ascend

The **vLLM-Ascend** plugin adds Ascend-specific kernels and the MoE-fusion operator to the standard vLLM framework. This is the recommended approach for single-node deployments.

Install the plugin and launch the server:

```bash

# Install the Ascend plugin (requires Ascend driver & toolkit)

pip install "vllm[ascend]"

# Launch the server

python -m vllm.entrypoints.openai.api_server \
    --model zai-org/GLM-5.2 \
    --dtype bfloat16 \
    --tensor-parallel-size 1 \
    --port 8080 \
    --reasoning_effort max

```

Query the deployed model using the OpenAI-compatible API:

```bash
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"zai-org/GLM-5.2","messages":[{"role":"user","content":"Explain the benefits of sparse attention on Ascend NPU."}]}'

```

### Method 2: SGLang Ascend Backend

**SGLang** offers an Ascend-optimized backend that implements the same fused operators with a flexible configuration system.

Install SGLang with Ascend support:

```bash
git clone https://github.com/sgl-project/sglang
cd sglang
pip install -e ".[ascend]"

```

Create a configuration file [`sgl_config.yaml`](https://github.com/zai-org/GLM-5/blob/main/sgl_config.yaml):

```yaml
model: "zai-org/GLM-5.2"
device: "ascend"
dtype: "bfloat16"
reasoning_effort: "max"

```

Start the server and query via the Python SDK:

```bash
sgl serve -c sgl_config.yaml

```

```python
from sglang import SGLangClient

client = SGLangClient("http://localhost:9000")
resp = client.chat(messages=[{"role":"user","content":"What is IndexShare?"}])
print(resp.choices[0].message.content)

```

### Method 3: Multi-Node Cluster with xLLM

For distributed deployments across multiple Ascend nodes, **xLLM** provides a quick-start script illustrating environment setup and multi-node configuration.

Install the xLLM Ascend plugin:

```bash
pip install "xllm[ascend]"

```

Create a launch script [`run_glm5.sh`](https://github.com/zai-org/GLM-5/blob/main/run_glm5.sh):

```bash
#!/bin/bash
export ASCEND_DEVICE_ID=$1          # device id per node

export MASTER_ADDR=$2               # IP of rank 0

export MASTER_PORT=29500
export WORLD_SIZE=$3                # total # of GPUs

export RANK=$4                      # local rank

python -m xllm.entrypoints.launch \
    --model zai-org/GLM-5.2 \
    --dtype bfloat16 \
    --reasoning_effort high \
    --port 8080

```

Launch on each node of a 4-node cluster:

```bash

# Node 0

./run_glm5.sh 0 10.0.0.1 4 0 &

# Node 1

./run_glm5.sh 1 10.0.0.1 4 1 &

# Node 2

./run_glm5.sh 2 10.0.0.1 4 2 &

# Node 3

./run_glm5.sh 3 10.0.0.1 4 3 &

```

## Key Optimizations for Ascend Hardware

When deploying GLM-5 on Ascend NPU platforms, the inference frameworks automatically offload heavy operations to the NPU while the host CPU orchestrates request scheduling and cache management.

The **MoE Mega-Fusion Operator** is particularly critical for Ascend performance. According to the source code in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md), this operator fuses expert routing, weighted computation, and result reduction into a single kernel, eliminating intermediate tensor reads/writes that would otherwise bottleneck the DMA-limited memory hierarchy.

Similarly, **Communication-Computation Fusion** splits AllReduce operations into ReduceScatter and AllGather phases, pipelining them with matrix multiplications to hide inter-device communication latency on multi-NPU clusters.

## Configuration Parameters

All three frameworks expose a **`reasoning_effort`** knob and an **`enable_thinking`** flag that control the model's internal speculative decoding pipeline.

- **`reasoning_effort`**: Set to `"max"` for highest quality reasoning or `"high"` for balanced performance. This parameter adjusts the depth of the speculative decoding pipeline.
- **`enable_thinking`**: Boolean flag to enable or disable the model's chain-of-thought generation capabilities.

These parameters are documented in the [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) under the "Serve GLM-5 Series Locally" section and can be passed via command-line arguments or configuration files.

## Summary

- **GLM-5** leverages **DeepSeek Sparse Attention (DSA)** and **IndexShare** optimizations specifically designed for Ascend NPU's sparse compute capabilities.
- Three primary frameworks support Ascend deployment: **vLLM-Ascend** for single-node setups, **SGLang** for flexible backend configurations, and **xLLM** for multi-node clusters.
- The **MoE Mega-Fusion Operator** and **Communication-Computation Fusion** eliminate memory bottlenecks and hide latency on Ascend's DMA-limited architecture.
- Use **`reasoning_effort`** (`max`/`high`) and **`enable_thinking`** flags to control the speculative decoding pipeline and reasoning depth.
- Reference [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) and [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) in the zai-org/GLM-5 repository for the latest deployment instructions and optimization settings.

## Frequently Asked Questions

### What makes GLM-5 compatible with Ascend NPUs?

GLM-5 implements **DeepSeek Sparse Attention (DSA)** with **IndexShare**-based sparse-attention layers that map efficiently to Ascend NPU's heterogeneous compute units. The model's sparse matrix kernels and fused MoE operators are specifically optimized for Ascend's memory hierarchy, as detailed in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md).

### How do I enable sparse attention optimizations on Ascend?

Sparse attention optimizations are automatically enabled when using supported frameworks like vLLM-Ascend, SGLang, or xLLM. These frameworks implement the **IndexShare** mechanism that reuses indexers across every four sparse-attention layers, reducing memory traffic. No manual configuration is required beyond installing the Ascend-specific plugin versions.

### What is the difference between reasoning_effort and enable_thinking?

The **`reasoning_effort`** parameter (`max` or `high`) controls the computational depth of the speculative decoding pipeline, affecting inference quality and latency. The **`enable_thinking`** flag is a boolean toggle that activates or deactivates the model's internal chain-of-thought generation. According to the [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), these parameters allow fine-grained control over the trade-off between reasoning quality and throughput.

### Can I deploy GLM-5 on a multi-node Ascend cluster?

Yes. The **xLLM** framework provides native support for multi-node deployment on Ascend clusters, as demonstrated in the [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) documentation. You can also use vLLM-Ascend with tensor parallelism across multiple NPUs by adjusting the `--tensor-parallel-size` parameter and setting appropriate environment variables for distributed communication.