# How to Optimize MoE Inference with Mega-Fusion Operators on Ascend NPU

> Optimize MoE inference on Ascend NPUs with Mega-Fusion operators. Fuse routing, expert computation, and reduction for up to 2.9x FLOP reduction and faster performance.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-21

---

**Enable the Mega-Fusion operator through framework-specific flags (`use_mega_fusion`, `mega_fusion`, or `XLLM_MEGAFUSION`) to fuse routing, expert computation, and reduction into a single kernel, eliminating redundant memory traffic and delivering up to 2.9× FLOP reduction on Ascend NPUs.**

The GLM-5 series (GLM-5, GLM-5.1, GLM-5.2) employs a **Mixture-of-Experts (MoE)** architecture that traditionally suffers from memory-bound bottlenecks during inference. By implementing Mega-Fusion operators specifically optimized for Ascend NPU hardware, as documented in the `zai-org/GLM-5` repository, you can collapse separate routing, computation, and reduction steps into unified kernels that minimize data movement and maximize compute unit utilization.

## Why Mega-Fusion Matters for MoE Performance

### Reduced Memory Traffic

Traditional MoE pipelines execute routing, expert weight fetching, computation, and result reduction as discrete steps with separate memory loads and stores. According to [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) (line 5), Mega-Fusion collapses these operations so tensor data flows through the hardware pipeline exactly once, eliminating redundant intermediate reads and writes of intermediate tensors.

### Better Compute Unit Utilization

By merging expert kernels into a single operation, the Ascend NPU maintains continuous utilization of matrix-multiply units without idle cycles typically wasted waiting for routing or reduction boundaries. The operator keeps the NPU's compute units busy while processing expert routing and weighted computation simultaneously.

### Lower Latency at Long Context

The operator shortens the critical path for token generation, enabling efficient processing of contexts up to 1 million tokens while maintaining competitive throughput with GPU backends. This optimization is particularly effective when combined with the IndexShare mechanism described in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (lines 25-26).

## How Mega-Fusion Works on Ascend NPU

The Mega-Fusion operator architecture, as implemented in the GLM-5 Ascend guide, fuses four critical components into a unified pipeline:

- **Routing**: Generation of top-k expert indices per token, removing the separate top-k kernel overhead
- **Weighted Expert Computation**: Expert sub-network calls (typically linear layers) multiplexed into a single large GEMM that processes selected experts
- **Reduction**: Weighted sum of expert outputs integrated directly into the MoE tensor computation, avoiding a second pass over output tensors
- **Communication-Computation Fusion**: AllReduce operations split into ReduceScatter and AllGather phases, pipelined with GEMM execution to hide communication latency (see [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md), line 6)

This architecture achieves a **2.9× FLOP reduction** at 1M context length when combined with IndexShare optimizations, making Ascend deployment competitive with GPU backends.

## Enabling Mega-Fusion in Inference Frameworks

The GLM-5 repository supports three Ascend-enabled inference frameworks. Each exposes Mega-Fusion through a simple configuration flag while internally handling operator selection, kernel compilation, and routing-weight packing.

### vLLM-Ascend

In `vLLM-Ascend`, set `use_mega_fusion=True` in the engine options to activate the fused MoE kernel:

```python
from vllm import LLM, SamplingParams

engine_opts = {
    "use_mega_fusion": True,      # Activates the fused MoE kernel

    "device": "ascend"
}

llm = LLM(
    model="zai-org/GLM-5.2",
    tokenizer="zai-org/GLM-5.2",
    engine_options=engine_opts,
)

sampler = SamplingParams(temperature=0.8, max_tokens=256)
output = llm.generate(
    prompts=["Explain the benefits of Mega-Fusion"], 
    sampling_params=sampler
)
print(output[0].text)

```

### SGLang

Configure via YAML or the [`ascend.json`](https://github.com/zai-org/GLM-5/blob/main/ascend.json) configuration file by setting `mega_fusion: true`:

```yaml
model:
  name: "zai-org/GLM-5.2"
  device: "ascend"

runtime:
  mega_fusion: true          # Enables the fused MoE kernel

  prefilling:
    enable: true

```

Launch the server with:

```bash
sglang serve --config ascend_config.yaml

```

### xLLM

Use environment variables to activate Mega-Fusion before starting the inference server:

```bash
export XLLM_MEGAFUSION=1          # Activate Mega-Fusion

export XLLM_DEVICE=ascend           # Target Ascend NPU

xllm-server --model zai-org/GLM-5.2 --port 8080

```

After the server starts, requests to the `/generate` endpoint automatically utilize the fused MoE kernel.

## Key Source Files and Reference Documentation

Understanding the implementation requires examining these specific files in the `zai-org/GLM-5` repository:

- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)**: Contains the high-level overview of Ascend-specific optimizations, including the Mega-Fusion operator architecture that fuses "expert routing, weighted computation, and result reduction" into a single unified operator.
- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (lines 25-26)**: Documents the IndexShare mechanism that complements Mega-Fusion to achieve the 2.9× FLOP reduction metric at 1M context.
- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (line 79)**: Lists supported Ascend inference frameworks (vLLM-Ascend, xLLM, SGLang).
- **[`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)**: Provides generic skill descriptions and architectural context for GLM-5 models.

## Summary

- **Enable Mega-Fusion** through framework-specific flags (`use_mega_fusion`, `mega_fusion`, or `XLLM_MEGAFUSION`) to fuse routing, expert computation, and reduction into single Ascend NPU kernels.
- **Eliminate memory bottlenecks** by reducing intermediate tensor reads/writes, achieving up to 2.9× FLOP reduction at 1M token contexts.
- **Deploy across three frameworks** (vLLM-Ascend, SGLang, xLLM) using simple configuration changes without modifying underlying MoE logic.
- **Reference implementation details** in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) and [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) to understand the communication-computation fusion and IndexShare optimizations that accompany Mega-Fusion.

## Frequently Asked Questions

### What specific operations does the Mega-Fusion operator combine?

The Mega-Fusion operator merges four distinct operations: (1) top-k expert routing index generation, (2) weighted expert computation via multiplexed GEMM calls, (3) reduction of expert outputs into the final MoE tensor, and (4) communication primitives (ReduceScatter/AllGather) pipelined with computation. According to the Ascend guide in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md), this fusion eliminates redundant reads and writes of intermediate tensors while significantly boosting computational efficiency.

### How much performance improvement can I expect when enabling Mega-Fusion?

According to the GLM-5 repository documentation, Mega-Fusion combined with IndexShare optimizations achieves a **2.9× FLOP reduction** when processing contexts up to 1 million tokens. The actual throughput improvement depends on batch size and context length, with the most significant gains appearing in long-context scenarios where memory bandwidth typically limits traditional MoE implementations.

### Do I need to modify the model weights or architecture to use Mega-Fusion?

No. The Mega-Fusion operator works transparently with standard GLM-5 checkpoint formats. The inference framework (vLLM-Ascend, SGLang, or xLLM) handles kernel compilation, operator selection, and routing-weight packing automatically when you enable the respective configuration flag. The model architecture remains unchanged; only the execution strategy on Ascend NPU hardware differs.

### Which Ascend NPU hardware generations support Mega-Fusion for GLM-5?

The GLM-5 repository documentation in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) describes these optimizations for modern Ascend NPUs capable of running the supported inference frameworks (vLLM-Ascend, SGLang, and xLLM). While specific chipset generations aren't explicitly version-locked in the documentation, the fused operators require Ascend hardware with unified memory architecture and compute units capable of executing combined GEMM-communication patterns.