How to Optimize MoE Inference with Mega-Fusion Operators on Ascend NPU
Enable the Mega-Fusion operator through framework-specific flags (use_mega_fusion, mega_fusion, or XLLM_MEGAFUSION) to fuse routing, expert computation, and reduction into a single kernel, eliminating redundant memory traffic and delivering up to 2.9× FLOP reduction on Ascend NPUs.
The GLM-5 series (GLM-5, GLM-5.1, GLM-5.2) employs a Mixture-of-Experts (MoE) architecture that traditionally suffers from memory-bound bottlenecks during inference. By implementing Mega-Fusion operators specifically optimized for Ascend NPU hardware, as documented in the zai-org/GLM-5 repository, you can collapse separate routing, computation, and reduction steps into unified kernels that minimize data movement and maximize compute unit utilization.
Why Mega-Fusion Matters for MoE Performance
Reduced Memory Traffic
Traditional MoE pipelines execute routing, expert weight fetching, computation, and result reduction as discrete steps with separate memory loads and stores. According to example/ascend.md (line 5), Mega-Fusion collapses these operations so tensor data flows through the hardware pipeline exactly once, eliminating redundant intermediate reads and writes of intermediate tensors.
Better Compute Unit Utilization
By merging expert kernels into a single operation, the Ascend NPU maintains continuous utilization of matrix-multiply units without idle cycles typically wasted waiting for routing or reduction boundaries. The operator keeps the NPU's compute units busy while processing expert routing and weighted computation simultaneously.
Lower Latency at Long Context
The operator shortens the critical path for token generation, enabling efficient processing of contexts up to 1 million tokens while maintaining competitive throughput with GPU backends. This optimization is particularly effective when combined with the IndexShare mechanism described in README.md (lines 25-26).
How Mega-Fusion Works on Ascend NPU
The Mega-Fusion operator architecture, as implemented in the GLM-5 Ascend guide, fuses four critical components into a unified pipeline:
- Routing: Generation of top-k expert indices per token, removing the separate top-k kernel overhead
- Weighted Expert Computation: Expert sub-network calls (typically linear layers) multiplexed into a single large GEMM that processes selected experts
- Reduction: Weighted sum of expert outputs integrated directly into the MoE tensor computation, avoiding a second pass over output tensors
- Communication-Computation Fusion: AllReduce operations split into ReduceScatter and AllGather phases, pipelined with GEMM execution to hide communication latency (see
example/ascend.md, line 6)
This architecture achieves a 2.9× FLOP reduction at 1M context length when combined with IndexShare optimizations, making Ascend deployment competitive with GPU backends.
Enabling Mega-Fusion in Inference Frameworks
The GLM-5 repository supports three Ascend-enabled inference frameworks. Each exposes Mega-Fusion through a simple configuration flag while internally handling operator selection, kernel compilation, and routing-weight packing.
vLLM-Ascend
In vLLM-Ascend, set use_mega_fusion=True in the engine options to activate the fused MoE kernel:
from vllm import LLM, SamplingParams
engine_opts = {
"use_mega_fusion": True, # Activates the fused MoE kernel
"device": "ascend"
}
llm = LLM(
model="zai-org/GLM-5.2",
tokenizer="zai-org/GLM-5.2",
engine_options=engine_opts,
)
sampler = SamplingParams(temperature=0.8, max_tokens=256)
output = llm.generate(
prompts=["Explain the benefits of Mega-Fusion"],
sampling_params=sampler
)
print(output[0].text)
SGLang
Configure via YAML or the ascend.json configuration file by setting mega_fusion: true:
model:
name: "zai-org/GLM-5.2"
device: "ascend"
runtime:
mega_fusion: true # Enables the fused MoE kernel
prefilling:
enable: true
Launch the server with:
sglang serve --config ascend_config.yaml
xLLM
Use environment variables to activate Mega-Fusion before starting the inference server:
export XLLM_MEGAFUSION=1 # Activate Mega-Fusion
export XLLM_DEVICE=ascend # Target Ascend NPU
xllm-server --model zai-org/GLM-5.2 --port 8080
After the server starts, requests to the /generate endpoint automatically utilize the fused MoE kernel.
Key Source Files and Reference Documentation
Understanding the implementation requires examining these specific files in the zai-org/GLM-5 repository:
example/ascend.md: Contains the high-level overview of Ascend-specific optimizations, including the Mega-Fusion operator architecture that fuses "expert routing, weighted computation, and result reduction" into a single unified operator.README.md(lines 25-26): Documents the IndexShare mechanism that complements Mega-Fusion to achieve the 2.9× FLOP reduction metric at 1M context.README.md(line 79): Lists supported Ascend inference frameworks (vLLM-Ascend, xLLM, SGLang).skills/glm-master-skill/SKILL.md: Provides generic skill descriptions and architectural context for GLM-5 models.
Summary
- Enable Mega-Fusion through framework-specific flags (
use_mega_fusion,mega_fusion, orXLLM_MEGAFUSION) to fuse routing, expert computation, and reduction into single Ascend NPU kernels. - Eliminate memory bottlenecks by reducing intermediate tensor reads/writes, achieving up to 2.9× FLOP reduction at 1M token contexts.
- Deploy across three frameworks (vLLM-Ascend, SGLang, xLLM) using simple configuration changes without modifying underlying MoE logic.
- Reference implementation details in
example/ascend.mdandREADME.mdto understand the communication-computation fusion and IndexShare optimizations that accompany Mega-Fusion.
Frequently Asked Questions
What specific operations does the Mega-Fusion operator combine?
The Mega-Fusion operator merges four distinct operations: (1) top-k expert routing index generation, (2) weighted expert computation via multiplexed GEMM calls, (3) reduction of expert outputs into the final MoE tensor, and (4) communication primitives (ReduceScatter/AllGather) pipelined with computation. According to the Ascend guide in example/ascend.md, this fusion eliminates redundant reads and writes of intermediate tensors while significantly boosting computational efficiency.
How much performance improvement can I expect when enabling Mega-Fusion?
According to the GLM-5 repository documentation, Mega-Fusion combined with IndexShare optimizations achieves a 2.9× FLOP reduction when processing contexts up to 1 million tokens. The actual throughput improvement depends on batch size and context length, with the most significant gains appearing in long-context scenarios where memory bandwidth typically limits traditional MoE implementations.
Do I need to modify the model weights or architecture to use Mega-Fusion?
No. The Mega-Fusion operator works transparently with standard GLM-5 checkpoint formats. The inference framework (vLLM-Ascend, SGLang, or xLLM) handles kernel compilation, operator selection, and routing-weight packing automatically when you enable the respective configuration flag. The model architecture remains unchanged; only the execution strategy on Ascend NPU hardware differs.
Which Ascend NPU hardware generations support Mega-Fusion for GLM-5?
The GLM-5 repository documentation in example/ascend.md describes these optimizations for modern Ascend NPUs capable of running the supported inference frameworks (vLLM-Ascend, SGLang, and xLLM). While specific chipset generations aren't explicitly version-locked in the documentation, the fused operators require Ascend hardware with unified memory architecture and compute units capable of executing combined GEMM-communication patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →