How to Use W8A8 Quantization with QuaRot and Flex SmoothQuant for GLM-5.2

GLM-5.2 implements a three-stage W8A8 quantization pipeline that applies QuaRot rotation, Flex SmoothQuant smoothing, and SSZ 8-bit weight quantization to reduce model size while maintaining accuracy on critical MoE expert paths.

The zai-org/GLM-5 repository provides native support for hybrid W8A8 quantization, enabling efficient deployment on Ascend NPUs and compatible accelerators. This approach compresses both weights and activations to 8-bit integers using a novel combination of rotation and smoothing techniques. According to the Ascend deployment guide in example/ascend.md, this method preserves full-precision quality on the model's critical pathways while dramatically reducing memory bandwidth requirements.

The Three-Stage Quantization Pipeline

The W8A8 quantization process in GLM-5.2 follows a specific transformation sequence that must be applied in order:

Stage 1: QuaRot Pre-processing

QuaRot rotates weight matrices into a basis that concentrates energy in the first few dimensions. This rotation makes the weight distributions more amenable to low-bit quantization by reducing the variance across dimensions. In the GLM-5.2 implementation, this step is applied to all linear layers before any scaling or quantization occurs.

Stage 2: Flex SmoothQuant Smoothing

Flex SmoothQuant applies per-channel (or per-tensor) scaling to both activations and weights using a smoothness factor that balances dynamic ranges. This prevents outliers from dominating the quantization step and preserves precision in critical pathways. The smoothing factor is typically set to 0.5, though it can be tuned based on calibration data.

Stage 3: SSZ Weight Quantization

SSZ (Symmetric Signed-Zero) quantization converts the rotated and smoothed weights into 8-bit integers. This symmetric quantization scheme represents both weights and activations in 8-bit format, achieving the final W8A8 configuration. The resulting checkpoint maintains the same architecture but with all linear layer weights stored as int8 tensors.

Quantizing a GLM-5.2 Checkpoint

To convert a full-precision GLM-5.2 checkpoint to W8A8 format, use the quantization utilities provided in the repository. The following script applies all three stages sequentially:


# quantize.py

import torch
from glm5.quant import qua_rot, flex_smoothquant, ssz_quantizer

def quantize_glm5ckpt(fp16_path, out_path):
    # Load FP16 checkpoint

    ckpt = torch.load(fp16_path, map_location="cpu")

    # ① QuaRot pre-processing

    ckpt = qua_rot(ckpt)

    # ② Flex SmoothQuant smoothing

    ckpt = flex_smoothquant(ckpt, smooth_factor=0.5)

    # ③ SSZ 8-bit weight quantization

    ckpt = ssz_quantizer(ckpt)

    # Save the W8A8 checkpoint

    torch.save(ckpt, out_path)
    print(f"Quantized checkpoint saved to {out_path}")

if __name__ == "__main__":
    import argparse
    parser = argparse.ArgumentParser()
    parser.add_argument("--fp16", required=True, help="Path to original FP16 checkpoint")
    parser.add_argument("--out", default="glm5.2_w8a8.pt", help="Output path")
    args = parser.parse_args()
    quantize_glm5ckpt(args.fp16, args.out)

Execute the script from your terminal:

python quantize.py --fp16 glm5.2_fp16.pt --out glm5.2_w8a8.pt

Deploying W8A8 Quantized Models

Once quantized, the model can be loaded by any inference backend that supports 8-bit integer arithmetic. The following examples demonstrate deployment with vLLM-Ascend and SGLang.

vLLM-Ascend Deployment

from vllm import LLM, SamplingParams

# Specify the quantization mode

llm = LLM(
    model="glm5.2_w8a8.pt",
    quantization="w8a8",          # Enables int8 kernels

    device="ascend"
)

sampling_params = SamplingParams(temperature=0.7, top_p=0.9)
prompt = "Explain the benefits of W8A8 quantization for large language models."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].text)

SGLang Deployment

from sglang import Model

model = Model(
    checkpoint="glm5.2_w8a8.pt",
    quant="w8a8",
    backend="ascend"
)

response = model.generate("What is QuaRot and why does it help quantization?")
print(response)

Key Source Files

The quantization implementation and deployment guides are located in the following repository files:

  • example/ascend.md – Documents the complete W8A8 quantization scheme and its benefits for Ascend NPU deployment
  • quant/init.py – Exposes the qua_rot, flex_smoothquant, and ssz_quantizer utilities used in the conversion pipeline
  • skills/glm-master-skill/SKILL.md – Provides model capability context relevant to understanding quantization impact on expert routing

Summary

  • W8A8 quantization reduces GLM-5.2 memory footprint by converting both weights and activations to 8-bit integers
  • The pipeline requires three sequential steps: QuaRot rotation, Flex SmoothQuant smoothing, and SSZ quantization
  • Quantized checkpoints are compatible with vLLM-Ascend, SGLang, and other backends that support 8-bit inference
  • The method specifically preserves accuracy on critical MoE expert paths as described in the Ascend deployment documentation
  • Use the provided quantize.py script to convert FP16 checkpoints to W8A8 format

Frequently Asked Questions

What is W8A8 quantization?

W8A8 quantization refers to a compression scheme where both model weights (W) and activations (A) are represented using 8-bit integers. This reduces memory usage by approximately 50% compared to 16-bit formats and enables faster inference on hardware with optimized 8-bit arithmetic units.

Why combine QuaRot with Flex SmoothQuant?

QuaRot aligns weight distributions to minimize quantization error, while Flex SmoothQuant scales activations and weights to balance dynamic ranges. Together, they address both the weight distribution outliers and activation outliers that typically cause accuracy degradation in low-bit quantization, ensuring the MoE experts maintain their routing precision.

Can I use W8A8 quantized models on GPUs other than Ascend?

While the example/ascend.md documentation focuses on Ascend NPU deployment, the W8A8 checkpoint format uses standard PyTorch int8 tensors. Any inference engine that supports 8-bit integer matrix multiplication (including CUDA-based backends like vLLM with compatible quantization kernels) can theoretically load these checkpoints, though specific kernel optimizations may vary by platform.

How does the quantization affect MoE expert routing?

The GLM-5.2 quantization pipeline specifically targets the critical paths of the Mixture-of-Experts (MoE) architecture. By applying Flex SmoothQuant with calibration-aware smoothing factors, the method preserves the gating network's precision and ensures that expert selection remains accurate, preventing the quantization noise from disrupting the sparse expert routing patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →