# How to Use W8A8 Quantization with QuaRot and Flex SmoothQuant for GLM-5.2

> Learn W8A8 quantization for GLM-5.2 using QuaRot and Flex SmoothQuant. Reduce model size while preserving accuracy with this efficient 3-stage pipeline. Optimize your LLM performance.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-21

---

**GLM-5.2 implements a three-stage W8A8 quantization pipeline that applies QuaRot rotation, Flex SmoothQuant smoothing, and SSZ 8-bit weight quantization to reduce model size while maintaining accuracy on critical MoE expert paths.**

The zai-org/GLM-5 repository provides native support for hybrid W8A8 quantization, enabling efficient deployment on Ascend NPUs and compatible accelerators. This approach compresses both weights and activations to 8-bit integers using a novel combination of rotation and smoothing techniques. According to the Ascend deployment guide in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md), this method preserves full-precision quality on the model's critical pathways while dramatically reducing memory bandwidth requirements.

## The Three-Stage Quantization Pipeline

The W8A8 quantization process in GLM-5.2 follows a specific transformation sequence that must be applied in order:

### Stage 1: QuaRot Pre-processing

**QuaRot** rotates weight matrices into a basis that concentrates energy in the first few dimensions. This rotation makes the weight distributions more amenable to low-bit quantization by reducing the variance across dimensions. In the GLM-5.2 implementation, this step is applied to all linear layers before any scaling or quantization occurs.

### Stage 2: Flex SmoothQuant Smoothing

**Flex SmoothQuant** applies per-channel (or per-tensor) scaling to both activations and weights using a smoothness factor that balances dynamic ranges. This prevents outliers from dominating the quantization step and preserves precision in critical pathways. The smoothing factor is typically set to 0.5, though it can be tuned based on calibration data.

### Stage 3: SSZ Weight Quantization

**SSZ (Symmetric Signed-Zero)** quantization converts the rotated and smoothed weights into 8-bit integers. This symmetric quantization scheme represents both weights and activations in 8-bit format, achieving the final W8A8 configuration. The resulting checkpoint maintains the same architecture but with all linear layer weights stored as `int8` tensors.

## Quantizing a GLM-5.2 Checkpoint

To convert a full-precision GLM-5.2 checkpoint to W8A8 format, use the quantization utilities provided in the repository. The following script applies all three stages sequentially:

```python

# quantize.py

import torch
from glm5.quant import qua_rot, flex_smoothquant, ssz_quantizer

def quantize_glm5ckpt(fp16_path, out_path):
    # Load FP16 checkpoint

    ckpt = torch.load(fp16_path, map_location="cpu")

    # ① QuaRot pre-processing

    ckpt = qua_rot(ckpt)

    # ② Flex SmoothQuant smoothing

    ckpt = flex_smoothquant(ckpt, smooth_factor=0.5)

    # ③ SSZ 8-bit weight quantization

    ckpt = ssz_quantizer(ckpt)

    # Save the W8A8 checkpoint

    torch.save(ckpt, out_path)
    print(f"Quantized checkpoint saved to {out_path}")

if __name__ == "__main__":
    import argparse
    parser = argparse.ArgumentParser()
    parser.add_argument("--fp16", required=True, help="Path to original FP16 checkpoint")
    parser.add_argument("--out", default="glm5.2_w8a8.pt", help="Output path")
    args = parser.parse_args()
    quantize_glm5ckpt(args.fp16, args.out)

```

Execute the script from your terminal:

```bash
python quantize.py --fp16 glm5.2_fp16.pt --out glm5.2_w8a8.pt

```

## Deploying W8A8 Quantized Models

Once quantized, the model can be loaded by any inference backend that supports 8-bit integer arithmetic. The following examples demonstrate deployment with vLLM-Ascend and SGLang.

### vLLM-Ascend Deployment

```python
from vllm import LLM, SamplingParams

# Specify the quantization mode

llm = LLM(
    model="glm5.2_w8a8.pt",
    quantization="w8a8",          # Enables int8 kernels

    device="ascend"
)

sampling_params = SamplingParams(temperature=0.7, top_p=0.9)
prompt = "Explain the benefits of W8A8 quantization for large language models."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].text)

```

### SGLang Deployment

```python
from sglang import Model

model = Model(
    checkpoint="glm5.2_w8a8.pt",
    quant="w8a8",
    backend="ascend"
)

response = model.generate("What is QuaRot and why does it help quantization?")
print(response)

```

## Key Source Files

The quantization implementation and deployment guides are located in the following repository files:

- **[example/ascend.md](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** – Documents the complete W8A8 quantization scheme and its benefits for Ascend NPU deployment
- **[quant/__init__.py](https://github.com/zai-org/GLM-5/blob/main/quant/__init__.py)** – Exposes the `qua_rot`, `flex_smoothquant`, and `ssz_quantizer` utilities used in the conversion pipeline
- **[skills/glm-master-skill/SKILL.md](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)** – Provides model capability context relevant to understanding quantization impact on expert routing

## Summary

- **W8A8 quantization** reduces GLM-5.2 memory footprint by converting both weights and activations to 8-bit integers
- The pipeline requires three sequential steps: **QuaRot rotation**, **Flex SmoothQuant smoothing**, and **SSZ quantization**
- Quantized checkpoints are compatible with vLLM-Ascend, SGLang, and other backends that support 8-bit inference
- The method specifically preserves accuracy on critical MoE expert paths as described in the Ascend deployment documentation
- Use the provided [`quantize.py`](https://github.com/zai-org/GLM-5/blob/main/quantize.py) script to convert FP16 checkpoints to W8A8 format

## Frequently Asked Questions

### What is W8A8 quantization?

W8A8 quantization refers to a compression scheme where both model **weights** (W) and **activations** (A) are represented using 8-bit integers. This reduces memory usage by approximately 50% compared to 16-bit formats and enables faster inference on hardware with optimized 8-bit arithmetic units.

### Why combine QuaRot with Flex SmoothQuant?

**QuaRot** aligns weight distributions to minimize quantization error, while **Flex SmoothQuant** scales activations and weights to balance dynamic ranges. Together, they address both the weight distribution outliers and activation outliers that typically cause accuracy degradation in low-bit quantization, ensuring the MoE experts maintain their routing precision.

### Can I use W8A8 quantized models on GPUs other than Ascend?

While the [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) documentation focuses on Ascend NPU deployment, the W8A8 checkpoint format uses standard PyTorch `int8` tensors. Any inference engine that supports 8-bit integer matrix multiplication (including CUDA-based backends like vLLM with compatible quantization kernels) can theoretically load these checkpoints, though specific kernel optimizations may vary by platform.

### How does the quantization affect MoE expert routing?

The GLM-5.2 quantization pipeline specifically targets the **critical paths** of the Mixture-of-Experts (MoE) architecture. By applying Flex SmoothQuant with calibration-aware smoothing factors, the method preserves the gating network's precision and ensures that expert selection remains accurate, preventing the quantization noise from disrupting the sparse expert routing patterns.