# Migrating from GLM-4 to GLM-5 Series Models: A Complete Guide

> Effortlessly migrate from GLM-4 to GLM-5 models. Update your identifier, configure new parameters, and upgrade your inference library for sparse attention. Start using GLM-5 today.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: migration-guide
- Published: 2026-06-19

---

**Migrating from GLM-4 to GLM-5 requires only updating your model identifier to `zai-org/GLM-5` and optionally configuring the new `reasoning_effort` and `enable_thinking` parameters, while upgrading your inference library to support the sparse attention architecture.**

The **GLM-S series** (GLM-5, GLM-5.1, GLM-5.2) in the `zai-org/GLM-5` repository introduces significant architectural improvements over the previous GLM-4 flagship, including DeepSeek Sparse Attention and IndexShare for efficient long-context processing. While the underlying capabilities have expanded to 744B parameters with 28.5T training tokens, the migration remains **backward-compatible** at the API level. You can upgrade existing applications by changing the model name and adjusting a few new generation parameters.

## Architectural Changes in GLM-5

The GLM-5 series implements several production-critical upgrades that improve throughput and reduce inference costs for long sequences.

**Scale and Efficiency.** The model scales to **744B parameters** (with 40B active per token) trained on 28.5T tokens, nearly doubling the capacity of GLM-4's 355B parameter architecture. According to the repository's [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (lines 45-47), this scale combines with **DeepSeek Sparse Attention (DSA)** to reduce FLOPs while maintaining 1M-token context windows.

**IndexShare Optimization.** Unlike GLM-4's independent index per attention block, GLM-5 implements **IndexShare**, where one index is reused every four sparse-attention layers. As documented in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (lines 25-26), this cuts per-token FLOPs by approximately 2.9× at 1M context lengths, significantly accelerating generation for very long sequences.

**Multi-Token Prediction (MTP).** GLM-5.2 improves the MTP layer used in speculative decoding, extending accepted speculative length by up to 20% compared to GLM-4's basic implementation. This enhancement particularly benefits "prefill-decode" pipelines where throughput is critical.

**Reasoning Controls.** The new **`reasoning_effort`** parameter (`max` by default, `high` optional) and **`enable_thinking`** flag provide fine-grained control over the reasoning budget. These parameters, referenced in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (lines 80-81), allow you to trade inference speed for answer quality or disable chain-of-thought reasoning entirely.

**Quantization Support.** Beyond GLM-4's FP16/BF16 support, GLM-5 offers **FP8** and hybrid W8A8 quantization checkpoints (e.g., `GLM-5-FP8`). This reduces memory footprint while preserving accuracy on critical reasoning paths, as noted in the model table (lines 64-68).

## Code-Level Migration Examples

Migration requires minimal code changes across major inference frameworks. Update your model identifier and pass the new reasoning parameters to the generation call.

### HuggingFace Transformers

Install Transformers version 0.5.12 or later, then update the model name and generation parameters:

```python
from transformers import AutoTokenizer, AutoModelForCausalLM

# Update to GLM-5 family identifier

model_name = "zai-org/GLM-5.2"  # or "zai-org/GLM-5", "zai-org/GLM-5.1"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
)

prompt = "Explain the benefits of sparse attention in large language models."
inputs = tokenizer(prompt, return_tensors="pt")

# Generate with new reasoning controls

output = model.generate(
    **inputs,
    max_new_tokens=256,
    reasoning_effort="high",    # Options: "max" (default), "high"

    enable_thinking=True,       # Set to False to disable reasoning

)

print(tokenizer.decode(output[0], skip_special_tokens=True))

```

The `reasoning_effort` and `enable_thinking` arguments are supported from Transformers v0.5.12+ (as indicated in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) lines 76-77).

### vLLM

For production inference with vLLM 0.23.0+, instantiate the model with the new identifier and pass reasoning parameters via `SamplingParams`:

```python
from vllm import LLM, SamplingParams

# Load GLM-5.2 checkpoint

llm = LLM(model="zai-org/GLM-5.2")

sampling_params = SamplingParams(
    max_tokens=256,
    reasoning_effort="high",    # Forwarded to the model's reasoning controller

    enable_thinking=True,
)

prompt = "Write a Python function that computes the factorial of n."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].outputs[0].text)

```

vLLM automatically handles tokenization and batch inference, forwarding the reasoning budget arguments directly to the underlying GLM-5 model.

### SGLang

With SGLang 0.5.13.post1 or later, use the `Engine` class to load GLM-5 and control device placement:

```python
import sglang as sgl

# Initialize engine with GLM-5.2

engine = sgl.Engine(
    model="zai-org/GLM-5.2",
    device="cuda",              # Use "ascend" for NPU deployments

    dtype="bfloat16",          # Use "fp8" for FP8 checkpoints

)

prompt = "Summarize the recent advances in sparse attention."
result = engine.generate(
    prompt,
    max_new_tokens=256,
    reasoning_effort="high",
    enable_thinking=True,
)

print(result.text)

```

SGLang abstracts low-level runtime details while supporting the same reasoning flags as other frameworks.

## Dependencies and Version Requirements

Ensure your inference libraries meet these minimum versions to support the GLM-5 architecture:

- **Transformers**: `>= 0.5.12` — Required for `reasoning_effort` and `enable_thinking` parameter support.
- **vLLM**: `>= 0.23.0` — Needed for optimized sparse attention and batch processing.
- **SGLang**: `>= 0.5.13.post1` — Provides the `Engine` interface with GLM-5 compatibility.
- **KTransformers** (optional): `>= 0.5.12` — For specialized quantization deployments.

Install the base requirements via `pip install -r requirements.txt` from the repository, which lists all dependencies including `transformers`, `vllm`, and `sglang` (referenced in [`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt)).

## Key Repository Files for Reference

When implementing your migration, consult these specific files in the `zai-org/GLM-5` repository:

- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)** (lines 45-47): Contains the core architectural description, scaling parameters, and reasoning-effort API documentation.
- **[`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)** (lines 13-14): Demonstrates GLM-5 integration with the Z-AI skill system for agentic applications.
- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** (lines 1-12): Details Ascend NPU deployment steps for GLM-5.2, including MoE fusion and quantization techniques.
- **[`resources/WECHAT.md`](https://github.com/zai-org/GLM-5/blob/main/resources/WECHAT.md)** (lines 4-5): Provides community support links via WeChat and Discord.

## Step-by-Step Migration Checklist

Follow this checklist to ensure a seamless transition:

1. **Update the model identifier** from GLM-4 to the GLM-5 family (`zai-org/GLM-5`, `zai-org/GLM-5.1`, or `zai-org/GLM-5.2`).
2. **Upgrade inference libraries** to the minimum required versions (Transformers ≥ 0.5.12, vLLM ≥ 0.23.0, or SGLang ≥ 0.5.13.post1).
3. **Configure reasoning parameters** by adding `reasoning_effort="high"` or `enable_thinking=False` to generation calls as needed.
4. **Optionally enable FP8 quantization** by loading a `*-FP8` checkpoint and setting the appropriate `dtype` for reduced memory usage.
5. **Verify hardware compatibility** for CUDA or Ascend NPU deployments using the examples in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md).

## Summary

Migrating from GLM-4 to the GLM-5 series requires minimal API changes while delivering substantial performance improvements:

- Update model identifiers to `zai-org/GLM-5*` and upgrade inference libraries to versions supporting the sparse attention architecture.
- Leverage **DeepSeek Sparse Attention (DSA)** and **IndexShare** for 2.9× FLOP reduction at 1M-token contexts.
- Use the new **`reasoning_effort`** and **`enable_thinking`** parameters to control the trade-off between generation speed and answer quality.
- Deploy **FP8 quantized checkpoints** to reduce memory footprint without sacrificing reasoning capabilities.
- Reference [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) (lines 45-47, 80-81) and [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) for implementation details and hardware-specific optimizations.

## Frequently Asked Questions

### Do I need to change my prompt format when migrating from GLM-4 to GLM-5?

No, the prompt format remains compatible. GLM-5 maintains backward compatibility with GLM-4's chat template and tokenization. You only need to update the model identifier and can optionally add the new `reasoning_effort` parameter to control the depth of chain-of-thought generation.

### What is the difference between GLM-5, GLM-5.1, and GLM-5.2?

The GLM-S series represents progressive iterations. GLM-5.2 specifically includes an improved **Multi-Token Prediction (MTP)** layer that extends speculative decoding length by up to 20% compared to earlier versions. All versions share the same 744B parameter architecture with DeepSeek Sparse Attention, but GLM-5.2 offers the highest throughput for prefill-decode pipelines.

### How do I deploy GLM-5 on Ascend NPU hardware instead of CUDA?

Refer to [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) (lines 1-12) in the repository, which documents specific steps for Ascend NPU deployment. When using SGLang, set `device="ascend"` in the `Engine` initialization, and ensure you have the appropriate MoE fusion and quantization configurations enabled for NPU optimization.

### Can I disable the reasoning capabilities to get faster responses?

Yes. Set `enable_thinking=False` in your generation call to disable the chain-of-thought reasoning. Alternatively, use `reasoning_effort="max"` (the default) for standard reasoning or `"high"` for more thorough analysis. This gives you explicit control over the latency-quality trade-off not available in GLM-4.