Migrating from GLM-4 to GLM-5 Series Models: A Complete Guide

Migrating from GLM-4 to GLM-5 requires only updating your model identifier to zai-org/GLM-5 and optionally configuring the new reasoning_effort and enable_thinking parameters, while upgrading your inference library to support the sparse attention architecture.

The GLM-S series (GLM-5, GLM-5.1, GLM-5.2) in the zai-org/GLM-5 repository introduces significant architectural improvements over the previous GLM-4 flagship, including DeepSeek Sparse Attention and IndexShare for efficient long-context processing. While the underlying capabilities have expanded to 744B parameters with 28.5T training tokens, the migration remains backward-compatible at the API level. You can upgrade existing applications by changing the model name and adjusting a few new generation parameters.

Architectural Changes in GLM-5

The GLM-5 series implements several production-critical upgrades that improve throughput and reduce inference costs for long sequences.

Scale and Efficiency. The model scales to 744B parameters (with 40B active per token) trained on 28.5T tokens, nearly doubling the capacity of GLM-4's 355B parameter architecture. According to the repository's README.md (lines 45-47), this scale combines with DeepSeek Sparse Attention (DSA) to reduce FLOPs while maintaining 1M-token context windows.

IndexShare Optimization. Unlike GLM-4's independent index per attention block, GLM-5 implements IndexShare, where one index is reused every four sparse-attention layers. As documented in README.md (lines 25-26), this cuts per-token FLOPs by approximately 2.9× at 1M context lengths, significantly accelerating generation for very long sequences.

Multi-Token Prediction (MTP). GLM-5.2 improves the MTP layer used in speculative decoding, extending accepted speculative length by up to 20% compared to GLM-4's basic implementation. This enhancement particularly benefits "prefill-decode" pipelines where throughput is critical.

Reasoning Controls. The new reasoning_effort parameter (max by default, high optional) and enable_thinking flag provide fine-grained control over the reasoning budget. These parameters, referenced in README.md (lines 80-81), allow you to trade inference speed for answer quality or disable chain-of-thought reasoning entirely.

Quantization Support. Beyond GLM-4's FP16/BF16 support, GLM-5 offers FP8 and hybrid W8A8 quantization checkpoints (e.g., GLM-5-FP8). This reduces memory footprint while preserving accuracy on critical reasoning paths, as noted in the model table (lines 64-68).

Code-Level Migration Examples

Migration requires minimal code changes across major inference frameworks. Update your model identifier and pass the new reasoning parameters to the generation call.

HuggingFace Transformers

Install Transformers version 0.5.12 or later, then update the model name and generation parameters:

from transformers import AutoTokenizer, AutoModelForCausalLM

# Update to GLM-5 family identifier

model_name = "zai-org/GLM-5.2"  # or "zai-org/GLM-5", "zai-org/GLM-5.1"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
)

prompt = "Explain the benefits of sparse attention in large language models."
inputs = tokenizer(prompt, return_tensors="pt")

# Generate with new reasoning controls

output = model.generate(
    **inputs,
    max_new_tokens=256,
    reasoning_effort="high",    # Options: "max" (default), "high"

    enable_thinking=True,       # Set to False to disable reasoning

)

print(tokenizer.decode(output[0], skip_special_tokens=True))

The reasoning_effort and enable_thinking arguments are supported from Transformers v0.5.12+ (as indicated in README.md lines 76-77).

vLLM

For production inference with vLLM 0.23.0+, instantiate the model with the new identifier and pass reasoning parameters via SamplingParams:

from vllm import LLM, SamplingParams

# Load GLM-5.2 checkpoint

llm = LLM(model="zai-org/GLM-5.2")

sampling_params = SamplingParams(
    max_tokens=256,
    reasoning_effort="high",    # Forwarded to the model's reasoning controller

    enable_thinking=True,
)

prompt = "Write a Python function that computes the factorial of n."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].outputs[0].text)

vLLM automatically handles tokenization and batch inference, forwarding the reasoning budget arguments directly to the underlying GLM-5 model.

SGLang

With SGLang 0.5.13.post1 or later, use the Engine class to load GLM-5 and control device placement:

import sglang as sgl

# Initialize engine with GLM-5.2

engine = sgl.Engine(
    model="zai-org/GLM-5.2",
    device="cuda",              # Use "ascend" for NPU deployments

    dtype="bfloat16",          # Use "fp8" for FP8 checkpoints

)

prompt = "Summarize the recent advances in sparse attention."
result = engine.generate(
    prompt,
    max_new_tokens=256,
    reasoning_effort="high",
    enable_thinking=True,
)

print(result.text)

SGLang abstracts low-level runtime details while supporting the same reasoning flags as other frameworks.

Dependencies and Version Requirements

Ensure your inference libraries meet these minimum versions to support the GLM-5 architecture:

  • Transformers: >= 0.5.12 — Required for reasoning_effort and enable_thinking parameter support.
  • vLLM: >= 0.23.0 — Needed for optimized sparse attention and batch processing.
  • SGLang: >= 0.5.13.post1 — Provides the Engine interface with GLM-5 compatibility.
  • KTransformers (optional): >= 0.5.12 — For specialized quantization deployments.

Install the base requirements via pip install -r requirements.txt from the repository, which lists all dependencies including transformers, vllm, and sglang (referenced in requirements.txt).

Key Repository Files for Reference

When implementing your migration, consult these specific files in the zai-org/GLM-5 repository:

  • README.md (lines 45-47): Contains the core architectural description, scaling parameters, and reasoning-effort API documentation.
  • skills/glm-master-skill/SKILL.md (lines 13-14): Demonstrates GLM-5 integration with the Z-AI skill system for agentic applications.
  • example/ascend.md (lines 1-12): Details Ascend NPU deployment steps for GLM-5.2, including MoE fusion and quantization techniques.
  • resources/WECHAT.md (lines 4-5): Provides community support links via WeChat and Discord.

Step-by-Step Migration Checklist

Follow this checklist to ensure a seamless transition:

  1. Update the model identifier from GLM-4 to the GLM-5 family (zai-org/GLM-5, zai-org/GLM-5.1, or zai-org/GLM-5.2).
  2. Upgrade inference libraries to the minimum required versions (Transformers ≥ 0.5.12, vLLM ≥ 0.23.0, or SGLang ≥ 0.5.13.post1).
  3. Configure reasoning parameters by adding reasoning_effort="high" or enable_thinking=False to generation calls as needed.
  4. Optionally enable FP8 quantization by loading a *-FP8 checkpoint and setting the appropriate dtype for reduced memory usage.
  5. Verify hardware compatibility for CUDA or Ascend NPU deployments using the examples in example/ascend.md.

Summary

Migrating from GLM-4 to the GLM-5 series requires minimal API changes while delivering substantial performance improvements:

  • Update model identifiers to zai-org/GLM-5* and upgrade inference libraries to versions supporting the sparse attention architecture.
  • Leverage DeepSeek Sparse Attention (DSA) and IndexShare for 2.9× FLOP reduction at 1M-token contexts.
  • Use the new reasoning_effort and enable_thinking parameters to control the trade-off between generation speed and answer quality.
  • Deploy FP8 quantized checkpoints to reduce memory footprint without sacrificing reasoning capabilities.
  • Reference README.md (lines 45-47, 80-81) and example/ascend.md for implementation details and hardware-specific optimizations.

Frequently Asked Questions

Do I need to change my prompt format when migrating from GLM-4 to GLM-5?

No, the prompt format remains compatible. GLM-5 maintains backward compatibility with GLM-4's chat template and tokenization. You only need to update the model identifier and can optionally add the new reasoning_effort parameter to control the depth of chain-of-thought generation.

What is the difference between GLM-5, GLM-5.1, and GLM-5.2?

The GLM-S series represents progressive iterations. GLM-5.2 specifically includes an improved Multi-Token Prediction (MTP) layer that extends speculative decoding length by up to 20% compared to earlier versions. All versions share the same 744B parameter architecture with DeepSeek Sparse Attention, but GLM-5.2 offers the highest throughput for prefill-decode pipelines.

How do I deploy GLM-5 on Ascend NPU hardware instead of CUDA?

Refer to example/ascend.md (lines 1-12) in the repository, which documents specific steps for Ascend NPU deployment. When using SGLang, set device="ascend" in the Engine initialization, and ensure you have the appropriate MoE fusion and quantization configurations enabled for NPU optimization.

Can I disable the reasoning capabilities to get faster responses?

Yes. Set enable_thinking=False in your generation call to disable the chain-of-thought reasoning. Alternatively, use reasoning_effort="max" (the default) for standard reasoning or "high" for more thorough analysis. This gives you explicit control over the latency-quality trade-off not available in GLM-4.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →