# GLM-5 | Z.ai | Knowledge Base | Instagit

GLM-5: From Vibe Coding to Agentic Engineering

GitHub Stars: 4.6k

Repository: https://github.com/zai-org/GLM-5

---

## Articles

### [How to Cache Expert Paths with IndexCache for GLM-5.2 Inference Optimization](/zai-org/GLM-5/how-to-cache-expert-paths-with-indexcache-for-glm-5-2-inference-optimization)

Optimize GLM-5.2 inference speed using IndexCache. This technique caches expert paths for MoE routing, cutting latency and redundant computations for large contexts.

- Tags: performance
- Published: 2026-06-21

### [How GLM-5.2 Closes the Gap with Frontier Models on Agentic Tasks](/zai-org/GLM-5/how-glm-5-2-closes-the-gap-with-frontier-models-on-agentic-tasks)

Discover how GLM-5.2 rivals frontier models on agentic tasks with advanced techniques like IndexShare sparse attention and speculative decoding. Experience efficient 1M-token inference.

- Tags: deep-dive
- Published: 2026-06-21

### [Achieving Best-in-Class Coding Performance on SWE-bench Pro with GLM-5.2](/zai-org/GLM-5/how-to-achieve-best-in-class-coding-performance-on-swe-bench-pro-with-glm-5-2)

Discover how GLM-5.2 achieves 62.1% best-in-class coding performance on SWE-bench Pro using IndexShare sparse attention and MTP speculative decoding for superior accuracy and context.

- Tags: performance
- Published: 2026-06-21

### [How to Optimize Memory Usage with Chunked Prefill in GLM-5.2](/zai-org/GLM-5/how-to-optimize-memory-usage-with-chunked-prefill-in-glm-5-2)

Optimize GLM-5.2 memory usage with Chunked Prefill. Learn how to set flags and tune chunk size to drastically reduce KV-cache memory for large prompts.

- Tags: performance
- Published: 2026-06-21

### [How to Use Unsloth for Fine-Tuning GLM-5.2: A Complete Guide](/zai-org/GLM-5/how-to-use-unsloth-for-fine-tuning-glm-5-2)

Learn to fine-tune GLM-5.2 with Unsloth. This guide details a memory-efficient LoRA workflow for consumer GPUs, preserving IndexShare sparsity and MTP speculative decoding.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Handle Prefill-Decode Disaggregation for Mixed Loads in GLM-5.2](/zai-org/GLM-5/how-to-handle-prefill-decode-disaggregation-for-mixed-loads-in-glm-5-2)

Optimize GLM-5.2 performance with prefill-decode disaggregation. Learn to manage mixed loads and prevent prompt blocking for high-concurrency deployments.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Deploy GLM-5.2 on Ascend NPU with xLLM: Complete Setup Guide](/zai-org/GLM-5/how-to-deploy-glm-5-2-on-ascend-npu-with-xllm)

Deploy GLM-5.2 on Ascend NPU using xLLM. Follow our guide to install drivers, xLLM with Ascend support, and launch the server for efficient W8A8 quantized MoE inference.

- Tags: how-to-guide
- Published: 2026-06-21

### [GLM-5 vs Claude Opus 4.5 on Vending Bench: Open Source Performance Benchmark](/zai-org/GLM-5/how-does-glm-5-compare-to-claude-opus-4-5-on-vending-bench)

Test GLM-5 against Claude Opus 4.5 on Vending Bench. See how GLM-5 ranks #1 among open-source models achieving excellent long-horizon evaluation results.

- Tags: performance
- Published: 2026-06-21

### [How to Switch Between Max and High Thinking Effort Modes in GLM-5.2](/zai-org/GLM-5/how-to-switch-between-max-and-high-thinking-effort-modes-in-glm-5-2)

Learn to switch between max and high thinking effort modes in GLM-5.2. Set reasoning_effort to "high" for lower latency or omit for default max mode. Optimize your GLM-5.2 performance now.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Configure DeepSeek Sparse Attention for GLM-5.2 Long-Context Capacity](/zai-org/GLM-5/how-to-configure-deepseek-sparse-attention-for-glm-5-2-long-context-capacity)

Configure DeepSeek Sparse Attention for GLM-5.2 to achieve 1 million token context with reduced FLOPs. Learn how to set attention_type deepseek for efficient long-context processing.

- Tags: how-to-guide
- Published: 2026-06-21

### [How Slime Async RL Infrastructure Improves Training Throughput for GLM-5.2](/zai-org/GLM-5/how-does-slime-async-rl-infrastructure-improve-training-throughput-for-glm-5-2)

Discover how Slime async RL infrastructure boosts GLM-5.2 training throughput. Learn how decoupling rollouts and updates enables continuous optimization of 744B parameters.

- Tags: performance
- Published: 2026-06-21

### [How Multi-Token Prediction (MTP) Improves Speculative Decoding Acceptance in GLM-5.2](/zai-org/GLM-5/how-multi-token-prediction-mtp-improves-speculative-decoding-acceptance-in-glm-5-2)

Discover how GLM-5.2's Multi-Token Prediction (MTP) boosts speculative decoding acceptance by 20% with attention pre-processing fusion and accelerated block-wise verification.

- Tags: deep-dive
- Published: 2026-06-21

### [How to Integrate KTransformers for Efficient GLM-5.2 Deployment](/zai-org/GLM-5/how-to-integrate-ktransformers-for-efficient-glm-5-2-deployment)

Integrate KTransformers with GLM-5.2 for 2-3x faster long-context inference. Leverage fused KV-cache kernels and IndexShare for efficient deployment.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Reproduce GLM-5.2 Terminal-Bench 2.1 Benchmark Results](/zai-org/GLM-5/how-to-reproduce-glm-5-2-terminal-bench-2-1-benchmark-results)

Reproduce GLM-5.2 Terminal-Bench 2.1 benchmark results by deploying the model with vLLM and running the harness. Achieve the published 81.0 mean score easily.

- Tags: how-to-guide
- Published: 2026-06-21

### [How GLM-5.1 Handles Long-Horizon Agentic Tasks: Architecture and Implementation](/zai-org/GLM-5/how-glm-5-1-handles-long-horizon-agentic-tasks)

Discover how GLM-5.1 tackles long-horizon agentic tasks with sparse attention, massive context windows, and reasoning controls for advanced planning and tool use over thousands of steps.

- Tags: architecture
- Published: 2026-06-21

### [How to Use W8A8 Quantization with QuaRot and Flex SmoothQuant for GLM-5.2](/zai-org/GLM-5/how-to-use-w8a8-quantization-with-quarot-and-flex-smoothquant-for-glm-5-2)

Learn W8A8 quantization for GLM-5.2 using QuaRot and Flex SmoothQuant. Reduce model size while preserving accuracy with this efficient 3-stage pipeline. Optimize your LLM performance.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Implement PD Separation for Throughput Stability in GLM-5.2](/zai-org/GLM-5/how-to-implement-pd-separation-for-throughput-stability-in-glm-5-2)

Stabilize GLM-5.2 throughput using PD separation. Run prefill and decode on separate streams with prefix caching to eliminate latency jitter in high-concurrency inference.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Enable Prefix Caching for GLM-5.2 to Reduce Decode Latency](/zai-org/GLM-5/how-to-enable-prefix-caching-for-glm-5-2-to-reduce-decode-latency)

Reduce GLM-5.2 decode latency by enabling prefix caching. Learn how to configure your serving framework like vLLM or SGLang to reuse KV-caches, skipping the prefill stage for faster token generation.

- Tags: how-to-guide
- Published: 2026-06-21

### [How GLM-5.2's IndexShare Architecture Reduces FLOPs by 2.9× at 1M Context](/zai-org/GLM-5/how-does-glm-5-2s-indexshare-architecture-reduce-flops)

Discover how GLM-5.2's IndexShare architecture cuts FLOPs by 2.9× at 1M context. Learn how reusing attention indexers minimizes computational costs for efficient LLM processing.

- Tags: performance
- Published: 2026-06-21

### [How to Optimize MoE Inference with Mega-Fusion Operators on Ascend NPU](/zai-org/GLM-5/how-to-optimize-moe-inference-with-mega-fusion-operators-on-ascend-npu)

Optimize MoE inference on Ascend NPUs with Mega-Fusion operators. Fuse routing, expert computation, and reduction for up to 2.9x FLOP reduction and faster performance.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Achieve 1 Million Token Long-Context with GLM-5.2: Complete Technical Guide](/zai-org/GLM-5/how-to-achieve-1-million-token-long-context-with-glm-5-2)

Unlock 1 million token context with GLM-5.2. Learn how IndexShare, DSA, and MTP speculative decoding enable massive context windows in vLLM, SGLang, and Transformers.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Use the reasoning_effort Parameter to Control Thinking Levels in GLM-5.2](/zai-org/GLM-5/how-to-use-reasoning-effort-parameter-to-control-thinking-levels-in-glm-5-2)

Control GLM-5.2 thinking levels with reasoning_effort. Set high for deep analysis, omit for max mode, or disable chain-of-thought for faster responses and lower latency.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Configure SGLang for Optimal GLM-5.2 Performance](/zai-org/GLM-5/how-to-configure-sglang-for-optimal-glm-5-2-performance)

Optimize GLM-5.2 performance with SGLang. Learn to configure SGLang by adjusting model flags and batch sizes for superior latency and throughput. Maximize your GPU resources.

- Tags: how-to-guide
- Published: 2026-06-21

### [How to Deploy GLM-5.2 Using vLLM for Production Inference](/zai-org/GLM-5/how-to-deploy-glm-5-2-using-vllm-for-production-inference)

Deploy GLM-5.2 for production inference using vLLM by setting up the OpenAI-compatible server. Learn how to leverage 1M token context and optimized attention kernels for efficient deployments.

- Tags: how-to-guide
- Published: 2026-06-21

### [Migrating from GLM-4 to GLM-5 Series Models: A Complete Guide](/zai-org/GLM-5/migrating-from-glm-4-to-glm-5-series-models)

Effortlessly migrate from GLM-4 to GLM-5 models. Update your identifier, configure new parameters, and upgrade your inference library for sparse attention. Start using GLM-5 today.

- Tags: migration-guide
- Published: 2026-06-19

### [Security Best Practices for Storing ZHIPU_API_KEY in GLM-5](/zai-org/GLM-5/security-best-practices-for-storing-zhihu-api-key)

Learn secure methods for storing your ZHIPU_API_KEY for GLM-5. Avoid committing credentials to source control and follow best practices for environment variable security.

- Tags: best-practices
- Published: 2026-06-19

### [Implementing Streaming Responses with GLM-5 in vLLM: A Complete Guide](/zai-org/GLM-5/implementing-streaming-responses-with-glm-5-in-vllm)

Implement GLM-5 streaming responses in vLLM with this comprehensive guide. Discover how setting "stream": true enables server-sent events for token generation.

- Tags: how-to-guide
- Published: 2026-06-19

### [Best Practices for Production Deployment of GLM-5](/zai-org/GLM-5/best-practices-for-production-deployment-of-glm-5)

Deploy GLM-5 in production efficiently. Optimize inference with vLLM or SGLang, configure hardware settings, and secure APIs. Master GLM-5 production deployment today.

- Tags: best-practices
- Published: 2026-06-19

### [Handling Long-Horizon Agentic Tasks with GLM-5.1: Architecture and Implementation](/zai-org/GLM-5/handling-long-horizon-agentic-tasks-with-glm-5.1)

Discover how GLM-5.1 handles long-horizon agentic tasks with massive context retention and sparse attention. Learn about its architecture and implementation for thousands of iterative tool calls.

- Tags: deep-dive
- Published: 2026-06-19

### [How Reasoning Effort Affects GLM-5 Latency and Quality: A Complete Guide](/zai-org/GLM-5/effect-of-reasoning-effort-on-glm-5-latency-and-quality)

Explore how reasoning effort impacts GLM-5 latency and quality. Learn the trade-offs between high and max modes for complex tasks and production workloads.

- Tags: deep-dive
- Published: 2026-06-19

### [Troubleshooting Slow Inference and OOM Errors with GLM-5: A Complete Optimization Guide](/zai-org/GLM-5/troubleshooting-slow-inference-or-oom-errors-with-glm-5)

Fix GLM-5 slow inference and OOM errors. Optimize performance by using FP16, caching tokens, limiting sequence length, and clearing GPU cache. Get faster results now.

- Tags: tutorial
- Published: 2026-06-19

### [Fine-tuning GLM-5 for Domain-Specific Tasks: A Complete Guide to GLM-S Adaptation](/zai-org/GLM-5/fine-tuning-glm-5-for-domain-specific-tasks)

Master fine-tuning GLM-5 for your domain with GLM-S. Learn how parameter-efficient fine-tuning using LoRA adapters preserves the MoE architecture for optimal results.

- Tags: how-to-guide
- Published: 2026-06-19

### [Minimum GPU Memory Requirements for GLM-5 Inference: Complete Hardware Guide](/zai-org/GLM-5/minimum-gpu-memory-requirements-for-glm-5-inference)

Discover GLM-5 inference GPU memory needs. Learn how BF16 needs 40GB, while FP8 quantization drops requirements to 24GB, enabling consumer hardware use.

- Tags: hardware-guide
- Published: 2026-06-19

### [How Slime Asynchronous RL Infrastructure Boosts Training Throughput in GLM-5](/zai-org/GLM-5/how-slime-asynchronous-rl-infrastructure-boosts-training-throughput)

Discover how Slime asynchronous RL infrastructure boosts GLM-5 training throughput by tenfold. Learn how decoupling trajectory generation accelerates large language model development.

- Tags: performance
- Published: 2026-06-19

### [Optimizing GLM-5 Inference Performance for 1M Token Context: Architecture and Implementation Guide](/zai-org/GLM-5/optimizing-glm-5-inference-performance-for-1m-token-context)

Boost GLM-5 inference performance for 1M token context using IndexShare and DSA kernels. Cut FLOPs by 2.9× while ensuring full context stability. Learn implementation details.

- Tags: performance
- Published: 2026-06-19

### [How DeepSeek Sparse Attention (DSA) Reduces GLM‑5 Deployment Costs](/zai-org/GLM-5/how-deepseek-sparse-attention-dsa-reduces-glm-5-deployment-costs)

Discover how DeepSeek Sparse Attention (DSA) dramatically cuts GLM-5 deployment costs. Learn how it reduces FLOPs and enables inference on 40GB GPUs, saving you money.

- Tags: performance
- Published: 2026-06-19

### [Deploying GLM-5 on Ascend NPU Platform: Optimization Guide for High-Throughput Inference](/zai-org/GLM-5/deploying-glm-5-on-ascend-npu-platform)

Deploy GLM-5 on Ascend NPUs for high-throughput inference. Optimize large language models using vLLM-Ascend, SGLang, or xLLM with fused MoE and sparse attention.

- Tags: how-to-guide
- Published: 2026-06-19

### [Choosing Between FP8 and BF16 Precision Models for GLM-5 Deployment: A Complete Guide](/zai-org/GLM-5/choosing-between-fp8-and-bf16-precision-models-for-glm-5-deployment)

Deploy GLM-5 models effectively. Learn when to choose FP8 for memory efficiency or BF16 for maximum accuracy and hardware compatibility. A complete guide for optimal deployment.

- Tags: deep-dive
- Published: 2026-06-19

### [How the MTP Layer Improves Speculative Decoding in GLM-5](/zai-org/GLM-5/how-does-mtp-layer-improve-speculative-decoding)

Discover how GLM-5's MTP layer boosts speculative decoding, extending acceptance by 20% and cutting FLOPs by 2.9× with efficient sparse-attention. Learn more!

- Tags: internals
- Published: 2026-06-19

### [How to Enable or Disable Thinking Mode in GLM-5: Complete Configuration Guide](/zai-org/GLM-5/enabling-or-disabling-thinking-mode-in-glm-5)

Learn how to enable or disable thinking mode in GLM-5 using enable_thinking and reasoning_effort. Configure GLM-5 for optimal output quality and inference speed.

- Tags: how-to-guide
- Published: 2026-06-19

### [How GLM-5's IndexShare Architecture Reduces FLOPs by 2.9× at 1M Token Context](/zai-org/GLM-5/how-does-glm-5s-indexshare-architecture-reduce-flops)

Discover how GLM-5's IndexShare architecture slashes FLOPs by 2.9× at 1M token context. Learn how it reuses indexers to cut redundant computation and reduce costs.

- Tags: performance
- Published: 2026-06-19

### [Controlling Reasoning Effort in GLM-5 API: max vs high](/zai-org/GLM-5/controlling-reasoning-effort-max-vs-high-in-glm-5-api)

Master GLM-5 API reasoning effort control. Choose between max for speed or high for deeper Chain-of-Thought analysis to optimize your results.

- Tags: how-to-guide
- Published: 2026-06-19

### [Deploying GLM-5 for Production Inference Using vLLM](/zai-org/GLM-5/deploying-glm-5-for-production-inference-using-vllm)

Deploy GLM-5 for production inference with vLLM. Leverage optimized CUDA kernels for DeepSeek Sparse Attention and MoE, supporting 1M token contexts and configurable reasoning budgets.

- Tags: how-to-guide
- Published: 2026-06-19

### [How to Deploy GLM-5 Locally with SGLang: A Complete Setup Guide](/zai-org/GLM-5/how-to-deploy-glm-5-locally-with-sglang)

Deploy GLM-5 locally with SGLang. Follow this complete setup guide to install SGLang, download the checkpoint, and run your model efficiently.

- Tags: getting-started
- Published: 2026-06-19

