GLM-5
GLM-5: From Vibe Coding to Agentic Engineering
Optimize GLM-5.2 inference speed using IndexCache. This technique caches expert paths for MoE routing, cutting latency and redundant computations for large contexts.
How GLM-5.2 Closes the Gap with Frontier Models on Agentic TasksDiscover how GLM-5.2 rivals frontier models on agentic tasks with advanced techniques like IndexShare sparse attention and speculative decoding. Experience efficient 1M-token inference.
Achieving Best-in-Class Coding Performance on SWE-bench Pro with GLM-5.2Discover how GLM-5.2 achieves 62.1% best-in-class coding performance on SWE-bench Pro using IndexShare sparse attention and MTP speculative decoding for superior accuracy and context.
How to Optimize Memory Usage with Chunked Prefill in GLM-5.2Optimize GLM-5.2 memory usage with Chunked Prefill. Learn how to set flags and tune chunk size to drastically reduce KV-cache memory for large prompts.
How to Use Unsloth for Fine-Tuning GLM-5.2: A Complete GuideLearn to fine-tune GLM-5.2 with Unsloth. This guide details a memory-efficient LoRA workflow for consumer GPUs, preserving IndexShare sparsity and MTP speculative decoding.
How to Handle Prefill-Decode Disaggregation for Mixed Loads in GLM-5.2Optimize GLM-5.2 performance with prefill-decode disaggregation. Learn to manage mixed loads and prevent prompt blocking for high-concurrency deployments.
How to Deploy GLM-5.2 on Ascend NPU with xLLM: Complete Setup GuideDeploy GLM-5.2 on Ascend NPU using xLLM. Follow our guide to install drivers, xLLM with Ascend support, and launch the server for efficient W8A8 quantized MoE inference.
GLM-5 vs Claude Opus 4.5 on Vending Bench: Open Source Performance BenchmarkTest GLM-5 against Claude Opus 4.5 on Vending Bench. See how GLM-5 ranks #1 among open-source models achieving excellent long-horizon evaluation results.
How to Switch Between Max and High Thinking Effort Modes in GLM-5.2Learn to switch between max and high thinking effort modes in GLM-5.2. Set reasoning_effort to "high" for lower latency or omit for default max mode. Optimize your GLM-5.2 performance now.
How to Configure DeepSeek Sparse Attention for GLM-5.2 Long-Context CapacityConfigure DeepSeek Sparse Attention for GLM-5.2 to achieve 1 million token context with reduced FLOPs. Learn how to set attention_type deepseek for efficient long-context processing.
How Slime Async RL Infrastructure Improves Training Throughput for GLM-5.2Discover how Slime async RL infrastructure boosts GLM-5.2 training throughput. Learn how decoupling rollouts and updates enables continuous optimization of 744B parameters.
How Multi-Token Prediction (MTP) Improves Speculative Decoding Acceptance in GLM-5.2Discover how GLM-5.2's Multi-Token Prediction (MTP) boosts speculative decoding acceptance by 20% with attention pre-processing fusion and accelerated block-wise verification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →