GLM-5 vs Claude Opus 4.5 on Vending Bench: Open Source Performance Benchmark
GLM-5 achieves a final balance of $4,432 on the Vending Bench 2 long-horizon evaluation, placing it within a few hundred dollars of Claude Opus 4.5 and securing the #1 rank among all open-source models.
The zai-org/GLM-5 repository (also referred to as the GLM-S family) represents a significant advancement in open-source language model capabilities on agentic benchmarks. When evaluated on the Vending Bench 2 simulation—a year-long test of resource management and strategic planning—GLM-5 narrows the gap to closed-source alternatives like Claude Opus 4.5 while maintaining full open-source accessibility.
Vending Bench 2 Performance Comparison
Final Account Balance Results
According to README.md at line 55, GLM-5 finishes the Vending Bench 2 simulation with a final account balance of $4,432. This result is very close to the performance reported for Claude Opus 4.5 (referred to as Claude Opus 4.S in the original paper), which remains the strongest closed-source baseline on this benchmark.
Open-Source Ranking
With this score, GLM-5 ranks #1 among all open-source models on the Vending Bench long-horizon evaluation. The repository documentation at resources/vending_bench.png provides a visual illustration of this placement, demonstrating that sparse attention architectures and advanced post-training can match proprietary systems on complex planning tasks.
Architectural Innovations Behind the Results
DeepSeek Sparse Attention (DSA) and IndexShare
GLM-5 utilizes DeepSeek Sparse Attention (DSA), a kernel-level optimization that reduces the computational cost of attending to long contexts while preserving full-context capability. Combined with IndexShare—described in arXiv 2603.12201—which shares a single indexer across every four sparse-attention layers, GLM-5 cuts per-token FLOPs by 2.9× at 1 million context length. This efficiency leaves compute budget available for the reasoning and planning steps critical to vending-machine operations.
MTP Speculative Decoding
The model implements MTP speculative decoding, extending the accepted speculative length by up to 20%. This reduces the number of forward passes required when generating long-form plans, such as weekly inventory ordering strategies, directly improving latency and throughput during the simulation.
RL Infrastructure "Slime"
As implemented in the source code, GLM-5 leverages an asynchronous RL infrastructure codenamed "slime" that accelerates post-training reinforcement learning fine-tuning. This infrastructure enables more policy-iteration steps, producing a model that better balances short-term profit against long-term sustainability—a core challenge in the vending-machine task where operators must manage inventory over thousands of simulation steps.
Reasoning-Effort Control
GLM-5 exposes a reasoning_effort parameter accepting "max" (default) or "high" values. When invoked with reasoning_effort="max", the model achieves the $4,432 benchmark score used for the official comparison. Switching to "high" can extract additional profit margin, trading latency for deeper chain-of-thought analysis.
Agentic Design and Tool Use
As documented in skills/glm-master-skill/SKILL.md, GLM-5 supports agentic behaviors including tool invocation, self-critique, and iterative replanning. These capabilities mirror the real-world workflow of vending-machine operators who must continuously monitor sales data, restock inventory, and adjust pricing based on market conditions.
Reproducing the Benchmark Results
The following Python snippet demonstrates how to invoke GLM-5 via the Z.ai API with the exact configuration used for the Vending Bench reproduction:
import requests
import json
# Configure API endpoint
API_URL = "https://api.z.ai/v1/chat/completions"
HEADERS = {
"Authorization": "Bearer <YOUR_API_KEY>",
"Content-Type": "application/json"
}
# Construct vending operator prompt
messages = [
{"role": "system", "content": "You are a vending-machine operator."},
{"role": "user", "content": "Given the current inventory and sales trend, how many cans of each product should I restock for the next week?"}
]
# Payload with max reasoning effort (default for benchmark)
payload = {
"model": "glm-5",
"messages": messages,
"reasoning_effort": "max",
"max_tokens": 1024,
"temperature": 0.2
}
# Execute request
resp = requests.post(API_URL, headers=HEADERS, json=payload)
result = resp.json()
print("Decision:", result["choices"][0]["message"]["content"])
To run with higher reasoning depth for additional optimization:
payload["reasoning_effort"] = "high"
Key Source Files and Documentation
README.md(line 55): Contains the official Vending Bench 2 result and direct comparison to Claude Opus 4.5resources/vending_bench.png: Visual illustration of GLM-5's benchmark placement relative to other modelsskills/glm-master-skill/SKILL.md: Documents the skill-level API powering agentic behaviors and tool useREADME_zh.md(line 55): Chinese localization of the benchmark resultsexample/ascend.md: Deployment guide for running GLM-5 on Ascend NPU hardware
Summary
- GLM-5 achieves a $4,432 final balance on Vending Bench 2, ranking #1 among open-source models
- The model operates within a few hundred dollars of Claude Opus 4.5, the leading closed-source baseline
- IndexShare and DeepSeek Sparse Attention reduce per-token FLOPs by 2.9× at 1M context, enabling efficient long-horizon planning
- MTP speculative decoding and asynchronous RL infrastructure ("slime") optimize inference speed and policy quality
- The
reasoning_effortparameter allows explicit trade-offs between latency and planning depth - Full reproduction code and deployment guides are available in the zai-org/GLM-5 repository
Frequently Asked Questions
What is the exact performance gap between GLM-5 and Claude Opus 4.5 on Vending Bench?
According to the source analysis, GLM-5 finishes the simulation with a final account balance of $4,432, which is described as "very close" to Claude Opus 4.5's result. The gap is limited to a few hundred dollars, making GLM-5 the top-performing open-source alternative while maintaining competitive proximity to the closed-source leader.
How does GLM-5 maintain a 1-million-token context without performance degradation?
GLM-5 implements DeepSeek Sparse Attention (DSA) combined with IndexShare architecture, which shares a single indexer across every four sparse-attention layers. This design reduces per-token FLOPs by 2.9× at 1 million context length, allowing the model to maintain the full context window throughout the year-long simulation without exploding memory or latency costs.
Can I reproduce the Vending Bench results using the public API?
Yes. The repository provides the exact API configuration used for benchmarking. By setting reasoning_effort to "max" (the default) and querying the glm-5 model endpoint as shown in README.md and the provided code examples, you can reproduce the $4,432 final balance result. The example/ascend.md file also documents local deployment on Ascend NPU hardware for offline reproduction.
What makes GLM-5 suitable for long-horizon agentic tasks like Vending Bench?
Beyond sparse attention efficiency, GLM-5 incorporates asynchronous RL infrastructure for improved post-training, MTP speculative decoding for faster plan generation, and agentic design patterns documented in skills/glm-master-skill/SKILL.md. These features enable tool use, self-critique, and iterative replanning—capabilities essential for managing the thousands of decision steps required in the vending-machine simulation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →