How to Reproduce GLM-5.2 Terminal-Bench 2.1 Benchmark Results
Deploy GLM-5.2 with the default "max" reasoning effort via vLLM or any supported backend, then run the Terminal-Bench 2.1 harness against the OpenAI-compatible endpoint to achieve the published 81.0 mean score.
GLM-5.2 is a 744-billion-parameter open-source language model that achieves state-of-the-art results on the Terminal-Bench 2.1 suite. According to the zai-org/GLM-5 repository, this model scores 81.0 compared to GLM-5.1's 62.0 on terminal command reasoning tasks. Reproducing these results requires deploying the model with specific inference settings and executing the external benchmark harness.
Prerequisites and Repository Setup
Start by cloning the official repository and installing the pinned dependencies required for model serving.
Clone the GLM-5 Repository
The repository contains deployment documentation, configuration files, and the requirements.txt specifying compatible versions of inference backends.
git clone https://github.com/zai-org/GLM-5.git
cd GLM-5
Install Server Dependencies
Install the exact dependency versions specified in requirements.txt, which pins torch, vllm, and the GLM-5 wheels to tested versions.
pip install -r requirements.txt
Deploy GLM-5.2 with Maximum Reasoning Effort
The reasoning_effort flag controls the compute budget allocated to the model. According to README.md at line 81, the default "max" setting must be used to reproduce the benchmark results. Using "high" instead will produce divergent scores.
Launch the vLLM Server
vLLM 0.23.0+ is the recommended backend. The following command exposes an OpenAI-compatible HTTP API at http://localhost:8000/v1/chat/completions:
vllm serve Zai-Org/GLM-5.2 --port 8000 --dtype bfloat16 --max-model-len 1048576
Alternative Inference Backends
While vLLM is recommended, the repository supports several backends that expose the same OpenAI-compatible API:
- SGLang
- Transformers
- K-Transformers
- Unsloth
- Ascend-NPU (see
example/ascend.mdfor NPU-specific instructions)
Install and Configure Terminal-Bench 2.1
Terminal-Bench is an external evaluation repository that tests models on realistic terminal commands including file edits, git operations, and package installations.
Clone and install the benchmark harness:
git clone https://github.com/terminal-bench/terminal-bench.git
cd terminal-bench
pip install -e .
The harness sends task sequences to your model endpoint and scores correctness, resource usage, and final task success.
Execute the Benchmark Evaluation
With the GLM-5.2 server running, invoke the evaluation script pointing to your local endpoint. The --reasoning-effort max parameter is critical to match the published 81.0 score documented in README.md at line 29.
python -m terminal_bench.run \
--api-url http://localhost:8000/v1/chat/completions \
--model glm-5.2 \
--benchmark-version 2.1 \
--reasoning-effort max
The script downloads the benchmark tasks, executes them against your deployed model, and aggregates the results.
Verify Your Results
Upon completion, the evaluation script outputs a line similar to:
Overall score: 81.0
This value matches the 81.0 score reported in the official documentation (README.md line 29), confirming successful reproduction of the GLM-5.2 Terminal-Bench 2.1 results. The significant improvement over GLM-5.1's 62.0 score demonstrates the enhanced terminal reasoning capabilities of the 744-billion-parameter model.
Optional Automation Scripts
For programmatic access or single-task debugging, use these snippets against the running server.
Python Client Example
import openai
client = openai.ChatCompletion.create(
api_base="http://localhost:8000/v1",
model="glm-5.2",
messages=[{"role": "user", "content": "List the files in the current directory"}],
temperature=0.0,
)
print(client.choices[0].message.content)
Shell Wrapper for Single Tasks
#!/usr/bin/env bash
TASK_ID=01 # change to any task from the 2.1 suite
python -m terminal_bench.run \
--api-url http://localhost:8000/v1/chat/completions \
--model glm-5.2 \
--benchmark-version 2.1 \
--task $TASK_ID \
--reasoning-effort max
Key Repository Files for Reference
Understanding these files helps debug reproduction issues:
README.md– Contains the benchmark claim (Terminal-Bench 2.1: 81.0) at line 29 and thereasoning_effortdocumentation at line 81.requirements.txt– Lists exact dependency versions for inference backends.example/ascend.md– Deployment instructions for Ascend NPU hardware.skills/glm-master-skill/SKILL.md– Describes the agentic skill interface used by GLM-5 during terminal tasks.
Summary
To successfully reproduce GLM-5.2 Terminal-Bench 2.1 benchmark results:
- Deploy the 744-billion-parameter model using vLLM or any supported backend with the default "max" reasoning effort.
- Ensure the server exposes an OpenAI-compatible API at a known endpoint.
- Install the external Terminal-Bench evaluation suite and execute with
--reasoning-effort max. - Verify the final output shows 81.0, matching the score documented in
README.mdline 29.
Frequently Asked Questions
What is the difference between "max" and "high" reasoning effort?
The reasoning_effort parameter controls the compute budget allocated to the model. According to README.md line 81, "max" is the default setting used for official benchmarks, while "high" uses a reduced budget. Running the evaluation with "high" instead of "max" will produce scores that diverge from the published 81.0 result.
Can I use a different inference backend besides vLLM?
Yes. The zai-org/GLM-5 repository officially supports SGLang, Transformers, K-Transformers, Unsloth, and Ascend-NPU. All backends expose a standard OpenAI-compatible HTTP API that Terminal-Bench can target, provided you maintain the "max" reasoning effort configuration.
Where is the 81.0 benchmark score documented in the repository?
The 81.0 Terminal-Bench 2.1 score is explicitly claimed in README.md at line 29, which highlights GLM-5.2's performance compared to GLM-5.1's 62.0. This documentation also references the resources/ directory containing visual summaries of these results.
Do I need specific hardware to run GLM-5.2 for this benchmark?
While the raw analysis does not specify minimum hardware requirements, deploying a 744-billion-parameter model typically requires multiple high-memory GPUs or NPUs. The example/ascend.md file provides specific guidance for Ascend NPU deployment, while standard CUDA setups should consult the requirements.txt for vLLM compatibility and memory optimization flags like --dtype bfloat16.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →