How to Reproduce GLM-5.2 Terminal-Bench 2.1 Benchmark Results

Deploy GLM-5.2 with the default "max" reasoning effort via vLLM or any supported backend, then run the Terminal-Bench 2.1 harness against the OpenAI-compatible endpoint to achieve the published 81.0 mean score.

GLM-5.2 is a 744-billion-parameter open-source language model that achieves state-of-the-art results on the Terminal-Bench 2.1 suite. According to the zai-org/GLM-5 repository, this model scores 81.0 compared to GLM-5.1's 62.0 on terminal command reasoning tasks. Reproducing these results requires deploying the model with specific inference settings and executing the external benchmark harness.

Prerequisites and Repository Setup

Start by cloning the official repository and installing the pinned dependencies required for model serving.

Clone the GLM-5 Repository

The repository contains deployment documentation, configuration files, and the requirements.txt specifying compatible versions of inference backends.

git clone https://github.com/zai-org/GLM-5.git
cd GLM-5

Install Server Dependencies

Install the exact dependency versions specified in requirements.txt, which pins torch, vllm, and the GLM-5 wheels to tested versions.

pip install -r requirements.txt

Deploy GLM-5.2 with Maximum Reasoning Effort

The reasoning_effort flag controls the compute budget allocated to the model. According to README.md at line 81, the default "max" setting must be used to reproduce the benchmark results. Using "high" instead will produce divergent scores.

Launch the vLLM Server

vLLM 0.23.0+ is the recommended backend. The following command exposes an OpenAI-compatible HTTP API at http://localhost:8000/v1/chat/completions:

vllm serve Zai-Org/GLM-5.2 --port 8000 --dtype bfloat16 --max-model-len 1048576

Alternative Inference Backends

While vLLM is recommended, the repository supports several backends that expose the same OpenAI-compatible API:

  • SGLang
  • Transformers
  • K-Transformers
  • Unsloth
  • Ascend-NPU (see example/ascend.md for NPU-specific instructions)

Install and Configure Terminal-Bench 2.1

Terminal-Bench is an external evaluation repository that tests models on realistic terminal commands including file edits, git operations, and package installations.

Clone and install the benchmark harness:

git clone https://github.com/terminal-bench/terminal-bench.git
cd terminal-bench
pip install -e .

The harness sends task sequences to your model endpoint and scores correctness, resource usage, and final task success.

Execute the Benchmark Evaluation

With the GLM-5.2 server running, invoke the evaluation script pointing to your local endpoint. The --reasoning-effort max parameter is critical to match the published 81.0 score documented in README.md at line 29.

python -m terminal_bench.run \
    --api-url http://localhost:8000/v1/chat/completions \
    --model glm-5.2 \
    --benchmark-version 2.1 \
    --reasoning-effort max

The script downloads the benchmark tasks, executes them against your deployed model, and aggregates the results.

Verify Your Results

Upon completion, the evaluation script outputs a line similar to:


Overall score: 81.0

This value matches the 81.0 score reported in the official documentation (README.md line 29), confirming successful reproduction of the GLM-5.2 Terminal-Bench 2.1 results. The significant improvement over GLM-5.1's 62.0 score demonstrates the enhanced terminal reasoning capabilities of the 744-billion-parameter model.

Optional Automation Scripts

For programmatic access or single-task debugging, use these snippets against the running server.

Python Client Example

import openai

client = openai.ChatCompletion.create(
    api_base="http://localhost:8000/v1",
    model="glm-5.2",
    messages=[{"role": "user", "content": "List the files in the current directory"}],
    temperature=0.0,
)

print(client.choices[0].message.content)

Shell Wrapper for Single Tasks

#!/usr/bin/env bash
TASK_ID=01  # change to any task from the 2.1 suite

python -m terminal_bench.run \
    --api-url http://localhost:8000/v1/chat/completions \
    --model glm-5.2 \
    --benchmark-version 2.1 \
    --task $TASK_ID \
    --reasoning-effort max

Key Repository Files for Reference

Understanding these files helps debug reproduction issues:

  • README.md – Contains the benchmark claim (Terminal-Bench 2.1: 81.0) at line 29 and the reasoning_effort documentation at line 81.
  • requirements.txt – Lists exact dependency versions for inference backends.
  • example/ascend.md – Deployment instructions for Ascend NPU hardware.
  • skills/glm-master-skill/SKILL.md – Describes the agentic skill interface used by GLM-5 during terminal tasks.

Summary

To successfully reproduce GLM-5.2 Terminal-Bench 2.1 benchmark results:

  • Deploy the 744-billion-parameter model using vLLM or any supported backend with the default "max" reasoning effort.
  • Ensure the server exposes an OpenAI-compatible API at a known endpoint.
  • Install the external Terminal-Bench evaluation suite and execute with --reasoning-effort max.
  • Verify the final output shows 81.0, matching the score documented in README.md line 29.

Frequently Asked Questions

What is the difference between "max" and "high" reasoning effort?

The reasoning_effort parameter controls the compute budget allocated to the model. According to README.md line 81, "max" is the default setting used for official benchmarks, while "high" uses a reduced budget. Running the evaluation with "high" instead of "max" will produce scores that diverge from the published 81.0 result.

Can I use a different inference backend besides vLLM?

Yes. The zai-org/GLM-5 repository officially supports SGLang, Transformers, K-Transformers, Unsloth, and Ascend-NPU. All backends expose a standard OpenAI-compatible HTTP API that Terminal-Bench can target, provided you maintain the "max" reasoning effort configuration.

Where is the 81.0 benchmark score documented in the repository?

The 81.0 Terminal-Bench 2.1 score is explicitly claimed in README.md at line 29, which highlights GLM-5.2's performance compared to GLM-5.1's 62.0. This documentation also references the resources/ directory containing visual summaries of these results.

Do I need specific hardware to run GLM-5.2 for this benchmark?

While the raw analysis does not specify minimum hardware requirements, deploying a 744-billion-parameter model typically requires multiple high-memory GPUs or NPUs. The example/ascend.md file provides specific guidance for Ascend NPU deployment, while standard CUDA setups should consult the requirements.txt for vLLM compatibility and memory optimization flags like --dtype bfloat16.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →