# How to Reproduce GLM-5.2 Terminal-Bench 2.1 Benchmark Results

> Reproduce GLM-5.2 Terminal-Bench 2.1 benchmark results by deploying the model with vLLM and running the harness. Achieve the published 81.0 mean score easily.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-21

---

**Deploy GLM-5.2 with the default "max" reasoning effort via vLLM or any supported backend, then run the Terminal-Bench 2.1 harness against the OpenAI-compatible endpoint to achieve the published 81.0 mean score.**

GLM-5.2 is a 744-billion-parameter open-source language model that achieves state-of-the-art results on the Terminal-Bench 2.1 suite. According to the `zai-org/GLM-5` repository, this model scores 81.0 compared to GLM-5.1's 62.0 on terminal command reasoning tasks. Reproducing these results requires deploying the model with specific inference settings and executing the external benchmark harness.

## Prerequisites and Repository Setup

Start by cloning the official repository and installing the pinned dependencies required for model serving.

### Clone the GLM-5 Repository

The repository contains deployment documentation, configuration files, and the [`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt) specifying compatible versions of inference backends.

```bash
git clone https://github.com/zai-org/GLM-5.git
cd GLM-5

```

### Install Server Dependencies

Install the exact dependency versions specified in [`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt), which pins `torch`, `vllm`, and the GLM-5 wheels to tested versions.

```bash
pip install -r requirements.txt

```

## Deploy GLM-5.2 with Maximum Reasoning Effort

The **reasoning_effort** flag controls the compute budget allocated to the model. According to [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) at line 81, the default "max" setting must be used to reproduce the benchmark results. Using "high" instead will produce divergent scores.

### Launch the vLLM Server

vLLM 0.23.0+ is the recommended backend. The following command exposes an OpenAI-compatible HTTP API at `http://localhost:8000/v1/chat/completions`:

```bash
vllm serve Zai-Org/GLM-5.2 --port 8000 --dtype bfloat16 --max-model-len 1048576

```

### Alternative Inference Backends

While vLLM is recommended, the repository supports several backends that expose the same OpenAI-compatible API:
- **SGLang**
- **Transformers**
- **K-Transformers**
- **Unsloth**
- **Ascend-NPU** (see [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) for NPU-specific instructions)

## Install and Configure Terminal-Bench 2.1

Terminal-Bench is an external evaluation repository that tests models on realistic terminal commands including file edits, git operations, and package installations.

Clone and install the benchmark harness:

```bash
git clone https://github.com/terminal-bench/terminal-bench.git
cd terminal-bench
pip install -e .

```

The harness sends task sequences to your model endpoint and scores correctness, resource usage, and final task success.

## Execute the Benchmark Evaluation

With the GLM-5.2 server running, invoke the evaluation script pointing to your local endpoint. The `--reasoning-effort max` parameter is critical to match the published 81.0 score documented in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) at line 29.

```bash
python -m terminal_bench.run \
    --api-url http://localhost:8000/v1/chat/completions \
    --model glm-5.2 \
    --benchmark-version 2.1 \
    --reasoning-effort max

```

The script downloads the benchmark tasks, executes them against your deployed model, and aggregates the results.

## Verify Your Results

Upon completion, the evaluation script outputs a line similar to:

```

Overall score: 81.0

```

This value matches the **81.0** score reported in the official documentation ([`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) line 29), confirming successful reproduction of the GLM-5.2 Terminal-Bench 2.1 results. The significant improvement over GLM-5.1's 62.0 score demonstrates the enhanced terminal reasoning capabilities of the 744-billion-parameter model.

## Optional Automation Scripts

For programmatic access or single-task debugging, use these snippets against the running server.

### Python Client Example

```python
import openai

client = openai.ChatCompletion.create(
    api_base="http://localhost:8000/v1",
    model="glm-5.2",
    messages=[{"role": "user", "content": "List the files in the current directory"}],
    temperature=0.0,
)

print(client.choices[0].message.content)

```

### Shell Wrapper for Single Tasks

```bash
#!/usr/bin/env bash
TASK_ID=01  # change to any task from the 2.1 suite

python -m terminal_bench.run \
    --api-url http://localhost:8000/v1/chat/completions \
    --model glm-5.2 \
    --benchmark-version 2.1 \
    --task $TASK_ID \
    --reasoning-effort max

```

## Key Repository Files for Reference

Understanding these files helps debug reproduction issues:

- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)** – Contains the benchmark claim (Terminal-Bench 2.1: 81.0) at line 29 and the `reasoning_effort` documentation at line 81.
- **[`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt)** – Lists exact dependency versions for inference backends.
- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** – Deployment instructions for Ascend NPU hardware.
- **[`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)** – Describes the agentic skill interface used by GLM-5 during terminal tasks.

## Summary

To successfully reproduce GLM-5.2 Terminal-Bench 2.1 benchmark results:

- Deploy the 744-billion-parameter model using **vLLM** or any supported backend with the default **"max"** reasoning effort.
- Ensure the server exposes an **OpenAI-compatible API** at a known endpoint.
- Install the external **Terminal-Bench** evaluation suite and execute with `--reasoning-effort max`.
- Verify the final output shows **81.0**, matching the score documented in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) line 29.

## Frequently Asked Questions

### What is the difference between "max" and "high" reasoning effort?

The **reasoning_effort** parameter controls the compute budget allocated to the model. According to [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) line 81, "max" is the default setting used for official benchmarks, while "high" uses a reduced budget. Running the evaluation with "high" instead of "max" will produce scores that diverge from the published 81.0 result.

### Can I use a different inference backend besides vLLM?

Yes. The `zai-org/GLM-5` repository officially supports **SGLang**, **Transformers**, **K-Transformers**, **Unsloth**, and **Ascend-NPU**. All backends expose a standard OpenAI-compatible HTTP API that Terminal-Bench can target, provided you maintain the "max" reasoning effort configuration.

### Where is the 81.0 benchmark score documented in the repository?

The **81.0** Terminal-Bench 2.1 score is explicitly claimed in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) at line 29, which highlights GLM-5.2's performance compared to GLM-5.1's 62.0. This documentation also references the `resources/` directory containing visual summaries of these results.

### Do I need specific hardware to run GLM-5.2 for this benchmark?

While the raw analysis does not specify minimum hardware requirements, deploying a 744-billion-parameter model typically requires multiple high-memory GPUs or NPUs. The [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) file provides specific guidance for Ascend NPU deployment, while standard CUDA setups should consult the [`requirements.txt`](https://github.com/zai-org/GLM-5/blob/main/requirements.txt) for vLLM compatibility and memory optimization flags like `--dtype bfloat16`.