# Terminal-Bench Performance Benchmarks: How Accuracy Is Measured in Qwen-Code

> Discover Terminal-Bench performance benchmarks for Qwen-Code. Learn how accuracy is precisely measured as the ratio of correctly solved to total attempted tasks.

- Repository: [Qwen/qwen-code](https://github.com/qwenlm/qwen-code)
- Tags: performance
- Published: 2026-02-19

---

**Terminal-Bench evaluates agent performance using three core metrics—accuracy, n_resolved, and n_unresolved—where accuracy is calculated as the ratio of correctly solved tasks to total tasks attempted.**

Terminal-Bench is an integration-testing harness in the `QwenLM/qwen-code` repository that runs agent-driven coding tasks and quantifies performance through structured JSON outputs. Understanding these **Terminal-Bench performance benchmarks** is essential for evaluating how well the Qwen-Code agent solves software engineering problems compared to a perfect oracle baseline.

## Understanding Terminal-Bench Performance Metrics

The benchmark asserts three primary metrics in the test suite located at [`integration-tests/terminal-bench/terminal-bench.test.ts`](https://github.com/QwenLM/qwen-code/blob/main/integration-tests/terminal-bench/terminal-bench.test.ts):

- **`accuracy`**: A floating-point ratio calculated as `n_resolved / (n_resolved + n_unresolved)`, ranging from `0.0` to `1.0`. An accuracy of `1.0` indicates perfect task completion.
- **`n_resolved`**: The integer count of tasks the agent completed without error.
- **`n_unresolved`**: The integer count of tasks that failed or timed out.

These metrics are captured in the [`results.json`](https://github.com/QwenLM/qwen-code/blob/main/results.json) file generated after each benchmark run. The test expectations at lines 202 and 321-323 validate that the results object contains these properties with valid numerical values.

## How Terminal-Bench Calculates Accuracy

After executing a task, Terminal-Bench creates a timestamped output directory containing [`results.json`](https://github.com/QwenLM/qwen-code/blob/main/results.json). The integration test reads this file to validate performance:

```ts
const resultsFile = join(latestDir, 'results.json');
const results = JSON.parse(readFileSync(resultsFile, 'utf-8'));

```

The **oracle agent** serves as a sanity check baseline. The test asserts that the oracle achieves `accuracy === 1.0`, `n_resolved === 1`, and `n_unresolved === 0`, demonstrating the ideal performance target (lines 202-204).

For the **Qwen-Code agent**, the benchmark requirements are less stringent but specific:

```ts
expect(results).toHaveProperty('accuracy');
expect(results.n_resolved).toBeGreaterThan(0);
expect(results.accuracy).toBeGreaterThan(0);   // line 323

```

This validation ensures the agent solves at least one task successfully while maintaining non-zero accuracy.

## Running Terminal-Bench Benchmarks

### Baseline Testing with the Oracle Agent

To establish the performance baseline using the perfect oracle agent on the `hello-world` task:

```bash

# Install terminal-bench (once)

uv tool install --python 3.12 terminal-bench

# Execute the oracle agent

tb run \
  --agent oracle \
  --dataset-path integration-tests/terminal-bench/ci-tasks \
  --task-id hello-world \
  --output-path /tmp/tb-output/oracle-hello \
  --n-concurrent 1

```

Inspect the perfect baseline results:

```bash
cat $(ls -d /tmp/tb-output/oracle-hello/*/ | tail -n1)/results.json

# → {"accuracy":1.0,"n_resolved":1,"n_unresolved":0,...}

```

### Evaluating the Qwen-Code Agent

To benchmark the Qwen-Code agent on complex tasks like `swe-bench-astropy-1`:

```bash

# Run the Qwen-Code agent (requires CLI installation via qwen-code-setup.sh.j2)

tb run \
  --agent-import-path integration-tests.terminal-bench.qwen_code:QwenCodeAgent \
  --agent-kwarg api_key=$OPENAI_API_KEY \
  --agent-kwarg version=latest \
  --dataset-path integration-tests/terminal-bench/ci-tasks \
  --task-id swe-bench-astropy-1 \
  --output-path /tmp/tb-output/qwen-astro \
  --n-concurrent 1

```

Check the generated metrics:

```bash
cat $(ls -d /tmp/tb-output/qwen-astro/*/ | tail -n1)/results.json

# → {"accuracy":0.75,"n_resolved":3,"n_unresolved":1,...}

```

The agent definition in [`integration-tests/terminal-bench/qwen_code.py`](https://github.com/QwenLM/qwen-code/blob/main/integration-tests/terminal-bench/qwen_code.py) handles API key management and command execution, while the setup script in `qwen-code-setup.sh.j2` installs the CLI using `npm install -g @qwen-code/qwen-code@${version}`.

## Key Implementation Files

- **[`integration-tests/terminal-bench/terminal-bench.test.ts`](https://github.com/QwenLM/qwen-code/blob/main/integration-tests/terminal-bench/terminal-bench.test.ts)**: Drives the benchmark suite, executes both oracle and Qwen-Code agents, and validates the three core metrics at lines 202, 321-323.
- **[`integration-tests/terminal-bench/qwen_code.py`](https://github.com/QwenLM/qwen-code/blob/main/integration-tests/terminal-bench/qwen_code.py)**: Python wrapper exposing the Qwen-Code agent to Terminal-Bench with configurable model parameters.
- **`integration-tests/terminal-bench/qwen-code-setup.sh.j2`**: Jinja2 template for installing the Qwen-Code CLI during test container initialization.
- **[`integration-tests/terminal-bench/ci-tasks/hello-world/task.yaml`](https://github.com/QwenLM/qwen-code/blob/main/integration-tests/terminal-bench/ci-tasks/hello-world/task.yaml)**: Simple CI task definition used for baseline validation.
- **[`integration-tests/terminal-bench/ci-tasks/swe-bench-astropy-1/task.yaml`](https://github.com/QwenLM/qwen-code/blob/main/integration-tests/terminal-bench/ci-tasks/swe-bench-astropy-1/task.yaml)**: Complex benchmark task for evaluating real-world agent performance.

## Summary

- Terminal-Bench measures agent performance through **accuracy**, **n_resolved**, and **n_unresolved** metrics stored in [`results.json`](https://github.com/QwenLM/qwen-code/blob/main/results.json).
- **Accuracy** is computed as the ratio of resolved tasks to total tasks, producing a float between 0 and 1.
- The **oracle agent** provides a perfect baseline (`accuracy === 1.0`) for validating the benchmark harness.
- The **Qwen-Code agent** must achieve `accuracy > 0` and `n_resolved > 0` to pass integration tests.
- Runtime limits are enforced via `DEFAULT_TIMEOUT_MS` but do not appear as separate metrics in the output.

## Frequently Asked Questions

### What is Terminal-Bench?

Terminal-Bench is an integration-testing harness that executes agent-driven coding tasks and records performance outcomes. According to the `QwenLM/qwen-code` source code, it runs tasks like `hello-world` and `swe-bench-astropy-1` while capturing success rates in a structured JSON format.

### How is accuracy calculated in Terminal-Bench?

Accuracy is calculated as `n_resolved / (n_resolved + n_unresolved)`, representing the proportion of tasks solved correctly. This floating-point value ranges from 0.0 (all tasks failed) to 1.0 (all tasks succeeded), as validated in [`terminal-bench.test.ts`](https://github.com/QwenLM/qwen-code/blob/main/terminal-bench.test.ts) at lines 321-323.

### What is the oracle agent in Terminal-Bench?

The oracle agent is a perfect reference implementation used to validate the benchmark infrastructure. It achieves `accuracy === 1.0`, `n_resolved === 1`, and `n_unresolved === 0` on test tasks, providing a baseline against which the Qwen-Code agent's performance is compared.

### Where are Terminal-Bench results stored?

Results are stored in a timestamped subdirectory within the specified output path, specifically in a [`results.json`](https://github.com/QwenLM/qwen-code/blob/main/results.json) file. The integration test reads this file using `join(latestDir, 'results.json')` to validate the accuracy and resolution metrics after each run.