Terminal-Bench Performance Benchmarks: How Accuracy Is Measured in Qwen-Code

Terminal-Bench evaluates agent performance using three core metrics—accuracy, n_resolved, and n_unresolved—where accuracy is calculated as the ratio of correctly solved tasks to total tasks attempted.

Terminal-Bench is an integration-testing harness in the QwenLM/qwen-code repository that runs agent-driven coding tasks and quantifies performance through structured JSON outputs. Understanding these Terminal-Bench performance benchmarks is essential for evaluating how well the Qwen-Code agent solves software engineering problems compared to a perfect oracle baseline.

Understanding Terminal-Bench Performance Metrics

The benchmark asserts three primary metrics in the test suite located at integration-tests/terminal-bench/terminal-bench.test.ts:

  • accuracy: A floating-point ratio calculated as n_resolved / (n_resolved + n_unresolved), ranging from 0.0 to 1.0. An accuracy of 1.0 indicates perfect task completion.
  • n_resolved: The integer count of tasks the agent completed without error.
  • n_unresolved: The integer count of tasks that failed or timed out.

These metrics are captured in the results.json file generated after each benchmark run. The test expectations at lines 202 and 321-323 validate that the results object contains these properties with valid numerical values.

How Terminal-Bench Calculates Accuracy

After executing a task, Terminal-Bench creates a timestamped output directory containing results.json. The integration test reads this file to validate performance:

const resultsFile = join(latestDir, 'results.json');
const results = JSON.parse(readFileSync(resultsFile, 'utf-8'));

The oracle agent serves as a sanity check baseline. The test asserts that the oracle achieves accuracy === 1.0, n_resolved === 1, and n_unresolved === 0, demonstrating the ideal performance target (lines 202-204).

For the Qwen-Code agent, the benchmark requirements are less stringent but specific:

expect(results).toHaveProperty('accuracy');
expect(results.n_resolved).toBeGreaterThan(0);
expect(results.accuracy).toBeGreaterThan(0);   // line 323

This validation ensures the agent solves at least one task successfully while maintaining non-zero accuracy.

Running Terminal-Bench Benchmarks

Baseline Testing with the Oracle Agent

To establish the performance baseline using the perfect oracle agent on the hello-world task:


# Install terminal-bench (once)

uv tool install --python 3.12 terminal-bench

# Execute the oracle agent

tb run \
  --agent oracle \
  --dataset-path integration-tests/terminal-bench/ci-tasks \
  --task-id hello-world \
  --output-path /tmp/tb-output/oracle-hello \
  --n-concurrent 1

Inspect the perfect baseline results:

cat $(ls -d /tmp/tb-output/oracle-hello/*/ | tail -n1)/results.json

# → {"accuracy":1.0,"n_resolved":1,"n_unresolved":0,...}

Evaluating the Qwen-Code Agent

To benchmark the Qwen-Code agent on complex tasks like swe-bench-astropy-1:


# Run the Qwen-Code agent (requires CLI installation via qwen-code-setup.sh.j2)

tb run \
  --agent-import-path integration-tests.terminal-bench.qwen_code:QwenCodeAgent \
  --agent-kwarg api_key=$OPENAI_API_KEY \
  --agent-kwarg version=latest \
  --dataset-path integration-tests/terminal-bench/ci-tasks \
  --task-id swe-bench-astropy-1 \
  --output-path /tmp/tb-output/qwen-astro \
  --n-concurrent 1

Check the generated metrics:

cat $(ls -d /tmp/tb-output/qwen-astro/*/ | tail -n1)/results.json

# → {"accuracy":0.75,"n_resolved":3,"n_unresolved":1,...}

The agent definition in integration-tests/terminal-bench/qwen_code.py handles API key management and command execution, while the setup script in qwen-code-setup.sh.j2 installs the CLI using npm install -g @qwen-code/qwen-code@${version}.

Key Implementation Files

Summary

  • Terminal-Bench measures agent performance through accuracy, n_resolved, and n_unresolved metrics stored in results.json.
  • Accuracy is computed as the ratio of resolved tasks to total tasks, producing a float between 0 and 1.
  • The oracle agent provides a perfect baseline (accuracy === 1.0) for validating the benchmark harness.
  • The Qwen-Code agent must achieve accuracy > 0 and n_resolved > 0 to pass integration tests.
  • Runtime limits are enforced via DEFAULT_TIMEOUT_MS but do not appear as separate metrics in the output.

Frequently Asked Questions

What is Terminal-Bench?

Terminal-Bench is an integration-testing harness that executes agent-driven coding tasks and records performance outcomes. According to the QwenLM/qwen-code source code, it runs tasks like hello-world and swe-bench-astropy-1 while capturing success rates in a structured JSON format.

How is accuracy calculated in Terminal-Bench?

Accuracy is calculated as n_resolved / (n_resolved + n_unresolved), representing the proportion of tasks solved correctly. This floating-point value ranges from 0.0 (all tasks failed) to 1.0 (all tasks succeeded), as validated in terminal-bench.test.ts at lines 321-323.

What is the oracle agent in Terminal-Bench?

The oracle agent is a perfect reference implementation used to validate the benchmark infrastructure. It achieves accuracy === 1.0, n_resolved === 1, and n_unresolved === 0 on test tasks, providing a baseline against which the Qwen-Code agent's performance is compared.

Where are Terminal-Bench results stored?

Results are stored in a timestamped subdirectory within the specified output path, specifically in a results.json file. The integration test reads this file using join(latestDir, 'results.json') to validate the accuracy and resolution metrics after each run.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →