Terminal-Bench Performance Benchmarks: How Accuracy Is Measured in Qwen-Code
Terminal-Bench evaluates agent performance using three core metrics—accuracy, n_resolved, and n_unresolved—where accuracy is calculated as the ratio of correctly solved tasks to total tasks attempted.
Terminal-Bench is an integration-testing harness in the QwenLM/qwen-code repository that runs agent-driven coding tasks and quantifies performance through structured JSON outputs. Understanding these Terminal-Bench performance benchmarks is essential for evaluating how well the Qwen-Code agent solves software engineering problems compared to a perfect oracle baseline.
Understanding Terminal-Bench Performance Metrics
The benchmark asserts three primary metrics in the test suite located at integration-tests/terminal-bench/terminal-bench.test.ts:
accuracy: A floating-point ratio calculated asn_resolved / (n_resolved + n_unresolved), ranging from0.0to1.0. An accuracy of1.0indicates perfect task completion.n_resolved: The integer count of tasks the agent completed without error.n_unresolved: The integer count of tasks that failed or timed out.
These metrics are captured in the results.json file generated after each benchmark run. The test expectations at lines 202 and 321-323 validate that the results object contains these properties with valid numerical values.
How Terminal-Bench Calculates Accuracy
After executing a task, Terminal-Bench creates a timestamped output directory containing results.json. The integration test reads this file to validate performance:
const resultsFile = join(latestDir, 'results.json');
const results = JSON.parse(readFileSync(resultsFile, 'utf-8'));
The oracle agent serves as a sanity check baseline. The test asserts that the oracle achieves accuracy === 1.0, n_resolved === 1, and n_unresolved === 0, demonstrating the ideal performance target (lines 202-204).
For the Qwen-Code agent, the benchmark requirements are less stringent but specific:
expect(results).toHaveProperty('accuracy');
expect(results.n_resolved).toBeGreaterThan(0);
expect(results.accuracy).toBeGreaterThan(0); // line 323
This validation ensures the agent solves at least one task successfully while maintaining non-zero accuracy.
Running Terminal-Bench Benchmarks
Baseline Testing with the Oracle Agent
To establish the performance baseline using the perfect oracle agent on the hello-world task:
# Install terminal-bench (once)
uv tool install --python 3.12 terminal-bench
# Execute the oracle agent
tb run \
--agent oracle \
--dataset-path integration-tests/terminal-bench/ci-tasks \
--task-id hello-world \
--output-path /tmp/tb-output/oracle-hello \
--n-concurrent 1
Inspect the perfect baseline results:
cat $(ls -d /tmp/tb-output/oracle-hello/*/ | tail -n1)/results.json
# → {"accuracy":1.0,"n_resolved":1,"n_unresolved":0,...}
Evaluating the Qwen-Code Agent
To benchmark the Qwen-Code agent on complex tasks like swe-bench-astropy-1:
# Run the Qwen-Code agent (requires CLI installation via qwen-code-setup.sh.j2)
tb run \
--agent-import-path integration-tests.terminal-bench.qwen_code:QwenCodeAgent \
--agent-kwarg api_key=$OPENAI_API_KEY \
--agent-kwarg version=latest \
--dataset-path integration-tests/terminal-bench/ci-tasks \
--task-id swe-bench-astropy-1 \
--output-path /tmp/tb-output/qwen-astro \
--n-concurrent 1
Check the generated metrics:
cat $(ls -d /tmp/tb-output/qwen-astro/*/ | tail -n1)/results.json
# → {"accuracy":0.75,"n_resolved":3,"n_unresolved":1,...}
The agent definition in integration-tests/terminal-bench/qwen_code.py handles API key management and command execution, while the setup script in qwen-code-setup.sh.j2 installs the CLI using npm install -g @qwen-code/qwen-code@${version}.
Key Implementation Files
integration-tests/terminal-bench/terminal-bench.test.ts: Drives the benchmark suite, executes both oracle and Qwen-Code agents, and validates the three core metrics at lines 202, 321-323.integration-tests/terminal-bench/qwen_code.py: Python wrapper exposing the Qwen-Code agent to Terminal-Bench with configurable model parameters.integration-tests/terminal-bench/qwen-code-setup.sh.j2: Jinja2 template for installing the Qwen-Code CLI during test container initialization.integration-tests/terminal-bench/ci-tasks/hello-world/task.yaml: Simple CI task definition used for baseline validation.integration-tests/terminal-bench/ci-tasks/swe-bench-astropy-1/task.yaml: Complex benchmark task for evaluating real-world agent performance.
Summary
- Terminal-Bench measures agent performance through accuracy, n_resolved, and n_unresolved metrics stored in
results.json. - Accuracy is computed as the ratio of resolved tasks to total tasks, producing a float between 0 and 1.
- The oracle agent provides a perfect baseline (
accuracy === 1.0) for validating the benchmark harness. - The Qwen-Code agent must achieve
accuracy > 0andn_resolved > 0to pass integration tests. - Runtime limits are enforced via
DEFAULT_TIMEOUT_MSbut do not appear as separate metrics in the output.
Frequently Asked Questions
What is Terminal-Bench?
Terminal-Bench is an integration-testing harness that executes agent-driven coding tasks and records performance outcomes. According to the QwenLM/qwen-code source code, it runs tasks like hello-world and swe-bench-astropy-1 while capturing success rates in a structured JSON format.
How is accuracy calculated in Terminal-Bench?
Accuracy is calculated as n_resolved / (n_resolved + n_unresolved), representing the proportion of tasks solved correctly. This floating-point value ranges from 0.0 (all tasks failed) to 1.0 (all tasks succeeded), as validated in terminal-bench.test.ts at lines 321-323.
What is the oracle agent in Terminal-Bench?
The oracle agent is a perfect reference implementation used to validate the benchmark infrastructure. It achieves accuracy === 1.0, n_resolved === 1, and n_unresolved === 0 on test tasks, providing a baseline against which the Qwen-Code agent's performance is compared.
Where are Terminal-Bench results stored?
Results are stored in a timestamped subdirectory within the specified output path, specifically in a results.json file. The integration test reads this file using join(latestDir, 'results.json') to validate the accuracy and resolution metrics after each run.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →