# Performance Benchmarks for Qwen-Agent: DeepPlanning and Code Interpreter Evaluation

> Explore Qwen-Agent's performance benchmarks: DeepPlanning and Code Interpreter evaluate agentic planning and code generation for diverse tasks. Discover quantitative metrics for travel, shopping, math, and visualization.

- Repository: [Qwen/Qwen-Agent](https://github.com/qwenlm/Qwen-Agent)
- Tags: deep-dive
- Published: 2026-03-09

---

**Qwen-Agent provides two official benchmark suites—DeepPlanning and Code Interpreter—that evaluate long-horizon agentic planning and code-generation capabilities with quantitative metrics for travel, shopping, math, and visualization tasks.**

The Qwen-Agent repository maintains rigorous **performance benchmarks for Qwen-Agent** to evaluate its agentic capabilities across complex, multi-step tasks. These standardized evaluations measure the framework's ability to handle constraint satisfaction, commonsense reasoning, and executable code generation. Developers can reproduce these assessments using the orchestration scripts and evaluation modules located in the `benchmark/` directory.

## DeepPlanning Benchmark: Travel and Shopping Evaluation

### Task Domains and Success Metrics

The **DeepPlanning** benchmark evaluates agentic planning through two distinct domains: **Travel** planning and **Shopping** tasks. This suite tests constraint satisfaction, commonsense reasoning, and personalization capabilities across scenarios requiring multiple tool invocations and long-horizon decision making.

For the Travel domain, the framework calculates a **composite_score** combining case accuracy, commonsense reasoning scores, and personalized recommendation alignment. The Shopping domain measures **match_rate** between planned and ground-truth item selections alongside **weighted_average_case_score** metrics.

### Sample Results for Qwen-Plus

According to the aggregated results documented in [`benchmark/deepplanning/README.md`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/README.md), the `qwen-plus` model demonstrates the following performance characteristics:

- **Travel domain**: composite_score of **0.2813**, case_acc of **0.0**, commonsense_score of **0.4292**, and personalized_score of **0.1333**
- **Shopping domain**: match_rate of **0.6209** and weighted_average_case_score of **0.1417**
- **Overall success rate** across both domains: approximately **56.7%**

These metrics indicate particular strength in commonsense reasoning (42.92%) while revealing challenges in strict case-level accuracy constraints for travel planning.

### Running the DeepPlanning Suite

Execute the unified orchestrator to evaluate models on both domains:

```bash

# Install dependencies

pip install -r benchmark/deepplanning/requirements.txt

# Configure models in benchmark/deepplanning/models_config.json

# Set API keys in .env file

# Launch unified runner

bash benchmark/deepplanning/run_all.sh

```

The [`run_all.sh`](https://github.com/QwenLM/Qwen-Agent/blob/main/run_all.sh) script parses [`models_config.json`](https://github.com/QwenLM/Qwen-Agent/blob/main/models_config.json), loads API credentials from `.env`, and iterates through configured models while invoking domain-specific [`run.sh`](https://github.com/QwenLM/Qwen-Agent/blob/main/run.sh) scripts for Travel and Shopping tasks. Intermediate results populate `travelplanning/result_report/` and `shoppingplanning/result_report/` directories.

## Code Interpreter Benchmark: Python Generation and Execution

### Evaluation Tasks and Accuracy Metrics

The **Code Interpreter** benchmark measures the ability to generate **executable Python code** and produce correct computational results across **Math** and **Visualization** tasks. The evaluation framework tracks both code executability and numerical accuracy using isolated Docker sandboxes for secure execution.

Key performance indicators include:

- Math problem-solving accuracy
- Visualization task completion rates (Hard and Easy variants)
- General executable rate across all code generation attempts

### Comparative Model Performance

As documented in [`benchmark/code_interpreter/README.md`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/code_interpreter/README.md), the evaluation provides comparative metrics across model architectures:

- **GPT-4**: Math **82.8%**, Vis-Hard **66.7%**, Vis-Easy **60.8%**, General **82.8%**
- **Qwen-72B-Chat**: Math **72.7%**, Vis-Hard **41.7%**, Vis-Easy **43.0%**, General **82.8%**
- **Qwen-7B-Chat**: Math **41.9%**, Vis-Hard **23.8%**, Vis-Easy **38.0%**, General **67.2%**

Notably, **Qwen-72B-Chat** achieves general executability of **82.8%**, matching GPT-4 performance, while delivering strong mathematical reasoning at **72.7%** accuracy. However, performance gaps widen on complex visualization tasks requiring sophisticated graphical generation.

### Executing the Code Interpreter Benchmark

Run the evaluation using the inference driver:

```bash

# Install dependencies

pip install -r benchmark/code_interpreter/requirements.txt

# Download and prepare dataset

cd benchmark/code_interpreter
wget https://qianwen-res.oss-cn-beijing.aliyuncs.com/assets/qwen_agent/benchmark_code_interpreter_data.zip
unzip benchmark_code_interpreter_data.zip
mkdir eval_data && mv eval_code_interpreter_v1.jsonl eval_data/

# Run evaluation

python inference_and_execute.py --model qwen-7b-chat

```

The [`inference_and_execute.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/inference_and_execute.py) script drives the LLM through the Agent API, executes generated code in Docker containers, and computes executability and correctness metrics. Result JSON files are saved to `benchmark/code_interpreter/output_data/`.

## Benchmark Architecture and Result Aggregation

### Unified Orchestrator Pattern

Both benchmark suites implement an **orchestrator → domain-specific runner → evaluation** pipeline. The [`benchmark/deepplanning/run_all.sh`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/run_all.sh) script serves as the central coordinator, while [`inference_and_execute.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/inference_and_execute.py) handles Code Interpreter orchestration.

For each evaluation domain, the architecture:

1. Loads specialized toolsets (search APIs for Travel, Docker containers for Code Interpreter)
2. Drives LLM interactions through the **Agent API** (`qwen_agent.agents.Assistant` or `FnCallAgent`)
3. Converts raw outputs to structured formats (JSON for travel plans)
4. Validates results using Python evaluation modules in `benchmark/deepplanning/*/evaluation/*.py`

### Aggregating Cross-Domain Results

The [`benchmark/deepplanning/aggregate_results.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/aggregate_results.py) utility merges per-domain JSON reports into comprehensive cross-domain summaries. After benchmark completion, aggregated results appear in `aggregated_results/<model>_aggregated.json`:

```python
import json
from pathlib import Path

# Load and analyze aggregated results

result_path = Path("aggregated_results/qwen-plus_aggregated.json")
with result_path.open() as f:
    report = json.load(f)

print(f"Travel composite score: {report['domains']['travel']['composite_score']}")
print(f"Shopping match rate: {report['domains']['shopping']['match_rate']}")

```

## Summary

- Qwen-Agent maintains two complementary **performance benchmarks**: DeepPlanning for agentic planning and Code Interpreter for code generation capabilities.
- DeepPlanning evaluates Travel and Shopping domains with metrics including composite scores, match rates, and overall success rates (approximately 56.7% for qwen-plus).
- Code Interpreter benchmarking compares models on Math and Visualization accuracy, with Qwen-72B-Chat achieving 72.7% math accuracy and 82.8% general executability.
- The benchmark architecture uses standardized orchestrators ([`run_all.sh`](https://github.com/QwenLM/Qwen-Agent/blob/main/run_all.sh), [`inference_and_execute.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/inference_and_execute.py)) and aggregation utilities ([`aggregate_results.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/aggregate_results.py)) for reproducible evaluation workflows.
- All benchmark configurations reside in the `benchmark/` directory, utilizing [`models_config.json`](https://github.com/QwenLM/Qwen-Agent/blob/main/models_config.json) for model endpoint management and `.env` files for credential handling.

## Frequently Asked Questions

### What specific metrics does the DeepPlanning benchmark use to evaluate Qwen-Agent?

The DeepPlanning benchmark employs domain-specific metrics including **composite_score**, **case_acc**, **commonsense_score**, and **personalized_score** for Travel tasks, alongside **match_rate** and **weighted_average_case_score** for Shopping tasks. These metrics evaluate constraint satisfaction accuracy, reasoning capabilities, and personalization alignment across multi-step planning scenarios requiring tool use and context management.

### How does Qwen-72B-Chat performance compare to GPT-4 on the Code Interpreter benchmark?

According to the results in [`benchmark/code_interpreter/README.md`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/code_interpreter/README.md), Qwen-72B-Chat achieves **72.7%** accuracy on Math tasks compared to GPT-4's **82.8%**, while matching GPT-4's **82.8%** general executable rate. However, Qwen-72B-Chat scores substantially lower on Visualization tasks (41.7% Hard, 43.0% Easy) compared to GPT-4 (66.7% Hard, 60.8% Easy), indicating relative challenges in complex graphical generation.

### Where are benchmark results stored after running the evaluation scripts?

DeepPlanning results are stored in `travelplanning/result_report/` and `shoppingplanning/result_report/` directories, with consolidated summaries in `aggregated_results/<model>_aggregated.json` generated by [`aggregate_results.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/aggregate_results.py). Code Interpreter outputs are saved to `benchmark/code_interpreter/output_data/` as JSON files containing execution traces, correctness flags, and performance metrics for each test case.

### Can developers evaluate custom LLM backends using the Qwen-Agent benchmark suite?

Yes. Configure custom model endpoints in [`benchmark/deepplanning/models_config.json`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/models_config.json) by specifying API URLs, authentication keys, and model identifiers. The [`run_all.sh`](https://github.com/QwenLM/Qwen-Agent/blob/main/run_all.sh) orchestrator and [`inference_and_execute.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/inference_and_execute.py) scripts route requests to these configured endpoints while maintaining the standardized evaluation protocol, enabling direct comparison of proprietary or fine-tuned models against the published baselines.