Performance Benchmarks for Qwen-Agent: DeepPlanning and Code Interpreter Evaluation
Qwen-Agent provides two official benchmark suites—DeepPlanning and Code Interpreter—that evaluate long-horizon agentic planning and code-generation capabilities with quantitative metrics for travel, shopping, math, and visualization tasks.
The Qwen-Agent repository maintains rigorous performance benchmarks for Qwen-Agent to evaluate its agentic capabilities across complex, multi-step tasks. These standardized evaluations measure the framework's ability to handle constraint satisfaction, commonsense reasoning, and executable code generation. Developers can reproduce these assessments using the orchestration scripts and evaluation modules located in the benchmark/ directory.
DeepPlanning Benchmark: Travel and Shopping Evaluation
Task Domains and Success Metrics
The DeepPlanning benchmark evaluates agentic planning through two distinct domains: Travel planning and Shopping tasks. This suite tests constraint satisfaction, commonsense reasoning, and personalization capabilities across scenarios requiring multiple tool invocations and long-horizon decision making.
For the Travel domain, the framework calculates a composite_score combining case accuracy, commonsense reasoning scores, and personalized recommendation alignment. The Shopping domain measures match_rate between planned and ground-truth item selections alongside weighted_average_case_score metrics.
Sample Results for Qwen-Plus
According to the aggregated results documented in benchmark/deepplanning/README.md, the qwen-plus model demonstrates the following performance characteristics:
- Travel domain: composite_score of 0.2813, case_acc of 0.0, commonsense_score of 0.4292, and personalized_score of 0.1333
- Shopping domain: match_rate of 0.6209 and weighted_average_case_score of 0.1417
- Overall success rate across both domains: approximately 56.7%
These metrics indicate particular strength in commonsense reasoning (42.92%) while revealing challenges in strict case-level accuracy constraints for travel planning.
Running the DeepPlanning Suite
Execute the unified orchestrator to evaluate models on both domains:
# Install dependencies
pip install -r benchmark/deepplanning/requirements.txt
# Configure models in benchmark/deepplanning/models_config.json
# Set API keys in .env file
# Launch unified runner
bash benchmark/deepplanning/run_all.sh
The run_all.sh script parses models_config.json, loads API credentials from .env, and iterates through configured models while invoking domain-specific run.sh scripts for Travel and Shopping tasks. Intermediate results populate travelplanning/result_report/ and shoppingplanning/result_report/ directories.
Code Interpreter Benchmark: Python Generation and Execution
Evaluation Tasks and Accuracy Metrics
The Code Interpreter benchmark measures the ability to generate executable Python code and produce correct computational results across Math and Visualization tasks. The evaluation framework tracks both code executability and numerical accuracy using isolated Docker sandboxes for secure execution.
Key performance indicators include:
- Math problem-solving accuracy
- Visualization task completion rates (Hard and Easy variants)
- General executable rate across all code generation attempts
Comparative Model Performance
As documented in benchmark/code_interpreter/README.md, the evaluation provides comparative metrics across model architectures:
- GPT-4: Math 82.8%, Vis-Hard 66.7%, Vis-Easy 60.8%, General 82.8%
- Qwen-72B-Chat: Math 72.7%, Vis-Hard 41.7%, Vis-Easy 43.0%, General 82.8%
- Qwen-7B-Chat: Math 41.9%, Vis-Hard 23.8%, Vis-Easy 38.0%, General 67.2%
Notably, Qwen-72B-Chat achieves general executability of 82.8%, matching GPT-4 performance, while delivering strong mathematical reasoning at 72.7% accuracy. However, performance gaps widen on complex visualization tasks requiring sophisticated graphical generation.
Executing the Code Interpreter Benchmark
Run the evaluation using the inference driver:
# Install dependencies
pip install -r benchmark/code_interpreter/requirements.txt
# Download and prepare dataset
cd benchmark/code_interpreter
wget https://qianwen-res.oss-cn-beijing.aliyuncs.com/assets/qwen_agent/benchmark_code_interpreter_data.zip
unzip benchmark_code_interpreter_data.zip
mkdir eval_data && mv eval_code_interpreter_v1.jsonl eval_data/
# Run evaluation
python inference_and_execute.py --model qwen-7b-chat
The inference_and_execute.py script drives the LLM through the Agent API, executes generated code in Docker containers, and computes executability and correctness metrics. Result JSON files are saved to benchmark/code_interpreter/output_data/.
Benchmark Architecture and Result Aggregation
Unified Orchestrator Pattern
Both benchmark suites implement an orchestrator → domain-specific runner → evaluation pipeline. The benchmark/deepplanning/run_all.sh script serves as the central coordinator, while inference_and_execute.py handles Code Interpreter orchestration.
For each evaluation domain, the architecture:
- Loads specialized toolsets (search APIs for Travel, Docker containers for Code Interpreter)
- Drives LLM interactions through the Agent API (
qwen_agent.agents.AssistantorFnCallAgent) - Converts raw outputs to structured formats (JSON for travel plans)
- Validates results using Python evaluation modules in
benchmark/deepplanning/*/evaluation/*.py
Aggregating Cross-Domain Results
The benchmark/deepplanning/aggregate_results.py utility merges per-domain JSON reports into comprehensive cross-domain summaries. After benchmark completion, aggregated results appear in aggregated_results/<model>_aggregated.json:
import json
from pathlib import Path
# Load and analyze aggregated results
result_path = Path("aggregated_results/qwen-plus_aggregated.json")
with result_path.open() as f:
report = json.load(f)
print(f"Travel composite score: {report['domains']['travel']['composite_score']}")
print(f"Shopping match rate: {report['domains']['shopping']['match_rate']}")
Summary
- Qwen-Agent maintains two complementary performance benchmarks: DeepPlanning for agentic planning and Code Interpreter for code generation capabilities.
- DeepPlanning evaluates Travel and Shopping domains with metrics including composite scores, match rates, and overall success rates (approximately 56.7% for qwen-plus).
- Code Interpreter benchmarking compares models on Math and Visualization accuracy, with Qwen-72B-Chat achieving 72.7% math accuracy and 82.8% general executability.
- The benchmark architecture uses standardized orchestrators (
run_all.sh,inference_and_execute.py) and aggregation utilities (aggregate_results.py) for reproducible evaluation workflows. - All benchmark configurations reside in the
benchmark/directory, utilizingmodels_config.jsonfor model endpoint management and.envfiles for credential handling.
Frequently Asked Questions
What specific metrics does the DeepPlanning benchmark use to evaluate Qwen-Agent?
The DeepPlanning benchmark employs domain-specific metrics including composite_score, case_acc, commonsense_score, and personalized_score for Travel tasks, alongside match_rate and weighted_average_case_score for Shopping tasks. These metrics evaluate constraint satisfaction accuracy, reasoning capabilities, and personalization alignment across multi-step planning scenarios requiring tool use and context management.
How does Qwen-72B-Chat performance compare to GPT-4 on the Code Interpreter benchmark?
According to the results in benchmark/code_interpreter/README.md, Qwen-72B-Chat achieves 72.7% accuracy on Math tasks compared to GPT-4's 82.8%, while matching GPT-4's 82.8% general executable rate. However, Qwen-72B-Chat scores substantially lower on Visualization tasks (41.7% Hard, 43.0% Easy) compared to GPT-4 (66.7% Hard, 60.8% Easy), indicating relative challenges in complex graphical generation.
Where are benchmark results stored after running the evaluation scripts?
DeepPlanning results are stored in travelplanning/result_report/ and shoppingplanning/result_report/ directories, with consolidated summaries in aggregated_results/<model>_aggregated.json generated by aggregate_results.py. Code Interpreter outputs are saved to benchmark/code_interpreter/output_data/ as JSON files containing execution traces, correctness flags, and performance metrics for each test case.
Can developers evaluate custom LLM backends using the Qwen-Agent benchmark suite?
Yes. Configure custom model endpoints in benchmark/deepplanning/models_config.json by specifying API URLs, authentication keys, and model identifiers. The run_all.sh orchestrator and inference_and_execute.py scripts route requests to these configured endpoints while maintaining the standardized evaluation protocol, enabling direct comparison of proprietary or fine-tuned models against the published baselines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →