How to Evaluate the Performance of Qwen-Agent: Complete Benchmarking Guide
To evaluate Qwen-Agent's performance, execute the three-stage benchmark pipeline—agent inference, plan conversion, and constraint evaluation—using benchmark/deepplanning/travelplanning/run.py, which calculates a composite score based on eight weighted commonsense dimensions and mandatory hard constraints.
Qwen-Agent is an open-source large language model agent framework developed by QwenLM. Evaluating its performance requires running the specialized travel planning benchmark that measures the agent's ability to generate feasible, cost-aware itineraries. This guide walks through the complete evaluation workflow, from running the benchmark pipeline to interpreting the composite scoring metrics.
Running the Qwen-Agent Benchmark Pipeline
The benchmark evaluates Qwen-Agent through a three-stage pipeline orchestrated by benchmark/deepplanning/travelplanning/run.py. This script handles agent inference, plan conversion, and final evaluation in sequence.
Execute the full pipeline with:
python benchmark/deepplanning/travelplanning/run.py \
--model qwen-plus \
--language en \
--workers 20 \
--output-dir ./results/qwenplus_en
Key parameters include:
--model– The model configuration name (must exist inmodels_config.json).--language– Test set language (enorzh), which selects the appropriate database and queries.--workers– Parallel threads for inference and evaluation.--output-dir– Directory for intermediate and final results.
The script sequentially calls run_agent_inference for generation, convert_reports for JSON structuring, and evaluate_plans for scoring, as implemented in the orchestration logic of run.py.
Core Evaluation Logic with evaluate_plans
The primary evaluation entry point is evaluate_plans in benchmark/deepplanning/travelplanning/evaluation/eval_converted.py. This function validates converted JSON plans against ground-truth test data and domain-specific databases.
from benchmark.deepplanning.travelplanning.evaluation.eval_converted import evaluate_plans
results = evaluate_plans(
plan_dir=Path("results/qwenplus_en/converted_plans"),
test_data_path=Path("benchmark/deepplanning/travelplanning/data/travelplanning_query_en.json"),
output_dir=Path("results/qwenplus_en/evaluation"),
database_dir=Path("benchmark/deepplanning/travelplanning/database/database_en"),
)
Inputs required:
- plan_dir – Directory containing converted JSON plans (one file per test sample).
- test_data_path – JSON file with metadata (
days,org, destination) for validation. - database_dir – CSV resources (hotels, restaurants, attractions) used by constraint checkers.
The function processes samples in parallel using ThreadPoolExecutor and returns a list of per-sample dictionaries containing scoring metrics.
Scoring Methodology: Commonsense and Hard Constraints
Qwen-Agent evaluation uses two distinct scoring systems: weighted commonsense dimensions and one-vote-veto hard constraints.
Commonsense Scoring
The calculate_weighted_score function aggregates results across eight dimensions defined in constraints_commonsense.py:
- Route Consistency
- Sandbox Compliance
- Itinerary Structure
- Time Feasibility
- Business Hours
- Duration Rationality
- Cost Calculation Accuracy
- Activity Diversity
Each dimension carries a weight of 0.125 and follows a one-vote-veto rule: the dimension scores 1.0 only if all its internal checks pass, otherwise 0.0. The final total_weighted_score is the weighted sum of these dimension scores.
Hard Constraint Scoring
The calculate_hard_score function applies a strict one-vote-veto across all mandatory rules defined in constraints_hard.py. If any hard constraint fails (e.g., missing intercity transport), the score becomes 0.0; otherwise 1.0.
Interpreting Evaluation Results
Each evaluated plan returns a dictionary with the following key metrics:
{
"sample_id": "42",
"total_weighted_score": 0.875,
"dimension_scores": {
"Route Consistency": 1,
"Sandbox Compliance": 1,
"Itinerary Structure": 1,
"Time Feasibility": 1,
"Business Hours": 0,
"Duration Rationality": 1,
"Cost Calculation Accuracy": 1,
"Activity Diversity": 1
},
"hard_score": 1,
"composite_score": 0.9375,
"case_acc": 0,
"details": {}
}
Metric definitions:
total_weighted_score– Weighted commonsense score (0 to 1).hard_score– Binary hard constraint satisfaction (0 or 1).composite_score– Average of commonsense and hard scores.case_acc– Strict all-pass indicator (1 only if both scores equal 1.0).
Use benchmark/deepplanning/travelplanning/evaluation/aggregate_results.py or custom pandas analysis to compute aggregate statistics across the full test set.
Practical Implementation: End-to-End Evaluation Script
When you have existing converted plans, bypass the inference stage and run evaluation directly:
#!/usr/bin/env python
# evaluate_qwen_agent.py
# -------------------------------------------------
# Run the full Qwen-Agent benchmark evaluation.
# -------------------------------------------------
import json
from pathlib import Path
from benchmark.deepplanning.travelplanning.evaluation.eval_converted import evaluate_plans
def main():
# Configuration
model = "qwen-plus"
language = "en"
base_dir = Path(__file__).parent
results_dir = base_dir / "results" / f"{model}_{language}"
# Paths
plan_dir = results_dir / "converted_plans"
test_data = base_dir / "benchmark" / "deepplanning" / "travelplanning" / "data" / f"travelplanning_query_{language}.json"
database_dir = base_dir / "benchmark" / "deepplanning" / "travelplanning" / "database" / f"database_{language}"
out_dir = results_dir / "evaluation"
# Execute evaluation
all_results = evaluate_plans(
plan_dir=plan_dir,
test_data_path=test_data,
output_dir=out_dir,
database_dir=database_dir,
)
# Compute summary statistics
composite_scores = [r["composite_score"] for r in all_results]
case_accs = [r["case_acc"] for r in all_results]
print("\n=== Qwen-Agent Benchmark Summary ===")
print(f"Samples evaluated: {len(all_results)}")
print(f"Avg. composite: {sum(composite_scores)/len(composite_scores):.3f}")
print(f"Overall case acc: {sum(case_accs)/len(case_accs):.3%}")
# Persist results
out_dir.mkdir(parents=True, exist_ok=True)
(out_dir / "full_results.json").write_text(json.dumps(all_results, indent=2))
if __name__ == "__main__":
main()
Run this script after generating the converted_plans directory to obtain quantitative performance metrics.
Customizing the Evaluation Framework
Extend or modify the evaluation by editing the constraint definitions:
- Add dimensions – Modify
EVALUATION_DIMENSIONSinconstraints_commonsense.pyand implement corresponding check functions. - Adjust weights – Change the
weightvalue for any dimension inEVALUATION_DIMENSIONS;calculate_weighted_scoreautomatically recalculates. - Update hard rules – Edit
constraints_hard.pyto add or remove mandatory checks while maintaining the one-vote-veto logic.
Summary
- The benchmark pipeline runs inference, conversion, and evaluation sequentially via
benchmark/deepplanning/travelplanning/run.py. - The
evaluate_plansfunction ineval_converted.pyprocesses structured plans against commonsense and hard constraints using parallel execution. - Commonsense scoring evaluates eight weighted dimensions with a one-vote-veto rule per dimension.
- Hard constraints apply a strict one-vote-veto across all mandatory rules, yielding a binary pass/fail score.
- Final performance is captured by
composite_score(average of both systems) andcase_acc(strict all-pass indicator).
Frequently Asked Questions
What is the difference between commonsense and hard constraint scores?
Commonsense scores measure itinerary quality across eight dimensions (route consistency, cost accuracy, etc.) using weighted averaging where each dimension requires perfect checks to score. Hard constraint scores use a strict one-vote-veto system where any violation of mandatory rules (e.g., missing transport) results in an immediate zero score.
How do I evaluate a custom model configuration?
Add your model configuration to models_config.json, then specify the model name via the --model parameter when running benchmark/deepplanning/travelplanning/run.py. Ensure your model API is compatible with the inference interface used by the pipeline.
What does a case_acc value of 1 indicate?
A case_acc of 1 indicates perfect performance where both the commonsense score equals 1.0 (all eight dimension checks passed) and the hard score equals 1.0 (no mandatory constraint violations), representing a completely valid travel plan according to all criteria.
Can I modify the evaluation weights for different dimensions?
Yes, edit the EVALUATION_DIMENSIONS dictionary in benchmark/deepplanning/travelplanning/evaluation/constraints_commonsense.py to adjust weights for any of the eight dimensions. The calculate_weighted_score function automatically recomputes the final score based on these updated weights.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →