How to Evaluate the Performance of Qwen-Agent: Complete Benchmarking Guide

To evaluate Qwen-Agent's performance, execute the three-stage benchmark pipeline—agent inference, plan conversion, and constraint evaluation—using benchmark/deepplanning/travelplanning/run.py, which calculates a composite score based on eight weighted commonsense dimensions and mandatory hard constraints.

Qwen-Agent is an open-source large language model agent framework developed by QwenLM. Evaluating its performance requires running the specialized travel planning benchmark that measures the agent's ability to generate feasible, cost-aware itineraries. This guide walks through the complete evaluation workflow, from running the benchmark pipeline to interpreting the composite scoring metrics.

Running the Qwen-Agent Benchmark Pipeline

The benchmark evaluates Qwen-Agent through a three-stage pipeline orchestrated by benchmark/deepplanning/travelplanning/run.py. This script handles agent inference, plan conversion, and final evaluation in sequence.

Execute the full pipeline with:

python benchmark/deepplanning/travelplanning/run.py \
    --model qwen-plus \
    --language en \
    --workers 20 \
    --output-dir ./results/qwenplus_en

Key parameters include:

  • --model – The model configuration name (must exist in models_config.json).
  • --language – Test set language (en or zh), which selects the appropriate database and queries.
  • --workers – Parallel threads for inference and evaluation.
  • --output-dir – Directory for intermediate and final results.

The script sequentially calls run_agent_inference for generation, convert_reports for JSON structuring, and evaluate_plans for scoring, as implemented in the orchestration logic of run.py.

Core Evaluation Logic with evaluate_plans

The primary evaluation entry point is evaluate_plans in benchmark/deepplanning/travelplanning/evaluation/eval_converted.py. This function validates converted JSON plans against ground-truth test data and domain-specific databases.

from benchmark.deepplanning.travelplanning.evaluation.eval_converted import evaluate_plans

results = evaluate_plans(
    plan_dir=Path("results/qwenplus_en/converted_plans"),
    test_data_path=Path("benchmark/deepplanning/travelplanning/data/travelplanning_query_en.json"),
    output_dir=Path("results/qwenplus_en/evaluation"),
    database_dir=Path("benchmark/deepplanning/travelplanning/database/database_en"),
)

Inputs required:

  • plan_dir – Directory containing converted JSON plans (one file per test sample).
  • test_data_path – JSON file with metadata (days, org, destination) for validation.
  • database_dir – CSV resources (hotels, restaurants, attractions) used by constraint checkers.

The function processes samples in parallel using ThreadPoolExecutor and returns a list of per-sample dictionaries containing scoring metrics.

Scoring Methodology: Commonsense and Hard Constraints

Qwen-Agent evaluation uses two distinct scoring systems: weighted commonsense dimensions and one-vote-veto hard constraints.

Commonsense Scoring

The calculate_weighted_score function aggregates results across eight dimensions defined in constraints_commonsense.py:

  • Route Consistency
  • Sandbox Compliance
  • Itinerary Structure
  • Time Feasibility
  • Business Hours
  • Duration Rationality
  • Cost Calculation Accuracy
  • Activity Diversity

Each dimension carries a weight of 0.125 and follows a one-vote-veto rule: the dimension scores 1.0 only if all its internal checks pass, otherwise 0.0. The final total_weighted_score is the weighted sum of these dimension scores.

Hard Constraint Scoring

The calculate_hard_score function applies a strict one-vote-veto across all mandatory rules defined in constraints_hard.py. If any hard constraint fails (e.g., missing intercity transport), the score becomes 0.0; otherwise 1.0.

Interpreting Evaluation Results

Each evaluated plan returns a dictionary with the following key metrics:

{
  "sample_id": "42",
  "total_weighted_score": 0.875,
  "dimension_scores": {
    "Route Consistency": 1,
    "Sandbox Compliance": 1,
    "Itinerary Structure": 1,
    "Time Feasibility": 1,
    "Business Hours": 0,
    "Duration Rationality": 1,
    "Cost Calculation Accuracy": 1,
    "Activity Diversity": 1
  },
  "hard_score": 1,
  "composite_score": 0.9375,
  "case_acc": 0,
  "details": {}
}

Metric definitions:

  • total_weighted_score – Weighted commonsense score (0 to 1).
  • hard_score – Binary hard constraint satisfaction (0 or 1).
  • composite_score – Average of commonsense and hard scores.
  • case_acc – Strict all-pass indicator (1 only if both scores equal 1.0).

Use benchmark/deepplanning/travelplanning/evaluation/aggregate_results.py or custom pandas analysis to compute aggregate statistics across the full test set.

Practical Implementation: End-to-End Evaluation Script

When you have existing converted plans, bypass the inference stage and run evaluation directly:

#!/usr/bin/env python

# evaluate_qwen_agent.py

# -------------------------------------------------

# Run the full Qwen-Agent benchmark evaluation.

# -------------------------------------------------

import json
from pathlib import Path
from benchmark.deepplanning.travelplanning.evaluation.eval_converted import evaluate_plans

def main():
    # Configuration

    model = "qwen-plus"
    language = "en"
    base_dir = Path(__file__).parent
    results_dir = base_dir / "results" / f"{model}_{language}"

    # Paths

    plan_dir = results_dir / "converted_plans"
    test_data = base_dir / "benchmark" / "deepplanning" / "travelplanning" / "data" / f"travelplanning_query_{language}.json"
    database_dir = base_dir / "benchmark" / "deepplanning" / "travelplanning" / "database" / f"database_{language}"
    out_dir = results_dir / "evaluation"

    # Execute evaluation

    all_results = evaluate_plans(
        plan_dir=plan_dir,
        test_data_path=test_data,
        output_dir=out_dir,
        database_dir=database_dir,
    )

    # Compute summary statistics

    composite_scores = [r["composite_score"] for r in all_results]
    case_accs = [r["case_acc"] for r in all_results]

    print("\n=== Qwen-Agent Benchmark Summary ===")
    print(f"Samples evaluated: {len(all_results)}")
    print(f"Avg. composite: {sum(composite_scores)/len(composite_scores):.3f}")
    print(f"Overall case acc: {sum(case_accs)/len(case_accs):.3%}")

    # Persist results

    out_dir.mkdir(parents=True, exist_ok=True)
    (out_dir / "full_results.json").write_text(json.dumps(all_results, indent=2))

if __name__ == "__main__":
    main()

Run this script after generating the converted_plans directory to obtain quantitative performance metrics.

Customizing the Evaluation Framework

Extend or modify the evaluation by editing the constraint definitions:

  • Add dimensions – Modify EVALUATION_DIMENSIONS in constraints_commonsense.py and implement corresponding check functions.
  • Adjust weights – Change the weight value for any dimension in EVALUATION_DIMENSIONS; calculate_weighted_score automatically recalculates.
  • Update hard rules – Edit constraints_hard.py to add or remove mandatory checks while maintaining the one-vote-veto logic.

Summary

  • The benchmark pipeline runs inference, conversion, and evaluation sequentially via benchmark/deepplanning/travelplanning/run.py.
  • The evaluate_plans function in eval_converted.py processes structured plans against commonsense and hard constraints using parallel execution.
  • Commonsense scoring evaluates eight weighted dimensions with a one-vote-veto rule per dimension.
  • Hard constraints apply a strict one-vote-veto across all mandatory rules, yielding a binary pass/fail score.
  • Final performance is captured by composite_score (average of both systems) and case_acc (strict all-pass indicator).

Frequently Asked Questions

What is the difference between commonsense and hard constraint scores?

Commonsense scores measure itinerary quality across eight dimensions (route consistency, cost accuracy, etc.) using weighted averaging where each dimension requires perfect checks to score. Hard constraint scores use a strict one-vote-veto system where any violation of mandatory rules (e.g., missing transport) results in an immediate zero score.

How do I evaluate a custom model configuration?

Add your model configuration to models_config.json, then specify the model name via the --model parameter when running benchmark/deepplanning/travelplanning/run.py. Ensure your model API is compatible with the inference interface used by the pipeline.

What does a case_acc value of 1 indicate?

A case_acc of 1 indicates perfect performance where both the commonsense score equals 1.0 (all eight dimension checks passed) and the hard score equals 1.0 (no mandatory constraint violations), representing a completely valid travel plan according to all criteria.

Can I modify the evaluation weights for different dimensions?

Yes, edit the EVALUATION_DIMENSIONS dictionary in benchmark/deepplanning/travelplanning/evaluation/constraints_commonsense.py to adjust weights for any of the eight dimensions. The calculate_weighted_score function automatically recomputes the final score based on these updated weights.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →