# How to Evaluate the Performance of Qwen-Agent: Complete Benchmarking Guide

> Learn how to evaluate Qwen-Agent performance with our comprehensive benchmarking guide. Explore the three-stage pipeline including agent inference, plan conversion, and constraint evaluation for a detailed analysis.

- Repository: [Qwen/Qwen-Agent](https://github.com/qwenlm/Qwen-Agent)
- Tags: performance
- Published: 2026-03-09

---

**To evaluate Qwen-Agent's performance, execute the three-stage benchmark pipeline—agent inference, plan conversion, and constraint evaluation—using [`benchmark/deepplanning/travelplanning/run.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/travelplanning/run.py), which calculates a composite score based on eight weighted commonsense dimensions and mandatory hard constraints.**

Qwen-Agent is an open-source large language model agent framework developed by QwenLM. Evaluating its performance requires running the specialized travel planning benchmark that measures the agent's ability to generate feasible, cost-aware itineraries. This guide walks through the complete evaluation workflow, from running the benchmark pipeline to interpreting the composite scoring metrics.

## Running the Qwen-Agent Benchmark Pipeline

The benchmark evaluates Qwen-Agent through a three-stage pipeline orchestrated by [`benchmark/deepplanning/travelplanning/run.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/travelplanning/run.py). This script handles agent inference, plan conversion, and final evaluation in sequence.

Execute the full pipeline with:

```bash
python benchmark/deepplanning/travelplanning/run.py \
    --model qwen-plus \
    --language en \
    --workers 20 \
    --output-dir ./results/qwenplus_en

```

Key parameters include:
- **`--model`** – The model configuration name (must exist in [`models_config.json`](https://github.com/QwenLM/Qwen-Agent/blob/main/models_config.json)).
- **`--language`** – Test set language (`en` or `zh`), which selects the appropriate database and queries.
- **`--workers`** – Parallel threads for inference and evaluation.
- **`--output-dir`** – Directory for intermediate and final results.

The script sequentially calls `run_agent_inference` for generation, `convert_reports` for JSON structuring, and `evaluate_plans` for scoring, as implemented in the orchestration logic of **run.py**.

## Core Evaluation Logic with `evaluate_plans`

The primary evaluation entry point is `evaluate_plans` in [`benchmark/deepplanning/travelplanning/evaluation/eval_converted.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/travelplanning/evaluation/eval_converted.py). This function validates converted JSON plans against ground-truth test data and domain-specific databases.

```python
from benchmark.deepplanning.travelplanning.evaluation.eval_converted import evaluate_plans

results = evaluate_plans(
    plan_dir=Path("results/qwenplus_en/converted_plans"),
    test_data_path=Path("benchmark/deepplanning/travelplanning/data/travelplanning_query_en.json"),
    output_dir=Path("results/qwenplus_en/evaluation"),
    database_dir=Path("benchmark/deepplanning/travelplanning/database/database_en"),
)

```

**Inputs required:**
- **plan_dir** – Directory containing converted JSON plans (one file per test sample).
- **test_data_path** – JSON file with metadata (`days`, `org`, destination) for validation.
- **database_dir** – CSV resources (hotels, restaurants, attractions) used by constraint checkers.

The function processes samples in parallel using `ThreadPoolExecutor` and returns a list of per-sample dictionaries containing scoring metrics.

## Scoring Methodology: Commonsense and Hard Constraints

Qwen-Agent evaluation uses two distinct scoring systems: **weighted commonsense dimensions** and **one-vote-veto hard constraints**.

### Commonsense Scoring

The `calculate_weighted_score` function aggregates results across **eight dimensions** defined in [`constraints_commonsense.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/constraints_commonsense.py):
- Route Consistency
- Sandbox Compliance
- Itinerary Structure
- Time Feasibility
- Business Hours
- Duration Rationality
- Cost Calculation Accuracy
- Activity Diversity

Each dimension carries a weight of `0.125` and follows a **one-vote-veto rule**: the dimension scores `1.0` only if all its internal checks pass, otherwise `0.0`. The final `total_weighted_score` is the weighted sum of these dimension scores.

### Hard Constraint Scoring

The `calculate_hard_score` function applies a strict one-vote-veto across all mandatory rules defined in [`constraints_hard.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/constraints_hard.py). If any hard constraint fails (e.g., missing intercity transport), the score becomes `0.0`; otherwise `1.0`.

## Interpreting Evaluation Results

Each evaluated plan returns a dictionary with the following key metrics:

```json
{
  "sample_id": "42",
  "total_weighted_score": 0.875,
  "dimension_scores": {
    "Route Consistency": 1,
    "Sandbox Compliance": 1,
    "Itinerary Structure": 1,
    "Time Feasibility": 1,
    "Business Hours": 0,
    "Duration Rationality": 1,
    "Cost Calculation Accuracy": 1,
    "Activity Diversity": 1
  },
  "hard_score": 1,
  "composite_score": 0.9375,
  "case_acc": 0,
  "details": {}
}

```

**Metric definitions:**
- **`total_weighted_score`** – Weighted commonsense score (0 to 1).
- **`hard_score`** – Binary hard constraint satisfaction (0 or 1).
- **`composite_score`** – Average of commonsense and hard scores.
- **`case_acc`** – Strict all-pass indicator (1 only if both scores equal 1.0).

Use [`benchmark/deepplanning/travelplanning/evaluation/aggregate_results.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/travelplanning/evaluation/aggregate_results.py) or custom pandas analysis to compute aggregate statistics across the full test set.

## Practical Implementation: End-to-End Evaluation Script

When you have existing converted plans, bypass the inference stage and run evaluation directly:

```python
#!/usr/bin/env python

# evaluate_qwen_agent.py

# -------------------------------------------------

# Run the full Qwen-Agent benchmark evaluation.

# -------------------------------------------------

import json
from pathlib import Path
from benchmark.deepplanning.travelplanning.evaluation.eval_converted import evaluate_plans

def main():
    # Configuration

    model = "qwen-plus"
    language = "en"
    base_dir = Path(__file__).parent
    results_dir = base_dir / "results" / f"{model}_{language}"

    # Paths

    plan_dir = results_dir / "converted_plans"
    test_data = base_dir / "benchmark" / "deepplanning" / "travelplanning" / "data" / f"travelplanning_query_{language}.json"
    database_dir = base_dir / "benchmark" / "deepplanning" / "travelplanning" / "database" / f"database_{language}"
    out_dir = results_dir / "evaluation"

    # Execute evaluation

    all_results = evaluate_plans(
        plan_dir=plan_dir,
        test_data_path=test_data,
        output_dir=out_dir,
        database_dir=database_dir,
    )

    # Compute summary statistics

    composite_scores = [r["composite_score"] for r in all_results]
    case_accs = [r["case_acc"] for r in all_results]

    print("\n=== Qwen-Agent Benchmark Summary ===")
    print(f"Samples evaluated: {len(all_results)}")
    print(f"Avg. composite: {sum(composite_scores)/len(composite_scores):.3f}")
    print(f"Overall case acc: {sum(case_accs)/len(case_accs):.3%}")

    # Persist results

    out_dir.mkdir(parents=True, exist_ok=True)
    (out_dir / "full_results.json").write_text(json.dumps(all_results, indent=2))

if __name__ == "__main__":
    main()

```

Run this script after generating the `converted_plans` directory to obtain quantitative performance metrics.

## Customizing the Evaluation Framework

Extend or modify the evaluation by editing the constraint definitions:

- **Add dimensions** – Modify `EVALUATION_DIMENSIONS` in [`constraints_commonsense.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/constraints_commonsense.py) and implement corresponding check functions.
- **Adjust weights** – Change the `weight` value for any dimension in `EVALUATION_DIMENSIONS`; `calculate_weighted_score` automatically recalculates.
- **Update hard rules** – Edit [`constraints_hard.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/constraints_hard.py) to add or remove mandatory checks while maintaining the one-vote-veto logic.

## Summary

- The benchmark pipeline runs inference, conversion, and evaluation sequentially via [`benchmark/deepplanning/travelplanning/run.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/travelplanning/run.py).
- The `evaluate_plans` function in [`eval_converted.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/eval_converted.py) processes structured plans against commonsense and hard constraints using parallel execution.
- **Commonsense scoring** evaluates eight weighted dimensions with a one-vote-veto rule per dimension.
- **Hard constraints** apply a strict one-vote-veto across all mandatory rules, yielding a binary pass/fail score.
- Final performance is captured by `composite_score` (average of both systems) and `case_acc` (strict all-pass indicator).

## Frequently Asked Questions

### What is the difference between commonsense and hard constraint scores?

**Commonsense scores** measure itinerary quality across eight dimensions (route consistency, cost accuracy, etc.) using weighted averaging where each dimension requires perfect checks to score. **Hard constraint scores** use a strict one-vote-veto system where any violation of mandatory rules (e.g., missing transport) results in an immediate zero score.

### How do I evaluate a custom model configuration?

Add your model configuration to [`models_config.json`](https://github.com/QwenLM/Qwen-Agent/blob/main/models_config.json), then specify the model name via the `--model` parameter when running [`benchmark/deepplanning/travelplanning/run.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/travelplanning/run.py). Ensure your model API is compatible with the inference interface used by the pipeline.

### What does a `case_acc` value of 1 indicate?

A `case_acc` of 1 indicates perfect performance where both the commonsense score equals 1.0 (all eight dimension checks passed) and the hard score equals 1.0 (no mandatory constraint violations), representing a completely valid travel plan according to all criteria.

### Can I modify the evaluation weights for different dimensions?

Yes, edit the `EVALUATION_DIMENSIONS` dictionary in [`benchmark/deepplanning/travelplanning/evaluation/constraints_commonsense.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/benchmark/deepplanning/travelplanning/evaluation/constraints_commonsense.py) to adjust weights for any of the eight dimensions. The `calculate_weighted_score` function automatically recomputes the final score based on these updated weights.