# How to Evaluate AI Agent Performance Using OSWorld, SWE-bench, and GAIA Benchmarks

> Learn how to evaluate AI agent performance using OSWorld, SWE-bench, and GAIA benchmarks. Discover key metrics and methods for assessing AI capabilities.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-17

---

**The bojieli/ai-agent-book repository evaluates AI agent performance by running stratified samples of easy, medium, and hard tasks across three major benchmark suites—OSWorld, SWE-bench, and GAIA—and aggregating binary pass/fail outcomes into pass-rate metrics.**

Modern autonomous agents must prove competence across diverse environments, from desktop operating systems to software engineering workflows and multi-modal reasoning tasks. The evaluation framework documented in this open-source book provides a reproducible methodology for measuring end-to-end agent capabilities using industry-standard benchmarks. By leveraging Docker-based evaluators and task-level verification, the repository captures whether agents actually complete objectives rather than merely generating plausible-looking outputs.

## The Three Core Benchmarks for AI Agent Evaluation

The repository focuses on three complementary benchmark families that cover distinct aspects of agent intelligence. Each suite provides official evaluators that return binary success signals and numeric rewards.

### OSWorld-Verified: Desktop OS Automation

**OSWorld-Verified** tests full-desktop operating system automation, requiring agents to manipulate file systems, navigate UIs, and configure network settings. Tasks range from simple file operations to complex multi-step workflows involving application installation and system configuration. According to the experiment results in [`chapter7/experiment-6-2-human-benchmark/results.json`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/experiment-6-2-human-benchmark/results.json), this benchmark achieved a **100% pass-rate** across all difficulty tiers in the sample experiment.

### SWE-bench Verified: Software Engineering Tasks

**SWE-bench Verified** evaluates code-fix capabilities by requiring agents to generate patches that resolve real GitHub issues. The benchmark uses containerized environments to verify that proposed changes actually fix the reported bug without breaking existing functionality. In the sample run documented in the results file, SWE-bench achieved a **66.67% pass-rate** (2/3 tasks), with failures concentrated in the hard difficulty tier due to hidden semantic contract violations.

### GAIA: Multi-Modal Reasoning and Tool Use

**GAIA** assesses multi-modal reasoning across text, images, and API interactions, requiring agents to gather evidence from diverse sources and synthesize correct answers. Like the other benchmarks, it employs a binary verification system where the **official_reward** equals `1.0` only when the final artifact exactly matches ground truth criteria. The sample experiment shows GAIA scoring **66.67%** (2/3 tasks), with specific failure modes emerging around data normalization requirements.

## Evaluation Metrics and Methodology

The repository employs a strict task-level evaluation protocol that prioritizes verifiable correctness over probabilistic accuracy.

### Task-Level Binary Metrics

Each task execution produces two primary signals:

- **Pass/Fail**: A binary outcome returned by the official benchmark evaluator
- **official_reward**: A numeric score where `1.0` indicates success and `0.0` indicates failure

These metrics are recorded in [`chapter7/experiment-6-2-human-benchmark/results.json`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/experiment-6-2-human-benchmark/results.json) alongside task metadata and execution timestamps. Unlike model-level metrics such as "Pass@1" that sample multiple attempts, these benchmarks require the agent to succeed on the first complete trajectory.

### Aggregated Pass-Rate Calculation

The primary comparative metric is the **pass-rate**, calculated as `passed / total_cases`. In the human-baseline experiment documented in the repository, the overall pass-rate across all three benchmarks was **13/18 ≈ 72.2%**. This aggregation allows direct comparison across benchmark families despite their different domain semantics.

### Stratified Sampling by Difficulty Tier

To surface meaningful failure patterns, the experiment selects one easy, one medium, and one hard task per benchmark before execution begins. This stratified approach prevents cherry-picking and ensures coverage of the full difficulty spectrum. The specific task selections are cryptographically locked in [`chapter7/experiment-6-2-human-benchmark/selection_manifest.json`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/experiment-6-2-human-benchmark/selection_manifest.json) using SHA-256 hashes to guarantee reproducibility.

## How the Evaluation Pipeline Works

The repository implements a three-phase evaluation pipeline that separates task setup, agent execution, and verification.

### Task Selection and Manifest Locking

Before any agent runs, the system randomly selects tasks from each difficulty tier and records them in [`selection_manifest.json`](https://github.com/bojieli/ai-agent-book/blob/main/selection_manifest.json). This file serves as the source of truth for which specific tasks constitute the evaluation set, preventing selection bias or task swapping during experiments.

### Human-Operator Execution and Trajectory Logging

During execution, the agent (in this case, a human operator using Codex) interacts with the target environment through native toolchains:

- **OSWorld**: pyautogui for UI automation within KVM virtual machines
- **SWE-bench**: Git operations and code editing within Docker containers  
- **GAIA**: HTTP API calls and file manipulation on dedicated servers

The system logs all actions, tool calls, and intermediate observations as **trajectories**, creating a complete audit trail of how the agent approached each task.

### Official Docker-Based Evaluators

Each benchmark provides a containerized evaluator that consumes the final artifact—whether a UI state, Git patch, or data file—and returns a strict binary **pass** flag. The evaluator runs exactly once per task in isolation, and its verdict is frozen in the results JSON. This external verification step ensures that agents cannot game the metric through output formatting tricks or partial solutions.

## Analyzing Results from the Experiment

You can replicate the metric calculations using the raw JSON output. The following Python script loads [`chapter7/experiment-6-2-human-benchmark/results.json`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/experiment-6-2-human-benchmark/results.json) and computes per-benchmark pass-rates:

```python
import json
from pathlib import Path
from collections import defaultdict

# Load the experiment JSON

RESULTS_PATH = Path(
    "chapter7/experiment-6-2-human-benchmark/results.json"
)
with RESULTS_PATH.open() as f:
    data = json.load(f)

# Aggregate pass/fail per benchmark

stats = defaultdict(lambda: {"passed": 0, "cases": 0})
for case in data["cases"]:
    bench = case["benchmark"]
    stats[bench]["cases"] += 1
    if case["result"] == "passed":
        stats[bench]["passed"] += 1

# Print a tidy table

print("Benchmark | Pass | Cases | Pass‑rate")
print("-" * 40)
for bench, vals in stats.items():
    rate = vals["passed"] / vals["cases"]
    print(f"{bench:10} | {vals['passed']:4} | {vals['cases']:5} | {rate:.2%}")

```

Running this script against the repository data yields the tier-wise breakdowns, revealing that OSWorld maintained perfect scores while SWE-bench struggled with hard-tier tasks involving ABI mismatches.

## Why Stratified Sampling Reveals Hidden Failure Modes

Aggregate pass-rates alone can mask critical competence gaps. The stratified approach in [`chapter7/experiment-6-2-human-benchmark/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/experiment-6-2-human-benchmark/README.md) exposes specific failure categories invisible in random sampling:

- **Normalization issues** in GAIA medium tasks where rounding rules determine success
- **UI termination problems** in OSWorld hard tasks involving AndroidWorld state management  
- **Hidden semantic contracts** in SWE-bench hard patches where function signatures appear correct but violate internal ABI requirements

These insights guide targeted improvements to agent prompting strategies, tool-call schemas, and environment interaction logic.

## Summary

- **OSWorld, SWE-bench, and GAIA** provide complementary coverage of OS automation, software engineering, and multi-modal reasoning capabilities.
- The **official_reward** and **pass/fail** metrics in [`results.json`](https://github.com/bojieli/ai-agent-book/blob/main/results.json) capture end-to-end task completion rather than intermediate step accuracy.
- **Stratified sampling** (easy/medium/hard) surfaces failure modes that aggregate scores obscure.
- **Docker-based evaluators** provide deterministic, reproducible verification independent of the agent's execution environment.
- The **72.2% overall pass-rate** in the sample experiment demonstrates current state-of-the-art performance while highlighting specific areas needing improvement.

## Frequently Asked Questions

### What distinguishes OSWorld from SWE-bench in evaluating agent capabilities?

OSWorld evaluates desktop environment manipulation requiring visual perception and UI interaction through tools like pyautogui, while SWE-bench tests code generation and repository management within containerized development environments. OSWorld tasks verify final system state, whereas SWE-bench executes test suites to validate patch correctness.

### How is the pass-rate metric calculated across different benchmarks?

The pass-rate equals the number of tasks marked `"passed"` divided by total cases attempted, as recorded in [`chapter7/experiment-6-2-human-benchmark/results.json`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/experiment-6-2-human-benchmark/results.json). Each benchmark's official evaluator returns a binary success flag that contributes equally to this aggregate metric, regardless of task domain differences.

### Why does the repository use stratified sampling instead of random task selection?

Stratified sampling ensures coverage across the full difficulty spectrum by mandating one easy, one medium, and one hard task per benchmark. This method, locked in [`selection_manifest.json`](https://github.com/bojieli/ai-agent-book/blob/main/selection_manifest.json), prevents algorithms from accidentally skipping hard tasks where failure modes typically concentrate, providing actionable insights for capability improvements.

### Where can I access the raw evaluation results and trajectories?

Machine-readable results are stored in [`chapter7/experiment-6-2-human-benchmark/results.json`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/experiment-6-2-human-benchmark/results.json), while methodological details and tier-wise analysis appear in the corresponding [`README.md`](https://github.com/bojieli/ai-agent-book/blob/main/README.md). The [`selection_manifest.json`](https://github.com/bojieli/ai-agent-book/blob/main/selection_manifest.json) file contains the SHA-256-locked task assignments ensuring experiment reproducibility.