How to Evaluate AI Agent Performance Using OSWorld, SWE-bench, and GAIA Benchmarks

The bojieli/ai-agent-book repository evaluates AI agent performance by running stratified samples of easy, medium, and hard tasks across three major benchmark suites—OSWorld, SWE-bench, and GAIA—and aggregating binary pass/fail outcomes into pass-rate metrics.

Modern autonomous agents must prove competence across diverse environments, from desktop operating systems to software engineering workflows and multi-modal reasoning tasks. The evaluation framework documented in this open-source book provides a reproducible methodology for measuring end-to-end agent capabilities using industry-standard benchmarks. By leveraging Docker-based evaluators and task-level verification, the repository captures whether agents actually complete objectives rather than merely generating plausible-looking outputs.

The Three Core Benchmarks for AI Agent Evaluation

The repository focuses on three complementary benchmark families that cover distinct aspects of agent intelligence. Each suite provides official evaluators that return binary success signals and numeric rewards.

OSWorld-Verified: Desktop OS Automation

OSWorld-Verified tests full-desktop operating system automation, requiring agents to manipulate file systems, navigate UIs, and configure network settings. Tasks range from simple file operations to complex multi-step workflows involving application installation and system configuration. According to the experiment results in chapter7/experiment-6-2-human-benchmark/results.json, this benchmark achieved a 100% pass-rate across all difficulty tiers in the sample experiment.

SWE-bench Verified: Software Engineering Tasks

SWE-bench Verified evaluates code-fix capabilities by requiring agents to generate patches that resolve real GitHub issues. The benchmark uses containerized environments to verify that proposed changes actually fix the reported bug without breaking existing functionality. In the sample run documented in the results file, SWE-bench achieved a 66.67% pass-rate (2/3 tasks), with failures concentrated in the hard difficulty tier due to hidden semantic contract violations.

GAIA: Multi-Modal Reasoning and Tool Use

GAIA assesses multi-modal reasoning across text, images, and API interactions, requiring agents to gather evidence from diverse sources and synthesize correct answers. Like the other benchmarks, it employs a binary verification system where the official_reward equals 1.0 only when the final artifact exactly matches ground truth criteria. The sample experiment shows GAIA scoring 66.67% (2/3 tasks), with specific failure modes emerging around data normalization requirements.

Evaluation Metrics and Methodology

The repository employs a strict task-level evaluation protocol that prioritizes verifiable correctness over probabilistic accuracy.

Task-Level Binary Metrics

Each task execution produces two primary signals:

  • Pass/Fail: A binary outcome returned by the official benchmark evaluator
  • official_reward: A numeric score where 1.0 indicates success and 0.0 indicates failure

These metrics are recorded in chapter7/experiment-6-2-human-benchmark/results.json alongside task metadata and execution timestamps. Unlike model-level metrics such as "Pass@1" that sample multiple attempts, these benchmarks require the agent to succeed on the first complete trajectory.

Aggregated Pass-Rate Calculation

The primary comparative metric is the pass-rate, calculated as passed / total_cases. In the human-baseline experiment documented in the repository, the overall pass-rate across all three benchmarks was 13/18 ≈ 72.2%. This aggregation allows direct comparison across benchmark families despite their different domain semantics.

Stratified Sampling by Difficulty Tier

To surface meaningful failure patterns, the experiment selects one easy, one medium, and one hard task per benchmark before execution begins. This stratified approach prevents cherry-picking and ensures coverage of the full difficulty spectrum. The specific task selections are cryptographically locked in chapter7/experiment-6-2-human-benchmark/selection_manifest.json using SHA-256 hashes to guarantee reproducibility.

How the Evaluation Pipeline Works

The repository implements a three-phase evaluation pipeline that separates task setup, agent execution, and verification.

Task Selection and Manifest Locking

Before any agent runs, the system randomly selects tasks from each difficulty tier and records them in selection_manifest.json. This file serves as the source of truth for which specific tasks constitute the evaluation set, preventing selection bias or task swapping during experiments.

Human-Operator Execution and Trajectory Logging

During execution, the agent (in this case, a human operator using Codex) interacts with the target environment through native toolchains:

  • OSWorld: pyautogui for UI automation within KVM virtual machines
  • SWE-bench: Git operations and code editing within Docker containers
  • GAIA: HTTP API calls and file manipulation on dedicated servers

The system logs all actions, tool calls, and intermediate observations as trajectories, creating a complete audit trail of how the agent approached each task.

Official Docker-Based Evaluators

Each benchmark provides a containerized evaluator that consumes the final artifact—whether a UI state, Git patch, or data file—and returns a strict binary pass flag. The evaluator runs exactly once per task in isolation, and its verdict is frozen in the results JSON. This external verification step ensures that agents cannot game the metric through output formatting tricks or partial solutions.

Analyzing Results from the Experiment

You can replicate the metric calculations using the raw JSON output. The following Python script loads chapter7/experiment-6-2-human-benchmark/results.json and computes per-benchmark pass-rates:

import json
from pathlib import Path
from collections import defaultdict

# Load the experiment JSON

RESULTS_PATH = Path(
    "chapter7/experiment-6-2-human-benchmark/results.json"
)
with RESULTS_PATH.open() as f:
    data = json.load(f)

# Aggregate pass/fail per benchmark

stats = defaultdict(lambda: {"passed": 0, "cases": 0})
for case in data["cases"]:
    bench = case["benchmark"]
    stats[bench]["cases"] += 1
    if case["result"] == "passed":
        stats[bench]["passed"] += 1

# Print a tidy table

print("Benchmark | Pass | Cases | Pass‑rate")
print("-" * 40)
for bench, vals in stats.items():
    rate = vals["passed"] / vals["cases"]
    print(f"{bench:10} | {vals['passed']:4} | {vals['cases']:5} | {rate:.2%}")

Running this script against the repository data yields the tier-wise breakdowns, revealing that OSWorld maintained perfect scores while SWE-bench struggled with hard-tier tasks involving ABI mismatches.

Why Stratified Sampling Reveals Hidden Failure Modes

Aggregate pass-rates alone can mask critical competence gaps. The stratified approach in chapter7/experiment-6-2-human-benchmark/README.md exposes specific failure categories invisible in random sampling:

  • Normalization issues in GAIA medium tasks where rounding rules determine success
  • UI termination problems in OSWorld hard tasks involving AndroidWorld state management
  • Hidden semantic contracts in SWE-bench hard patches where function signatures appear correct but violate internal ABI requirements

These insights guide targeted improvements to agent prompting strategies, tool-call schemas, and environment interaction logic.

Summary

  • OSWorld, SWE-bench, and GAIA provide complementary coverage of OS automation, software engineering, and multi-modal reasoning capabilities.
  • The official_reward and pass/fail metrics in results.json capture end-to-end task completion rather than intermediate step accuracy.
  • Stratified sampling (easy/medium/hard) surfaces failure modes that aggregate scores obscure.
  • Docker-based evaluators provide deterministic, reproducible verification independent of the agent's execution environment.
  • The 72.2% overall pass-rate in the sample experiment demonstrates current state-of-the-art performance while highlighting specific areas needing improvement.

Frequently Asked Questions

What distinguishes OSWorld from SWE-bench in evaluating agent capabilities?

OSWorld evaluates desktop environment manipulation requiring visual perception and UI interaction through tools like pyautogui, while SWE-bench tests code generation and repository management within containerized development environments. OSWorld tasks verify final system state, whereas SWE-bench executes test suites to validate patch correctness.

How is the pass-rate metric calculated across different benchmarks?

The pass-rate equals the number of tasks marked "passed" divided by total cases attempted, as recorded in chapter7/experiment-6-2-human-benchmark/results.json. Each benchmark's official evaluator returns a binary success flag that contributes equally to this aggregate metric, regardless of task domain differences.

Why does the repository use stratified sampling instead of random task selection?

Stratified sampling ensures coverage across the full difficulty spectrum by mandating one easy, one medium, and one hard task per benchmark. This method, locked in selection_manifest.json, prevents algorithms from accidentally skipping hard tasks where failure modes typically concentrate, providing actionable insights for capability improvements.

Where can I access the raw evaluation results and trajectories?

Machine-readable results are stored in chapter7/experiment-6-2-human-benchmark/results.json, while methodological details and tier-wise analysis appear in the corresponding README.md. The selection_manifest.json file contains the SHA-256-locked task assignments ensuring experiment reproducibility.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →