How Cua-Bench Evaluates Agents on OSWorld, ScreenSpot, and Windows Arena Benchmarks

Cua-Bench executes agents inside specialized sandboxes—QEMU VMs for OSWorld, static screenshot datasets for ScreenSpot, and Docker-based Windows containers for Windows Arena—then computes success metrics through task verifiers, IoU calculations, and end-to-end completion checks.

Cua-Bench is the built-in benchmarking suite for the Cua (Computer-Use Agent) framework. It provides a unified evaluation pipeline that provisions the appropriate runtime environment for each benchmark family, feeds screenshots to the agent, and aggregates performance statistics using benchmark-specific scoring logic.

OSWorld Evaluation via QEMU/KVM Transport

OSWorld tests full-desktop automation across 369+ Linux tasks including browsing, document editing, and coding. Cua-Bench evaluates agents on OSWorld by launching a QEMU/KVM virtual machine that runs the official OSWorld Flask server.

VM Provisioning and the OSWorld Transport Layer

Cua-Bench downloads the OSWorld Ubuntu .qcow2 image from HuggingFace and caches it locally. The VM is started with the OSWorld Flask server pre-installed. Agent communication flows through OSWorldTransport (libs/python/cua-sandbox/cua_sandbox/transport/osworld.py), which implements two critical endpoints:

  • /screenshot – Returns a PNG of the current desktop state.
  • /execute_action – Receives JSON actions (click, type, scroll) and executes them via pyautogui inside the VM.

This transport layer abstracts the VM boundary, allowing the generic agent loop to treat the remote desktop as a local environment.

The Agent Loop and Task Verification

The ComputerAgent (or any subclass) is instantiated with the OSWorld transport. The evaluation loop proceeds as follows:

  1. The agent receives a screenshot via the transport.
  2. It predicts an action and sends it back via /execute_action.
  3. The loop continues until the HUD task reports success or a step limit is reached.

Each OSWorld task includes a verifier function that checks specific success conditions—such as verifying a document was saved, a webpage opened, or a UI element exists. Cua-Bench records whether the verifier passed and aggregates the success rate, step count, and wall-clock time.

Running OSWorld Benchmarks

Execute OSWorld evaluation using the Cua-Bench CLI:

cua-bench run tasks/osworld --agent opencua --output osworld-results/

This command invokes the cb run entry point (libs/python/cua-bench/cua_bench/cli/main.py), which constructs the QEMU sandbox, loads the OSWorld tasks, and drives the evaluation loop.

ScreenSpot v2 and Pro Evaluation

ScreenSpot benchmarks test GUI grounding and click-prediction accuracy on static screenshots. Cua-Bench evaluates agents on both ScreenSpot-v2 and ScreenSpot-Pro (high-resolution variant) using dataset-driven scripts.

Dataset Loading and Click Prediction

The benchmark drivers (ss-v2.py and ss-pro.py in libs/python/agent/benchmarks/) load datasets via HuggingFace:

  • lmms-lab/ScreenSpot-v2
  • lmms-lab/ScreenSpot-Pro

For each sample, the script feeds the screenshot to the agent's predict_click method, which returns an (x, y) coordinate representing the predicted click location.

IoU-Based Accuracy Metrics

Cua-Bench converts the predicted coordinate into a 1×1 pixel bounding box and computes the Intersection-over-Union (IoU) against the ground-truth bounding box. A prediction is counted as correct if IoU ≥ 0.5. The benchmark computes overall accuracy and per-class breakdowns.

Results are written to Markdown reports (screenspot_v2_results.md or screenspot_pro_results.md) in the output folder. Shared utilities for IoU calculation and aggregation live in libs/python/agent/benchmarks/utils.py.

ScreenSpot Benchmark Execution

Run ScreenSpot evaluations with specific task targets:


# ScreenSpot v2

cua-bench run tasks/screenspot-v2 --agent opencua --output ss-v2-results/

# ScreenSpot-Pro (high-res)

cua-bench run tasks/screenspot-pro --agent opencua --output ss-pro-results/

Windows Arena Evaluation on Dockerized Windows VMs

Windows Arena tests end-to-end tasks on Windows 10 environments across applications like Chrome, LibreOffice, and VLC. Cua-Bench evaluates agents on Windows Arena using a containerized Windows VM with a custom control server.

WinArena Container Architecture

The trycua/winarena:latest Docker image contains a Windows 10 VM plus the WinArena server, which exposes the same /status and /screenshot API as the Linux computer-server on port 8000. The WAASetupController (libs/cua-bench/tasks/winarena_adapter/setup_controller.py) handles VM provisioning, including optional Windows ISO downloads and application pre-installation (--install-apps).

Agent communication uses the standard ComputerTransport (libs/python/cua-sandbox/cua_sandbox/transport/websocket.py) or VNC-SSH transport to interact with the WinArena server.

Task Evaluators and Success Criteria

Task definitions reside in libs/cua-bench/tasks/winarena_adapter/evaluators/. Each evaluator executes a series of actions (clicks, typing, file operations) and verifies the result—such as confirming a PDF was saved or a video played successfully.

Cua-Bench collects success flags and reports task-success rate, average steps, and elapsed time in a formatted table.

Running Windows Arena Benchmarks

Launch Windows Arena evaluation with setup flags:

cua-bench run tasks/winarena_adapter \
    --agent opencua \
    --setup --download-iso \
    --output winarena-results/

The --setup flag triggers the WAASetupController to provision the container, while --download-iso fetches the Windows installation media if not present.

Unified CLI and Shared Infrastructure

All three benchmark families share the same Cua-Bench CLI architecture. The cb run command parses the task name and constructs the appropriate Sandbox abstraction (libs/python/cua-sandbox/cua_sandbox/sandbox.py), which selects between OSWorldTransport (Linux QEMU), standard ComputerTransport (Windows container), or direct dataset loading (ScreenSpot).

Platform selection logic is defined in libs/python/cua-bench/cua_bench/cli/commands/platform.py, which maps CLI flags like --platform linux-qemu or --platform winarena to the correct transport and sandbox implementations.

Summary

  • OSWorld: Cua-Bench provisions a QEMU VM running the OSWorld Flask server, uses OSWorldTransport to exchange screenshots and actions, and scores tasks via verifier functions that check desktop state.
  • ScreenSpot: Static datasets drive click-prediction evaluation; accuracy is measured by IoU ≥ 0.5 between predicted and ground-truth bounding boxes, with results aggregated in Markdown reports.
  • Windows Arena: A Docker-based Windows VM runs the WinArena server on port 8000; evaluators execute actions and verify success criteria across real Windows applications.
  • Unified Interface: The cua-bench run command handles all three benchmarks through a shared entry point and sandbox abstraction layer.

Frequently Asked Questions

How does Cua-Bench handle action execution inside the OSWorld VM?

Cua-Bench sends actions as JSON payloads to the /execute_action endpoint of the OSWorld Flask server, which translates them into pyautogui commands executed directly on the Ubuntu desktop.

What metric determines success in ScreenSpot evaluations?

ScreenSpot uses Intersection-over-Union (IoU). The agent's predicted (x, y) coordinate is converted to a 1×1 box; if the IoU with the ground-truth bounding box is 0.5 or higher, the prediction is marked correct.

Can Cua-Bench evaluate agents on Windows Arena without pre-installing applications?

Yes. The WAASetupController supports the --install-apps flag during setup, which automatically pre-installs benchmark applications like Chrome and LibreOffice into the Windows container before evaluation begins.

Where does Cua-Bench store benchmark results and logs?

Results are written to the directory specified by the --output flag. OSWorld and Windows Arena produce success-rate tables and timing statistics, while ScreenSpot generates Markdown reports (screenspot_v2_results.md or screenspot_pro_results.md) containing per-class accuracy breakdowns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →