# How Cua-Bench Evaluates Agents on OSWorld, ScreenSpot, and Windows Arena Benchmarks

> Cua-Bench evaluates agents on OSWorld, ScreenSpot, and Windows Arena benchmarks using specialized sandboxes and advanced verification methods. Learn how it computes success metrics.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: how-to-guide
- Published: 2026-04-27

---

**Cua-Bench executes agents inside specialized sandboxes—QEMU VMs for OSWorld, static screenshot datasets for ScreenSpot, and Docker-based Windows containers for Windows Arena—then computes success metrics through task verifiers, IoU calculations, and end-to-end completion checks.**

Cua-Bench is the built-in benchmarking suite for the Cua (Computer-Use Agent) framework. It provides a unified evaluation pipeline that provisions the appropriate runtime environment for each benchmark family, feeds screenshots to the agent, and aggregates performance statistics using benchmark-specific scoring logic.

## OSWorld Evaluation via QEMU/KVM Transport

OSWorld tests full-desktop automation across 369+ Linux tasks including browsing, document editing, and coding. Cua-Bench evaluates agents on OSWorld by launching a QEMU/KVM virtual machine that runs the official OSWorld Flask server.

### VM Provisioning and the OSWorld Transport Layer

Cua-Bench downloads the OSWorld Ubuntu `.qcow2` image from HuggingFace and caches it locally. The VM is started with the OSWorld Flask server pre-installed. Agent communication flows through **`OSWorldTransport`** ([`libs/python/cua-sandbox/cua_sandbox/transport/osworld.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-sandbox/cua_sandbox/transport/osworld.py)), which implements two critical endpoints:

- `/screenshot` – Returns a PNG of the current desktop state.
- `/execute_action` – Receives JSON actions (`click`, `type`, `scroll`) and executes them via `pyautogui` inside the VM.

This transport layer abstracts the VM boundary, allowing the generic agent loop to treat the remote desktop as a local environment.

### The Agent Loop and Task Verification

The `ComputerAgent` (or any subclass) is instantiated with the OSWorld transport. The evaluation loop proceeds as follows:

1. The agent receives a screenshot via the transport.
2. It predicts an action and sends it back via `/execute_action`.
3. The loop continues until the HUD task reports success or a step limit is reached.

Each OSWorld task includes a **verifier function** that checks specific success conditions—such as verifying a document was saved, a webpage opened, or a UI element exists. Cua-Bench records whether the verifier passed and aggregates the **success rate**, **step count**, and **wall-clock time**.

### Running OSWorld Benchmarks

Execute OSWorld evaluation using the Cua-Bench CLI:

```bash
cua-bench run tasks/osworld --agent opencua --output osworld-results/

```

This command invokes the `cb run` entry point ([`libs/python/cua-bench/cua_bench/cli/main.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-bench/cua_bench/cli/main.py)), which constructs the QEMU sandbox, loads the OSWorld tasks, and drives the evaluation loop.

## ScreenSpot v2 and Pro Evaluation

ScreenSpot benchmarks test GUI grounding and click-prediction accuracy on static screenshots. Cua-Bench evaluates agents on both **ScreenSpot-v2** and **ScreenSpot-Pro** (high-resolution variant) using dataset-driven scripts.

### Dataset Loading and Click Prediction

The benchmark drivers ([`ss-v2.py`](https://github.com/trycua/cua/blob/main/ss-v2.py) and [`ss-pro.py`](https://github.com/trycua/cua/blob/main/ss-pro.py) in `libs/python/agent/benchmarks/`) load datasets via HuggingFace:

- `lmms-lab/ScreenSpot-v2`
- `lmms-lab/ScreenSpot-Pro`

For each sample, the script feeds the screenshot to the agent's **`predict_click`** method, which returns an `(x, y)` coordinate representing the predicted click location.

### IoU-Based Accuracy Metrics

Cua-Bench converts the predicted coordinate into a 1×1 pixel bounding box and computes the **Intersection-over-Union (IoU)** against the ground-truth bounding box. A prediction is counted as correct if **IoU ≥ 0.5**. The benchmark computes overall accuracy and per-class breakdowns.

Results are written to Markdown reports ([`screenspot_v2_results.md`](https://github.com/trycua/cua/blob/main/screenspot_v2_results.md) or [`screenspot_pro_results.md`](https://github.com/trycua/cua/blob/main/screenspot_pro_results.md)) in the output folder. Shared utilities for IoU calculation and aggregation live in [`libs/python/agent/benchmarks/utils.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/benchmarks/utils.py).

### ScreenSpot Benchmark Execution

Run ScreenSpot evaluations with specific task targets:

```bash

# ScreenSpot v2

cua-bench run tasks/screenspot-v2 --agent opencua --output ss-v2-results/

# ScreenSpot-Pro (high-res)

cua-bench run tasks/screenspot-pro --agent opencua --output ss-pro-results/

```

## Windows Arena Evaluation on Dockerized Windows VMs

Windows Arena tests end-to-end tasks on Windows 10 environments across applications like Chrome, LibreOffice, and VLC. Cua-Bench evaluates agents on Windows Arena using a containerized Windows VM with a custom control server.

### WinArena Container Architecture

The `trycua/winarena:latest` Docker image contains a Windows 10 VM plus the **WinArena server**, which exposes the same `/status` and `/screenshot` API as the Linux computer-server on port 8000. The **`WAASetupController`** ([`libs/cua-bench/tasks/winarena_adapter/setup_controller.py`](https://github.com/trycua/cua/blob/main/libs/cua-bench/tasks/winarena_adapter/setup_controller.py)) handles VM provisioning, including optional Windows ISO downloads and application pre-installation (`--install-apps`).

Agent communication uses the standard **`ComputerTransport`** ([`libs/python/cua-sandbox/cua_sandbox/transport/websocket.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-sandbox/cua_sandbox/transport/websocket.py)) or VNC-SSH transport to interact with the WinArena server.

### Task Evaluators and Success Criteria

Task definitions reside in `libs/cua-bench/tasks/winarena_adapter/evaluators/`. Each evaluator executes a series of actions (clicks, typing, file operations) and verifies the result—such as confirming a PDF was saved or a video played successfully.

Cua-Bench collects success flags and reports **task-success rate**, **average steps**, and **elapsed time** in a formatted table.

### Running Windows Arena Benchmarks

Launch Windows Arena evaluation with setup flags:

```bash
cua-bench run tasks/winarena_adapter \
    --agent opencua \
    --setup --download-iso \
    --output winarena-results/

```

The `--setup` flag triggers the `WAASetupController` to provision the container, while `--download-iso` fetches the Windows installation media if not present.

## Unified CLI and Shared Infrastructure

All three benchmark families share the same **Cua-Bench CLI** architecture. The `cb run` command parses the task name and constructs the appropriate **Sandbox** abstraction ([`libs/python/cua-sandbox/cua_sandbox/sandbox.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-sandbox/cua_sandbox/sandbox.py)), which selects between `OSWorldTransport` (Linux QEMU), standard `ComputerTransport` (Windows container), or direct dataset loading (ScreenSpot).

Platform selection logic is defined in [`libs/python/cua-bench/cua_bench/cli/commands/platform.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-bench/cua_bench/cli/commands/platform.py), which maps CLI flags like `--platform linux-qemu` or `--platform winarena` to the correct transport and sandbox implementations.

## Summary

- **OSWorld**: Cua-Bench provisions a QEMU VM running the OSWorld Flask server, uses `OSWorldTransport` to exchange screenshots and actions, and scores tasks via verifier functions that check desktop state.
- **ScreenSpot**: Static datasets drive click-prediction evaluation; accuracy is measured by IoU ≥ 0.5 between predicted and ground-truth bounding boxes, with results aggregated in Markdown reports.
- **Windows Arena**: A Docker-based Windows VM runs the WinArena server on port 8000; evaluators execute actions and verify success criteria across real Windows applications.
- **Unified Interface**: The `cua-bench run` command handles all three benchmarks through a shared entry point and sandbox abstraction layer.

## Frequently Asked Questions

### How does Cua-Bench handle action execution inside the OSWorld VM?

Cua-Bench sends actions as JSON payloads to the `/execute_action` endpoint of the OSWorld Flask server, which translates them into `pyautogui` commands executed directly on the Ubuntu desktop.

### What metric determines success in ScreenSpot evaluations?

ScreenSpot uses **Intersection-over-Union (IoU)**. The agent's predicted `(x, y)` coordinate is converted to a 1×1 box; if the IoU with the ground-truth bounding box is 0.5 or higher, the prediction is marked correct.

### Can Cua-Bench evaluate agents on Windows Arena without pre-installing applications?

Yes. The `WAASetupController` supports the `--install-apps` flag during setup, which automatically pre-installs benchmark applications like Chrome and LibreOffice into the Windows container before evaluation begins.

### Where does Cua-Bench store benchmark results and logs?

Results are written to the directory specified by the `--output` flag. OSWorld and Windows Arena produce success-rate tables and timing statistics, while ScreenSpot generates Markdown reports ([`screenspot_v2_results.md`](https://github.com/trycua/cua/blob/main/screenspot_v2_results.md) or [`screenspot_pro_results.md`](https://github.com/trycua/cua/blob/main/screenspot_pro_results.md)) containing per-class accuracy breakdowns.