# How to Add a Custom Benchmark Task to Cua-Bench: Complete Implementation Guide

> Learn how to add a custom benchmark task to Cua-Bench. Follow this implementation guide to create and validate your own tasks for the Cua-Bench framework.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: how-to-guide
- Published: 2026-04-27

---

**To add a custom benchmark task to Cua-Bench, create a directory containing a [`main.py`](https://github.com/trycua/cua/blob/main/main.py) file that exposes four optional functions decorated with `@cb.tasks_config`, `@cb.setup_task`, `@cb.evaluate_task`, and optionally `@cb.solve_task`, then validate your implementation using `cb task info` and execute it with `cb run`.**

Cua-Bench is the open-source benchmark suite for computer-use agents maintained by the trycua organization. Adding your own custom benchmark task enables you to evaluate agent performance on domain-specific workflows ranging from simple file operations to complex multi-step GUI interactions.

## The Four Core Decorators

Cua-Bench discovers benchmark environments by importing a directory containing a **[`main.py`](https://github.com/trycua/cua/blob/main/main.py)** file. According to the source code in [[`cua_bench/decorators.py`](https://github.com/trycua/cua/blob/main/cua_bench/decorators.py)](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/decorators.py), you must implement four optional functions, each marked with a specific decorator that adds hidden attributes (`_td_type`, `_td_split`) for the core loader:

- **`@cb.tasks_config(split="train")`**: Returns a list of `cb.Task` objects describing every task variant. The function signature must be `def load() -> list[cb.Task]`.

- **`@cb.setup_task(split="train")`**: Runs **once** when the environment resets to create windows, files, or configure the system. The signature is `async def start(task_cfg: cb.Task, session: cb.DesktopSession)`.

- **`@cb.evaluate_task(split="train")`**: Called after the agent finishes to compute rewards. Must return `async def evaluate(task_cfg: cb.Task, session: cb.DesktopSession) -> list[float]`.

- **`@cb.solve_task(split="train")`** *(optional)*: Provides a reference solution runnable via `cb interact … --oracle`. Signature matches setup: `async def solve(task_cfg: cb.Task, session: cb.DesktopSession)`.

When a task directory is passed to the CLI, Cua-Bench executes `Environment.make_from_module` in [[`cua_bench/environment.py`](https://github.com/trycua/cua/blob/main/cua_bench/environment.py)](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/environment.py) (lines 99-130), which scans every module attribute and wires any callable with `_td_type` set into the `Environment` instance.

## Step 1: Scaffold a New Task Environment

Use the CLI to generate boilerplate files. The command is implemented in [[`cli/commands/task.py`](https://github.com/trycua/cua/blob/main/cli/commands/task.py)](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py#L15-L78):

```bash
cb task create my_new_task_env

```

This creates:

- [`pyproject.toml`](https://github.com/trycua/cua/blob/main/pyproject.toml) — Python package configuration
- [`main.py`](https://github.com/trycua/cua/blob/main/main.py) — Stub implementations of the four decorator functions
- `gui/` — Directory containing [`index.html`](https://github.com/trycua/cua/blob/main/index.html) for simulated browser tasks

## Step 2: Define Task Variants with `@cb.tasks_config`

Edit [`main.py`](https://github.com/trycua/cua/blob/main/main.py) to declare your task configurations. The **`@cb.tasks_config`** decorator marks the function that returns task variants:

```python
import cua_bench as cb

@cb.tasks_config(split="train")
def load():
    """Return list of task configurations."""
    return [
        cb.Task(
            description="Count the number of words in the 'short.txt' file on the Desktop.",
            metadata={"filename": "short.txt", "expected_words": 5},
            computer={
                "provider": "native",  # or "simulated" for headless browser

                "setup_config": {
                    "os_type": "linux",
                    "width": 1024,
                    "height": 768,
                },
            },
        ),
        cb.Task(
            description="Count the number of words in the 'long.txt' file on the Desktop.",
            metadata={"filename": "long.txt", "expected_words": 23},
            computer={"provider": "native", "setup_config": {"os_type": "linux"}},
        )
    ]

```

The **`cb.Task`** constructor requires `description`, optional `metadata` for passing parameters between functions, and a `computer` dictionary specifying the execution environment.

## Step 3: Configure the Environment with `@cb.setup_task`

Implement **`@cb.setup_task`** to prepare the system state before the agent acts:

```python
@cb.setup_task(split="train")
async def start(task_cfg: cb.Task, session: cb.DesktopSession):
    """Create the target file when the task starts."""
    filename = f"~/Desktop/{task_cfg.metadata['filename']}"
    
    content_map = {
        "short.txt": "Cua-Bench is a benchmark suite.",
        "long.txt": "Cua-Bench provides a collection of computer-use reinforcement-learning tasks."
    }
    content = content_map[task_cfg.metadata["filename"]]
    
    await session.run_command(f'echo "{content}" > {filename}')

```

The **`session`** parameter provides a `DesktopSession` object for executing commands in the VM or container.

## Step 4: Implement Scoring with `@cb.evaluate_task`

Add **`@cb.evaluate_task`** to calculate rewards after the agent completes its actions:

```python
@cb.evaluate_task(split="train")
async def evaluate(task_cfg: cb.Task, session: cb.DesktopSession) -> list[float]:
    """Check if the word count is correct."""
    filename = f"~/Desktop/{task_cfg.metadata['filename']}"
    result = await session.run_command(f"wc -w {filename} | awk '{{print $1}}'")
    
    actual_count = int(result["stdout"].strip())
    expected = task_cfg.metadata["expected_words"]
    
    return [1.0] if actual_count == expected else [0.0]

```

The function must return a **list of floats** representing rewards for the task variant.

## Step 5: Add a Reference Solution (Optional)

Provide an oracle implementation using **`@cb.solve_task`** for debugging and baseline measurements:

```python
@cb.solve_task(split="train")
async def solve(task_cfg: cb.Task, session: cb.DesktopSession):
    """Reference solution that outputs the expected answer."""
    expected = task_cfg.metadata["expected_words"]
    await session.run_command(f'echo "Answer: {expected}"')

```

Run this with `cb interact my_new_task_env --oracle` to verify your task logic works correctly before testing against agents.

## Validation and Execution

Validate your task configuration using the CLI (implemented in [[`cli/commands/task.py`](https://github.com/trycua/cua/blob/main/cli/commands/task.py)](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py#L21-L124)):

```bash

# Display task summary and variant count

cb task info my_new_task_env

# Run the reference solution

cb interact my_new_task_env --oracle

# Execute in benchmark mode

cb run my_new_task_env --variant-id 0

```

The runner in [[`cli/commands/run.py`](https://github.com/trycua/cua/blob/main/cli/commands/run.py)](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/run.py) creates an `Environment` via `make_from_module`, calls `reset()`, then runs `step()` until completion or `max_steps` exhaustion.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [[`cua_bench/decorators.py`](https://github.com/trycua/cua/blob/main/cua_bench/decorators.py)](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/decorators.py) | Defines `@cb.tasks_config`, `@cb.setup_task`, `@cb.evaluate_task`, `@cb.solve_task`. |
| [[`cua_bench/environment.py`](https://github.com/trycua/cua/blob/main/cua_bench/environment.py)](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/environment.py) | Contains `Environment.make_from_module` (lines 99-130) for loading task modules. |
| [[`cli/commands/task.py`](https://github.com/trycua/cua/blob/main/cli/commands/task.py)](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py) | Implements `cb task create` and `cb task info` commands. |
| [`templates/starter_env/`](https://github.com/trycua/cua/tree/main/libs/cua-bench/templates/starter_env) | Boilerplate used by the scaffold command. |

## Summary

- **Cua-Bench** discovers tasks by importing [`main.py`](https://github.com/trycua/cua/blob/main/main.py) files and scanning for decorated functions using `_td_type` attributes.
- **Four decorators** control the lifecycle: `@cb.tasks_config` (define variants), `@cb.setup_task` (prepare environment), `@cb.evaluate_task` (score results), and `@cb.solve_task` (reference solution).
- **Scaffold quickly** using `cb task create <directory>` to generate [`pyproject.toml`](https://github.com/trycua/cua/blob/main/pyproject.toml), [`main.py`](https://github.com/trycua/cua/blob/main/main.py), and `gui/` files.
- **Return rewards** as a list of floats from the evaluate function to integrate with the benchmark runner.
- **Validate** with `cb task info` and `cb interact --oracle` before full benchmark runs.

## Frequently Asked Questions

### How do I switch between native VMs and simulated browsers?

In the `cb.Task` `computer` dictionary, set `"provider": "native"` to run inside a Docker/QEMU VM, or `"provider": "simulated"` to render the [`gui/index.html`](https://github.com/trycua/cua/blob/main/gui/index.html) in a headless Playwright browser. The native provider displays the actual file system inside the VM, while simulated runs HTML-based tasks without Docker.

### Can I pass custom parameters between setup and evaluation?

Yes. Store arbitrary data in the `metadata` dictionary when constructing `cb.Task` in your `@cb.tasks_config` function. This dictionary is passed to both `@cb.setup_task` and `@cb.evaluate_task` via the `task_cfg.metadata` attribute, allowing you to communicate file paths, expected values, or configuration flags.

### Where should I place custom tasks to include them in the benchmark suite?

Place your task directory anywhere under `libs/cua-bench/tasks/` or pass a custom folder path to the benchmark runner. The runner walks the directory tree, automatically loading any environment containing a [`main.py`](https://github.com/trycua/cua/blob/main/main.py) file and aggregating scores across all discovered tasks.

### How do I debug a failing evaluation?

Run the reference solution using `cb interact <task_dir> --oracle` to verify that the setup creates the correct state and that the evaluation logic produces expected rewards. Check the `Environment.make_from_module` implementation in [`environment.py`](https://github.com/trycua/cua/blob/main/environment.py) (lines 99-130) to ensure your decorated functions are being detected correctly if the task fails to load.