How to Add a Custom Benchmark Task to Cua-Bench: Complete Implementation Guide

To add a custom benchmark task to Cua-Bench, create a directory containing a main.py file that exposes four optional functions decorated with @cb.tasks_config, @cb.setup_task, @cb.evaluate_task, and optionally @cb.solve_task, then validate your implementation using cb task info and execute it with cb run.

Cua-Bench is the open-source benchmark suite for computer-use agents maintained by the trycua organization. Adding your own custom benchmark task enables you to evaluate agent performance on domain-specific workflows ranging from simple file operations to complex multi-step GUI interactions.

The Four Core Decorators

Cua-Bench discovers benchmark environments by importing a directory containing a main.py file. According to the source code in [cua_bench/decorators.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/decorators.py), you must implement four optional functions, each marked with a specific decorator that adds hidden attributes (_td_type, _td_split) for the core loader:

  • @cb.tasks_config(split="train"): Returns a list of cb.Task objects describing every task variant. The function signature must be def load() -> list[cb.Task].

  • @cb.setup_task(split="train"): Runs once when the environment resets to create windows, files, or configure the system. The signature is async def start(task_cfg: cb.Task, session: cb.DesktopSession).

  • @cb.evaluate_task(split="train"): Called after the agent finishes to compute rewards. Must return async def evaluate(task_cfg: cb.Task, session: cb.DesktopSession) -> list[float].

  • @cb.solve_task(split="train") (optional): Provides a reference solution runnable via cb interact … --oracle. Signature matches setup: async def solve(task_cfg: cb.Task, session: cb.DesktopSession).

When a task directory is passed to the CLI, Cua-Bench executes Environment.make_from_module in [cua_bench/environment.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/environment.py) (lines 99-130), which scans every module attribute and wires any callable with _td_type set into the Environment instance.

Step 1: Scaffold a New Task Environment

Use the CLI to generate boilerplate files. The command is implemented in [cli/commands/task.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py#L15-L78):

cb task create my_new_task_env

This creates:

  • pyproject.toml — Python package configuration
  • main.py — Stub implementations of the four decorator functions
  • gui/ — Directory containing index.html for simulated browser tasks

Step 2: Define Task Variants with @cb.tasks_config

Edit main.py to declare your task configurations. The @cb.tasks_config decorator marks the function that returns task variants:

import cua_bench as cb

@cb.tasks_config(split="train")
def load():
    """Return list of task configurations."""
    return [
        cb.Task(
            description="Count the number of words in the 'short.txt' file on the Desktop.",
            metadata={"filename": "short.txt", "expected_words": 5},
            computer={
                "provider": "native",  # or "simulated" for headless browser

                "setup_config": {
                    "os_type": "linux",
                    "width": 1024,
                    "height": 768,
                },
            },
        ),
        cb.Task(
            description="Count the number of words in the 'long.txt' file on the Desktop.",
            metadata={"filename": "long.txt", "expected_words": 23},
            computer={"provider": "native", "setup_config": {"os_type": "linux"}},
        )
    ]

The cb.Task constructor requires description, optional metadata for passing parameters between functions, and a computer dictionary specifying the execution environment.

Step 3: Configure the Environment with @cb.setup_task

Implement @cb.setup_task to prepare the system state before the agent acts:

@cb.setup_task(split="train")
async def start(task_cfg: cb.Task, session: cb.DesktopSession):
    """Create the target file when the task starts."""
    filename = f"~/Desktop/{task_cfg.metadata['filename']}"
    
    content_map = {
        "short.txt": "Cua-Bench is a benchmark suite.",
        "long.txt": "Cua-Bench provides a collection of computer-use reinforcement-learning tasks."
    }
    content = content_map[task_cfg.metadata["filename"]]
    
    await session.run_command(f'echo "{content}" > {filename}')

The session parameter provides a DesktopSession object for executing commands in the VM or container.

Step 4: Implement Scoring with @cb.evaluate_task

Add @cb.evaluate_task to calculate rewards after the agent completes its actions:

@cb.evaluate_task(split="train")
async def evaluate(task_cfg: cb.Task, session: cb.DesktopSession) -> list[float]:
    """Check if the word count is correct."""
    filename = f"~/Desktop/{task_cfg.metadata['filename']}"
    result = await session.run_command(f"wc -w {filename} | awk '{{print $1}}'")
    
    actual_count = int(result["stdout"].strip())
    expected = task_cfg.metadata["expected_words"]
    
    return [1.0] if actual_count == expected else [0.0]

The function must return a list of floats representing rewards for the task variant.

Step 5: Add a Reference Solution (Optional)

Provide an oracle implementation using @cb.solve_task for debugging and baseline measurements:

@cb.solve_task(split="train")
async def solve(task_cfg: cb.Task, session: cb.DesktopSession):
    """Reference solution that outputs the expected answer."""
    expected = task_cfg.metadata["expected_words"]
    await session.run_command(f'echo "Answer: {expected}"')

Run this with cb interact my_new_task_env --oracle to verify your task logic works correctly before testing against agents.

Validation and Execution

Validate your task configuration using the CLI (implemented in [cli/commands/task.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py#L21-L124)):


# Display task summary and variant count

cb task info my_new_task_env

# Run the reference solution

cb interact my_new_task_env --oracle

# Execute in benchmark mode

cb run my_new_task_env --variant-id 0

The runner in [cli/commands/run.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/run.py) creates an Environment via make_from_module, calls reset(), then runs step() until completion or max_steps exhaustion.

Key Implementation Files

File Purpose
[cua_bench/decorators.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/decorators.py) Defines @cb.tasks_config, @cb.setup_task, @cb.evaluate_task, @cb.solve_task.
[cua_bench/environment.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/environment.py) Contains Environment.make_from_module (lines 99-130) for loading task modules.
[cli/commands/task.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py) Implements cb task create and cb task info commands.
templates/starter_env/ Boilerplate used by the scaffold command.

Summary

  • Cua-Bench discovers tasks by importing main.py files and scanning for decorated functions using _td_type attributes.
  • Four decorators control the lifecycle: @cb.tasks_config (define variants), @cb.setup_task (prepare environment), @cb.evaluate_task (score results), and @cb.solve_task (reference solution).
  • Scaffold quickly using cb task create <directory> to generate pyproject.toml, main.py, and gui/ files.
  • Return rewards as a list of floats from the evaluate function to integrate with the benchmark runner.
  • Validate with cb task info and cb interact --oracle before full benchmark runs.

Frequently Asked Questions

How do I switch between native VMs and simulated browsers?

In the cb.Task computer dictionary, set "provider": "native" to run inside a Docker/QEMU VM, or "provider": "simulated" to render the gui/index.html in a headless Playwright browser. The native provider displays the actual file system inside the VM, while simulated runs HTML-based tasks without Docker.

Can I pass custom parameters between setup and evaluation?

Yes. Store arbitrary data in the metadata dictionary when constructing cb.Task in your @cb.tasks_config function. This dictionary is passed to both @cb.setup_task and @cb.evaluate_task via the task_cfg.metadata attribute, allowing you to communicate file paths, expected values, or configuration flags.

Where should I place custom tasks to include them in the benchmark suite?

Place your task directory anywhere under libs/cua-bench/tasks/ or pass a custom folder path to the benchmark runner. The runner walks the directory tree, automatically loading any environment containing a main.py file and aggregating scores across all discovered tasks.

How do I debug a failing evaluation?

Run the reference solution using cb interact <task_dir> --oracle to verify that the setup creates the correct state and that the evaluation logic produces expected rewards. Check the Environment.make_from_module implementation in environment.py (lines 99-130) to ensure your decorated functions are being detected correctly if the task fails to load.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →