How to Add a Custom Benchmark Task to Cua-Bench: Complete Implementation Guide
To add a custom benchmark task to Cua-Bench, create a directory containing a main.py file that exposes four optional functions decorated with @cb.tasks_config, @cb.setup_task, @cb.evaluate_task, and optionally @cb.solve_task, then validate your implementation using cb task info and execute it with cb run.
Cua-Bench is the open-source benchmark suite for computer-use agents maintained by the trycua organization. Adding your own custom benchmark task enables you to evaluate agent performance on domain-specific workflows ranging from simple file operations to complex multi-step GUI interactions.
The Four Core Decorators
Cua-Bench discovers benchmark environments by importing a directory containing a main.py file. According to the source code in [cua_bench/decorators.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/decorators.py), you must implement four optional functions, each marked with a specific decorator that adds hidden attributes (_td_type, _td_split) for the core loader:
-
@cb.tasks_config(split="train"): Returns a list ofcb.Taskobjects describing every task variant. The function signature must bedef load() -> list[cb.Task]. -
@cb.setup_task(split="train"): Runs once when the environment resets to create windows, files, or configure the system. The signature isasync def start(task_cfg: cb.Task, session: cb.DesktopSession). -
@cb.evaluate_task(split="train"): Called after the agent finishes to compute rewards. Must returnasync def evaluate(task_cfg: cb.Task, session: cb.DesktopSession) -> list[float]. -
@cb.solve_task(split="train")(optional): Provides a reference solution runnable viacb interact … --oracle. Signature matches setup:async def solve(task_cfg: cb.Task, session: cb.DesktopSession).
When a task directory is passed to the CLI, Cua-Bench executes Environment.make_from_module in [cua_bench/environment.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/environment.py) (lines 99-130), which scans every module attribute and wires any callable with _td_type set into the Environment instance.
Step 1: Scaffold a New Task Environment
Use the CLI to generate boilerplate files. The command is implemented in [cli/commands/task.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py#L15-L78):
cb task create my_new_task_env
This creates:
pyproject.toml— Python package configurationmain.py— Stub implementations of the four decorator functionsgui/— Directory containingindex.htmlfor simulated browser tasks
Step 2: Define Task Variants with @cb.tasks_config
Edit main.py to declare your task configurations. The @cb.tasks_config decorator marks the function that returns task variants:
import cua_bench as cb
@cb.tasks_config(split="train")
def load():
"""Return list of task configurations."""
return [
cb.Task(
description="Count the number of words in the 'short.txt' file on the Desktop.",
metadata={"filename": "short.txt", "expected_words": 5},
computer={
"provider": "native", # or "simulated" for headless browser
"setup_config": {
"os_type": "linux",
"width": 1024,
"height": 768,
},
},
),
cb.Task(
description="Count the number of words in the 'long.txt' file on the Desktop.",
metadata={"filename": "long.txt", "expected_words": 23},
computer={"provider": "native", "setup_config": {"os_type": "linux"}},
)
]
The cb.Task constructor requires description, optional metadata for passing parameters between functions, and a computer dictionary specifying the execution environment.
Step 3: Configure the Environment with @cb.setup_task
Implement @cb.setup_task to prepare the system state before the agent acts:
@cb.setup_task(split="train")
async def start(task_cfg: cb.Task, session: cb.DesktopSession):
"""Create the target file when the task starts."""
filename = f"~/Desktop/{task_cfg.metadata['filename']}"
content_map = {
"short.txt": "Cua-Bench is a benchmark suite.",
"long.txt": "Cua-Bench provides a collection of computer-use reinforcement-learning tasks."
}
content = content_map[task_cfg.metadata["filename"]]
await session.run_command(f'echo "{content}" > {filename}')
The session parameter provides a DesktopSession object for executing commands in the VM or container.
Step 4: Implement Scoring with @cb.evaluate_task
Add @cb.evaluate_task to calculate rewards after the agent completes its actions:
@cb.evaluate_task(split="train")
async def evaluate(task_cfg: cb.Task, session: cb.DesktopSession) -> list[float]:
"""Check if the word count is correct."""
filename = f"~/Desktop/{task_cfg.metadata['filename']}"
result = await session.run_command(f"wc -w {filename} | awk '{{print $1}}'")
actual_count = int(result["stdout"].strip())
expected = task_cfg.metadata["expected_words"]
return [1.0] if actual_count == expected else [0.0]
The function must return a list of floats representing rewards for the task variant.
Step 5: Add a Reference Solution (Optional)
Provide an oracle implementation using @cb.solve_task for debugging and baseline measurements:
@cb.solve_task(split="train")
async def solve(task_cfg: cb.Task, session: cb.DesktopSession):
"""Reference solution that outputs the expected answer."""
expected = task_cfg.metadata["expected_words"]
await session.run_command(f'echo "Answer: {expected}"')
Run this with cb interact my_new_task_env --oracle to verify your task logic works correctly before testing against agents.
Validation and Execution
Validate your task configuration using the CLI (implemented in [cli/commands/task.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py#L21-L124)):
# Display task summary and variant count
cb task info my_new_task_env
# Run the reference solution
cb interact my_new_task_env --oracle
# Execute in benchmark mode
cb run my_new_task_env --variant-id 0
The runner in [cli/commands/run.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/run.py) creates an Environment via make_from_module, calls reset(), then runs step() until completion or max_steps exhaustion.
Key Implementation Files
| File | Purpose |
|---|---|
[cua_bench/decorators.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/decorators.py) |
Defines @cb.tasks_config, @cb.setup_task, @cb.evaluate_task, @cb.solve_task. |
[cua_bench/environment.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/environment.py) |
Contains Environment.make_from_module (lines 99-130) for loading task modules. |
[cli/commands/task.py](https://github.com/trycua/cua/blob/main/libs/cua-bench/cua_bench/cli/commands/task.py) |
Implements cb task create and cb task info commands. |
templates/starter_env/ |
Boilerplate used by the scaffold command. |
Summary
- Cua-Bench discovers tasks by importing
main.pyfiles and scanning for decorated functions using_td_typeattributes. - Four decorators control the lifecycle:
@cb.tasks_config(define variants),@cb.setup_task(prepare environment),@cb.evaluate_task(score results), and@cb.solve_task(reference solution). - Scaffold quickly using
cb task create <directory>to generatepyproject.toml,main.py, andgui/files. - Return rewards as a list of floats from the evaluate function to integrate with the benchmark runner.
- Validate with
cb task infoandcb interact --oraclebefore full benchmark runs.
Frequently Asked Questions
How do I switch between native VMs and simulated browsers?
In the cb.Task computer dictionary, set "provider": "native" to run inside a Docker/QEMU VM, or "provider": "simulated" to render the gui/index.html in a headless Playwright browser. The native provider displays the actual file system inside the VM, while simulated runs HTML-based tasks without Docker.
Can I pass custom parameters between setup and evaluation?
Yes. Store arbitrary data in the metadata dictionary when constructing cb.Task in your @cb.tasks_config function. This dictionary is passed to both @cb.setup_task and @cb.evaluate_task via the task_cfg.metadata attribute, allowing you to communicate file paths, expected values, or configuration flags.
Where should I place custom tasks to include them in the benchmark suite?
Place your task directory anywhere under libs/cua-bench/tasks/ or pass a custom folder path to the benchmark runner. The runner walks the directory tree, automatically loading any environment containing a main.py file and aggregating scores across all discovered tasks.
How do I debug a failing evaluation?
Run the reference solution using cb interact <task_dir> --oracle to verify that the setup creates the correct state and that the evaluation logic produces expected rewards. Check the Environment.make_from_module implementation in environment.py (lines 99-130) to ensure your decorated functions are being detected correctly if the task fails to load.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →