How to Contribute Custom Environment Adapters to the Cua-Bench Registry
To contribute a custom environment adapter to the Cua-Bench registry, implement a Python class inheriting from EnvironmentAdapter with setup, launch, and evaluate methods, define a task.yaml metadata file, generate the registry locally using cb dump-task-registry, and submit a pull request to the separate cua-bench-registry repository.
Cua-Bench is the benchmarking framework developed by trycua that standardizes how computer-use agents interact with diverse sandboxed environments. Contributing custom environment adapters allows you to extend the platform to support proprietary applications, custom operating systems, or specialized GUI workflows. This guide walks through the complete process for contributing custom environment adapters to the Cua-Bench registry, referencing the actual source code in the trycua/cua repository.
Step 1: Create the Adapter Implementation
Every benchmark environment in Cua-Bench is controlled by an environment adapter—a Python module that knows how to provision, control, and evaluate a specific sandbox. Create a new package under libs/cua-bench/tasks/ that implements the required interface.
Your adapter must expose a class inheriting from cua_bench.adapter.EnvironmentAdapter (see the reference implementation in libs/cua-bench/tasks/winarena_adapter/controller_adapter.py). At minimum, implement these three methods:
| Method | Purpose |
|---|---|
setup(self, sandbox: Sandbox) |
Install required software, download containers or ISOs, and configure the VM. |
launch(self, sandbox: Sandbox) |
Start the environment and return the URL or port the agent should connect to. |
evaluate(self, sandbox: Sandbox, run_id: str) |
Execute the benchmark task, collect metrics, and return a result dictionary. |
cleanup(self, sandbox: Sandbox) (optional) |
Tear down resources after the run completes. |
The Sandbox object passed to each method is imported from the core cua package and provides secure shell access and process management within the isolated environment.
Step 2: Add the Task Definition
Alongside your adapter code, create a task.yaml file that registers the dataset, environment name, and static assets with the discovery system. The directory structure must mirror the existing Windows Arena adapter:
libs/
└─ cua-bench/
└─ tasks/
└─ my_custom_env/
├─ __init__.py # Registers the adapter class
├─ controller_adapter.py # Implements the adapter logic
├─ task.yaml # Metadata for the registry
└─ assets/ # Optional UI resources (icons, screenshots)
The __init__.py file should expose your adapter class to Cua-Bench’s import mechanism, similar to libs/cua-bench/tasks/winarena_adapter/__init__.py.
The task.yaml declares the adapter entry point and environment constraints:
name: my-custom-env
description: |
A custom Linux desktop environment that runs MyApp.
environment:
os: linux
type: gui
adapter: my_custom_env.controller_adapter.MyCustomEnvAdapter
Step 3: Generate the Task Registry
Cua-Bench uses a centralized task_registry.json file that the benchmark UI consumes. After adding your adapter, generate this registry locally to verify discovery.
Navigate to the bench library and install the package in editable mode:
cd libs/cua-bench
uv tool install -e .
Then execute the dump script:
cb dump-task-registry
This runs libs/cua-bench/cua_bench/scripts/dump_task_registry.py, which walks every task.yaml under libs/cua-bench/tasks/ and writes the consolidated JSON. Verify that your new environment appears in the generated task_registry.json before proceeding.
Step 4: Write Unit Tests
Add comprehensive tests under libs/cua-bench/tasks/my_custom_env/tests/ that spin up an ephemeral sandbox and exercise each adapter method. The repository contains examples in libs/cua-bench/tasks/winarena_adapter/tests/ demonstrating how to mock sandbox interactions and assert on adapter outputs.
Run pytest from the repository root to ensure your tests pass locally and will succeed in CI:
pytest libs/cua-bench/tasks/my_custom_env/tests/
Step 5: Submit to the Registry Repository
The public registry lives in a separate repository at https://github.com/trycua/cua-bench-registry. To publish your adapter:
- Fork the registry repository.
- Add your task folder under the
datasets/directory (allowing the registry CI to recompute the JSON) or copy the verifiedtask_registry.jsonfrom your local run. - Include a
README.mddescribing usage, licensing, and any special requirements (e.g., API keys or proprietary binaries). - Open a pull request against the
mainbranch.
The registry CI runs dump_task_registry.py again, validates the JSON schema against your task.yaml, and executes your tests. Once merged, the adapter becomes visible on the public Cua-Bench UI at https://cuabench.ai/registry and can be referenced in benchmark commands:
cb run dataset my_custom_env --agent cua-agent
Implementation Example
Below is a complete adapter implementation for a bespoke Linux desktop application:
# libs/cua-bench/tasks/my_custom_env/controller_adapter.py
from cua_bench.adapter import EnvironmentAdapter
from cua import Sandbox
class MyCustomEnvAdapter(EnvironmentAdapter):
"""Adapter for a bespoke Linux desktop application."""
async def setup(self, sandbox: Sandbox) -> None:
# Install dependencies inside the sandbox
await sandbox.shell.run("apt-get update && apt-get install -y my-app")
async def launch(self, sandbox: Sandbox) -> str:
# Start the GUI and return the URL the agent should connect to
await sandbox.shell.run("my-app --headless &")
return f"http://{sandbox.host}:{sandbox.port}"
async def evaluate(self, sandbox: Sandbox, run_id: str) -> dict:
# Perform the benchmark task and return metrics
result = await sandbox.shell.run("my-app --benchmark")
return {"run_id": run_id, "score": float(result.stdout.strip())}
Key Files Reference
| File | Role |
|---|---|
libs/cua-bench/tasks/winarena_adapter/controller_adapter.py |
Reference implementation of a full-featured environment adapter. |
libs/cua-bench/tasks/winarena_adapter/__init__.py |
Registers the adapter class with the Cua-Bench discovery mechanism. |
libs/cua-bench/cua_bench/scripts/dump_task_registry.py |
Generates task_registry.json from all task.yaml files. |
docs/content/docs/cuabench/guide/fundamentals/adapters.mdx |
User-facing documentation explaining the adapter contract. |
Summary
- Environment adapters inherit from
cua_bench.adapter.EnvironmentAdapterand implementsetup,launch, andevaluatemethods. - Place adapter code under
libs/cua-bench/tasks/with atask.yamlfile defining metadata and the adapter class path. - Run
cb dump-task-registrylocally to validate that Cua-Bench discovers your environment and generates valid JSON. - Submit contributions to the separate
cua-bench-registryrepository, where CI validates the schema and runs your tests. - Once merged, environments are publicly visible at
cuabench.ai/registryand accessible via thecb runCLI.
Frequently Asked Questions
What is the difference between the main Cua repository and the Cua-Bench Registry?
The trycua/cua repository contains the core benchmarking framework and adapter interface, while the cua-bench-registry repository hosts the canonical task definitions and generated task_registry.json that powers the public UI. You develop and test adapters locally in the main repo, then publish them by submitting a pull request to the registry repository.
Do I need to implement all four adapter methods?
You must implement setup, launch, and evaluate. The cleanup method is optional but recommended for environments that require explicit resource teardown, such as stopping services or deleting temporary files after a benchmark run.
How does Cua-Bench discover my custom adapter?
The dump_task_registry.py script recursively searches libs/cua-bench/tasks/ for task.yaml files. Each YAML file’s adapter field points to a Python class path that the system imports dynamically. Ensure your __init__.py correctly exports the adapter class so the registry generation succeeds.
Can I contribute an adapter for a proprietary or closed-source environment?
Yes. The adapter code and metadata in the registry can reference external resources or require authentication tokens. As long as you provide the controller logic that knows how to provision and evaluate the sandbox, Cua-Bench can integrate it. Document any licensing restrictions or access requirements in your README.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →