How to Contribute Custom Environment Adapters to the Cua-Bench Registry

To contribute a custom environment adapter to the Cua-Bench registry, implement a Python class inheriting from EnvironmentAdapter with setup, launch, and evaluate methods, define a task.yaml metadata file, generate the registry locally using cb dump-task-registry, and submit a pull request to the separate cua-bench-registry repository.

Cua-Bench is the benchmarking framework developed by trycua that standardizes how computer-use agents interact with diverse sandboxed environments. Contributing custom environment adapters allows you to extend the platform to support proprietary applications, custom operating systems, or specialized GUI workflows. This guide walks through the complete process for contributing custom environment adapters to the Cua-Bench registry, referencing the actual source code in the trycua/cua repository.

Step 1: Create the Adapter Implementation

Every benchmark environment in Cua-Bench is controlled by an environment adapter—a Python module that knows how to provision, control, and evaluate a specific sandbox. Create a new package under libs/cua-bench/tasks/ that implements the required interface.

Your adapter must expose a class inheriting from cua_bench.adapter.EnvironmentAdapter (see the reference implementation in libs/cua-bench/tasks/winarena_adapter/controller_adapter.py). At minimum, implement these three methods:

Method Purpose
setup(self, sandbox: Sandbox) Install required software, download containers or ISOs, and configure the VM.
launch(self, sandbox: Sandbox) Start the environment and return the URL or port the agent should connect to.
evaluate(self, sandbox: Sandbox, run_id: str) Execute the benchmark task, collect metrics, and return a result dictionary.
cleanup(self, sandbox: Sandbox) (optional) Tear down resources after the run completes.

The Sandbox object passed to each method is imported from the core cua package and provides secure shell access and process management within the isolated environment.

Step 2: Add the Task Definition

Alongside your adapter code, create a task.yaml file that registers the dataset, environment name, and static assets with the discovery system. The directory structure must mirror the existing Windows Arena adapter:

libs/
└─ cua-bench/
   └─ tasks/
      └─ my_custom_env/
         ├─ __init__.py          # Registers the adapter class

         ├─ controller_adapter.py # Implements the adapter logic

         ├─ task.yaml            # Metadata for the registry

         └─ assets/              # Optional UI resources (icons, screenshots)

The __init__.py file should expose your adapter class to Cua-Bench’s import mechanism, similar to libs/cua-bench/tasks/winarena_adapter/__init__.py.

The task.yaml declares the adapter entry point and environment constraints:

name: my-custom-env
description: |
  A custom Linux desktop environment that runs MyApp.
environment:
  os: linux
  type: gui
adapter: my_custom_env.controller_adapter.MyCustomEnvAdapter

Step 3: Generate the Task Registry

Cua-Bench uses a centralized task_registry.json file that the benchmark UI consumes. After adding your adapter, generate this registry locally to verify discovery.

Navigate to the bench library and install the package in editable mode:

cd libs/cua-bench
uv tool install -e .

Then execute the dump script:

cb dump-task-registry

This runs libs/cua-bench/cua_bench/scripts/dump_task_registry.py, which walks every task.yaml under libs/cua-bench/tasks/ and writes the consolidated JSON. Verify that your new environment appears in the generated task_registry.json before proceeding.

Step 4: Write Unit Tests

Add comprehensive tests under libs/cua-bench/tasks/my_custom_env/tests/ that spin up an ephemeral sandbox and exercise each adapter method. The repository contains examples in libs/cua-bench/tasks/winarena_adapter/tests/ demonstrating how to mock sandbox interactions and assert on adapter outputs.

Run pytest from the repository root to ensure your tests pass locally and will succeed in CI:

pytest libs/cua-bench/tasks/my_custom_env/tests/

Step 5: Submit to the Registry Repository

The public registry lives in a separate repository at https://github.com/trycua/cua-bench-registry. To publish your adapter:

  1. Fork the registry repository.
  2. Add your task folder under the datasets/ directory (allowing the registry CI to recompute the JSON) or copy the verified task_registry.json from your local run.
  3. Include a README.md describing usage, licensing, and any special requirements (e.g., API keys or proprietary binaries).
  4. Open a pull request against the main branch.

The registry CI runs dump_task_registry.py again, validates the JSON schema against your task.yaml, and executes your tests. Once merged, the adapter becomes visible on the public Cua-Bench UI at https://cuabench.ai/registry and can be referenced in benchmark commands:

cb run dataset my_custom_env --agent cua-agent

Implementation Example

Below is a complete adapter implementation for a bespoke Linux desktop application:


# libs/cua-bench/tasks/my_custom_env/controller_adapter.py

from cua_bench.adapter import EnvironmentAdapter
from cua import Sandbox

class MyCustomEnvAdapter(EnvironmentAdapter):
    """Adapter for a bespoke Linux desktop application."""

    async def setup(self, sandbox: Sandbox) -> None:
        # Install dependencies inside the sandbox

        await sandbox.shell.run("apt-get update && apt-get install -y my-app")

    async def launch(self, sandbox: Sandbox) -> str:
        # Start the GUI and return the URL the agent should connect to

        await sandbox.shell.run("my-app --headless &")
        return f"http://{sandbox.host}:{sandbox.port}"

    async def evaluate(self, sandbox: Sandbox, run_id: str) -> dict:
        # Perform the benchmark task and return metrics

        result = await sandbox.shell.run("my-app --benchmark")
        return {"run_id": run_id, "score": float(result.stdout.strip())}

Key Files Reference

File Role
libs/cua-bench/tasks/winarena_adapter/controller_adapter.py Reference implementation of a full-featured environment adapter.
libs/cua-bench/tasks/winarena_adapter/__init__.py Registers the adapter class with the Cua-Bench discovery mechanism.
libs/cua-bench/cua_bench/scripts/dump_task_registry.py Generates task_registry.json from all task.yaml files.
docs/content/docs/cuabench/guide/fundamentals/adapters.mdx User-facing documentation explaining the adapter contract.

Summary

  • Environment adapters inherit from cua_bench.adapter.EnvironmentAdapter and implement setup, launch, and evaluate methods.
  • Place adapter code under libs/cua-bench/tasks/ with a task.yaml file defining metadata and the adapter class path.
  • Run cb dump-task-registry locally to validate that Cua-Bench discovers your environment and generates valid JSON.
  • Submit contributions to the separate cua-bench-registry repository, where CI validates the schema and runs your tests.
  • Once merged, environments are publicly visible at cuabench.ai/registry and accessible via the cb run CLI.

Frequently Asked Questions

What is the difference between the main Cua repository and the Cua-Bench Registry?

The trycua/cua repository contains the core benchmarking framework and adapter interface, while the cua-bench-registry repository hosts the canonical task definitions and generated task_registry.json that powers the public UI. You develop and test adapters locally in the main repo, then publish them by submitting a pull request to the registry repository.

Do I need to implement all four adapter methods?

You must implement setup, launch, and evaluate. The cleanup method is optional but recommended for environments that require explicit resource teardown, such as stopping services or deleting temporary files after a benchmark run.

How does Cua-Bench discover my custom adapter?

The dump_task_registry.py script recursively searches libs/cua-bench/tasks/ for task.yaml files. Each YAML file’s adapter field points to a Python class path that the system imports dynamically. Ensure your __init__.py correctly exports the adapter class so the registry generation succeeds.

Can I contribute an adapter for a proprietary or closed-source environment?

Yes. The adapter code and metadata in the registry can reference external resources or require authentication tokens. As long as you provide the controller logic that knows how to provision and evaluate the sandbox, Cua-Bench can integrate it. Document any licensing restrictions or access requirements in your README.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →