How to Implement Custom Action Spaces for Specific Automation Tasks in Cua

You implement custom action spaces in Cua by defining a plain-text function signature block, injecting it into the agent's prompt template, parsing the LLM's structured response, and mapping parsed actions to sandbox commands using response-item factories.

Cua is an open-source computer-use automation framework that drives GUI interactions through language models. Unlike rigid automation tools, Cua's action space—the set of functions available to the model—is not hard-coded in the core engine but exists as a configurable text block within agent loops. This architecture allows you to tailor available actions for specific automation scenarios without modifying the underlying inference logic.

How the Action Space Mechanism Works

Cua processes automation through a six-step pipeline that separates LLM prompting from sandbox execution:

  1. Action-space string declaration: A plain-text block (e.g., UITARS_ACTION_SPACE in libs/python/agent/cua_agent/loops/uitars.py) defines available functions, their signatures, and descriptions. This string is interpolated into the prompt template at the {action_space} placeholder.

  2. Prompt template composition: The template (e.g., UITARS_PROMPT_TEMPLATE or GLM45V_PROMPT_TEMPLATE) combines the task instruction, action space, and contextual screenshots before sending to the LLM.

  3. Model response generation: The LLM returns a structured function call matching one of the declared signatures, such as click(start_box='<|box_start|>(100,200)<|box_end|>').

  4. Text parsing: The parse_action function (for UITARS) or parse_glm_response (for GLM-4.5V) in uitars.py converts the textual call into a dictionary with function and args keys.

  5. Command construction: Factory functions like make_click_item or make_type_item in libs/python/agent/cua_agent/responses.py transform the parsed dictionary into a ComputerCall payload that the sandbox understands.

  6. Agent registration: The @register_agent decorator in the loop file makes the custom agent discoverable by the CLI and server.

Because the action space is simply a string, you can add, remove, or modify actions without touching core inference code. You only need to keep the action-space declaration, the parser, and the response factory synchronized.

Step-by-Step Guide to Creating a Custom Action Space

1. Create a New Loop Module

Create a Python file in libs/python/agent/cua_agent/loops/ (e.g., my_custom_loop.py). This module will house your action space definition, prompt template, and agent logic.

2. Define the Action Space String

Declare your available actions as a triple-quoted string following the function signature format. Each line represents one callable function the model may invoke.


# libs/python/agent/cua_agent/loops/my_custom_loop.py

MY_ACTION_SPACE = """
click(start_box='<|box_start|>(x,y)<|box_end|>')
type(content='')                     # type text into the active element

screenshot(region='<|box_start|>(x1,y1,x2,y2)<|box_end|>')  # capture a sub-region

"""

Include parameter names, types in comments, and special token formats (like bounding box markers) that your parser will recognize.

3. Configure the Prompt Template

Build a prompt template that injects the action space via the {action_space} placeholder.

MY_PROMPT_TEMPLATE = """You are a GUI agent. Perform the next step.

## Action Space

{action_space}

## Task

{instruction}

## Previous Screenshots

{screenshots}
"""

When constructing the final prompt, format the string with action_space=MY_ACTION_SPACE and the current task instruction.

4. Parse the Model Output

For standard function calls (func(arg='value')), reuse the generic parse_action implementation from uitars.py. If your action space uses custom syntax or token formats (like GLM-4.5V's specific bounding box markers), implement a dedicated parser similar to parse_glm_response.

The parser must return a dictionary structured as {"function": "func_name", "args": {"arg1": "value"}}.

5. Implement Response Factories

Add builder functions in libs/python/agent/cua_agent/responses.py (or a local helper module) that convert parsed arguments into sandbox-compatible payloads.

def make_screenshot_item(region: str) -> dict:
    # region format: "x1,y1,x2,y2"

    x1, y1, x2, y2 = map(int, region.split(','))
    return {
        "type": "computer_call",
        "action": "screenshot",
        "params": {"region": [x1, y1, x2, y2]},
    }

Existing factories like make_click_item and make_type_item in responses.py serve as reference implementations.

6. Wire the Execution Logic

Within your agent loop, connect the parser output to the appropriate factory based on the function name.

from ..responses import make_click_item, make_type_item, make_screenshot_item

def build_computer_call(parsed: dict):
    fn = parsed["function"]
    args = parsed["args"]
    
    if fn == "click":
        return make_click_item(**args)
    elif fn == "type":
        return make_type_item(content=args["content"])
    elif fn == "screenshot":
        return make_screenshot_item(region=args["region"])
    else:
        raise ValueError(f"Unsupported action: {fn}")

7. Register the Agent

Expose your loop to the Cua CLI and server using the @register_agent decorator, specifying the agent name and capabilities.

from ..decorators import register_agent

@register_agent(name="my_custom_agent", capabilities=["step"])
async def my_custom_agent_loop(messages, model, **cfg):
    # Build prompt with injected action space

    prompt = MY_PROMPT_TEMPLATE.format(
        action_space=MY_ACTION_SPACE,
        instruction=messages[-1]["content"],
        screenshots=cfg.get("screenshots", [])
    )
    
    # Send to LLM via AsyncAgentConfig

    response = await cfg["agent_config"].predict_step(
        messages=[{"role": "user", "content": prompt}],
        model=model
    )
    
    # Parse and convert to computer call

    action_dict = parse_action(response["output"])
    computer_call = build_computer_call(action_dict)
    
    return {"output": [computer_call], "usage": response.get("usage", [])}

8. Test the Implementation

Invoke your custom agent from the command line to verify the sandbox receives correct JSON payloads:

cua run --agent my_custom_agent --task "Take a screenshot of the top-left corner"

Complete Implementation Example

Here is a condensed example combining the declaration, parsing, and response logic:


# libs/python/agent/cua_agent/loops/my_custom_loop.py

from ..decorators import register_agent
from ..responses import make_click_item, make_type_item

MY_ACTION_SPACE = """
click(start_box='<|box_start|>(x,y)<|box_end|>')
type(content='')                     
screenshot(region='<|box_start|>(x1,y1,x2,y2)<|box_end|>')  
"""

MY_PROMPT_TEMPLATE = """You are a GUI agent. ...

## Action Space

{action_space}

## Task

{instruction}
"""

def parse_action(action_str: str):
    # Simplified parser implementation

    import re
    match = re.match(r'(\w+)\((.*?)\)', action_str.strip())
    if not match:
        return None
    fn_name, args_str = match.groups()
    args = {}
    for arg in args_str.split(','):
        if '=' in arg:
            k, v = arg.split('=', 1)
            args[k.strip()] = v.strip().strip("'\"")
    return {"function": fn_name, "args": args}

def make_screenshot_item(region: str) -> dict:
    x1, y1, x2, y2 = map(int, region.split(','))
    return {
        "type": "computer_call",
        "action": "screenshot",
        "params": {"region": [x1, y1, x2, y2]},
    }

@register_agent(name="my_custom_agent", capabilities=["step"])
async def my_custom_agent_loop(messages, model, **cfg):
    prompt = MY_PROMPT_TEMPLATE.format(
        action_space=MY_ACTION_SPACE,
        instruction=messages[-1]["content"],
    )
    
    response = await cfg["agent_config"].predict_step(
        messages=[{"role": "user", "content": prompt}], 
        model=model
    )
    
    parsed = parse_action(response["output"])
    if parsed["function"] == "screenshot":
        item = make_screenshot_item(parsed["args"]["region"])
    elif parsed["function"] == "click":
        item = make_click_item(**parsed["args"])
    else:
        item = make_type_item(content=parsed["args"]["content"])
    
    return {"output": [item], "usage": response.get("usage", [])}

Key Files in the Action Space Workflow

Understanding these source files is critical when implementing custom action spaces:

Summary

  • Action spaces are plain-text strings, not hard-coded logic, residing in agent loop files like uitars.py.
  • Three components must stay synchronized: the action-space declaration, the parser (e.g., parse_action), and the response factory (e.g., make_click_item).
  • New actions require: adding the function signature to the action-space string, implementing a builder in responses.py, and updating the execution logic in your loop.
  • Registration is mandatory: Use @register_agent to make custom agents discoverable by the Cua CLI.
  • Reference existing implementations: Study uitars.py for standard parsing and glm45v.py for custom token formats.

Frequently Asked Questions

How do I add a completely new action type that doesn't exist in the base Cua implementation?

Define the function signature in your custom ACTION_SPACE string, then implement a corresponding factory function in responses.py (or your loop module) that returns a dictionary with "type": "computer_call", the action name, and any required parameters. Update your loop's execution logic to route the parsed function name to this new factory. As long as the sandbox recognizes the action name in the payload, the automation will execute.

Can I remove standard actions like click or type from the action space?

Yes. Since the action space is simply a string injected into the prompt, removing a function declaration from the ACTION_SPACE string prevents the model from calling it. Ensure you also remove or modify the corresponding conditional branches in your execution logic that would otherwise handle those parsed actions, or raise appropriate errors if the model somehow generates them.

What is the difference between parse_action and parse_glm_response?

parse_action (in libs/python/agent/cua_agent/loops/uitars.py) handles standard Python-like function calls with simple key='value' arguments. parse_glm_response (specific to GLM-4.5V agents) processes responses that use custom bounding-box token formats like <|box_start|>(x,y)<|box_end|>. Use parse_action for standard signatures; implement a custom parser only if your action space uses unique tokenization or syntax that the generic regex cannot capture.

Where should I place my custom action space definition files?

Create new loop files under libs/python/agent/cua_agent/loops/ (e.g., my_custom_loop.py) alongside existing implementations like uitars.py and glm45v.py. This location ensures your module can import from ..responses and ..decorators using relative imports, and keeps your custom agent aligned with the project's architecture for registration and discovery.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →