How to Implement Custom Action Spaces for Specific Automation Tasks in Cua
You implement custom action spaces in Cua by defining a plain-text function signature block, injecting it into the agent's prompt template, parsing the LLM's structured response, and mapping parsed actions to sandbox commands using response-item factories.
Cua is an open-source computer-use automation framework that drives GUI interactions through language models. Unlike rigid automation tools, Cua's action space—the set of functions available to the model—is not hard-coded in the core engine but exists as a configurable text block within agent loops. This architecture allows you to tailor available actions for specific automation scenarios without modifying the underlying inference logic.
How the Action Space Mechanism Works
Cua processes automation through a six-step pipeline that separates LLM prompting from sandbox execution:
-
Action-space string declaration: A plain-text block (e.g.,
UITARS_ACTION_SPACEinlibs/python/agent/cua_agent/loops/uitars.py) defines available functions, their signatures, and descriptions. This string is interpolated into the prompt template at the{action_space}placeholder. -
Prompt template composition: The template (e.g.,
UITARS_PROMPT_TEMPLATEorGLM45V_PROMPT_TEMPLATE) combines the task instruction, action space, and contextual screenshots before sending to the LLM. -
Model response generation: The LLM returns a structured function call matching one of the declared signatures, such as
click(start_box='<|box_start|>(100,200)<|box_end|>'). -
Text parsing: The
parse_actionfunction (for UITARS) orparse_glm_response(for GLM-4.5V) inuitars.pyconverts the textual call into a dictionary withfunctionandargskeys. -
Command construction: Factory functions like
make_click_itemormake_type_iteminlibs/python/agent/cua_agent/responses.pytransform the parsed dictionary into aComputerCallpayload that the sandbox understands. -
Agent registration: The
@register_agentdecorator in the loop file makes the custom agent discoverable by the CLI and server.
Because the action space is simply a string, you can add, remove, or modify actions without touching core inference code. You only need to keep the action-space declaration, the parser, and the response factory synchronized.
Step-by-Step Guide to Creating a Custom Action Space
1. Create a New Loop Module
Create a Python file in libs/python/agent/cua_agent/loops/ (e.g., my_custom_loop.py). This module will house your action space definition, prompt template, and agent logic.
2. Define the Action Space String
Declare your available actions as a triple-quoted string following the function signature format. Each line represents one callable function the model may invoke.
# libs/python/agent/cua_agent/loops/my_custom_loop.py
MY_ACTION_SPACE = """
click(start_box='<|box_start|>(x,y)<|box_end|>')
type(content='') # type text into the active element
screenshot(region='<|box_start|>(x1,y1,x2,y2)<|box_end|>') # capture a sub-region
"""
Include parameter names, types in comments, and special token formats (like bounding box markers) that your parser will recognize.
3. Configure the Prompt Template
Build a prompt template that injects the action space via the {action_space} placeholder.
MY_PROMPT_TEMPLATE = """You are a GUI agent. Perform the next step.
## Action Space
{action_space}
## Task
{instruction}
## Previous Screenshots
{screenshots}
"""
When constructing the final prompt, format the string with action_space=MY_ACTION_SPACE and the current task instruction.
4. Parse the Model Output
For standard function calls (func(arg='value')), reuse the generic parse_action implementation from uitars.py. If your action space uses custom syntax or token formats (like GLM-4.5V's specific bounding box markers), implement a dedicated parser similar to parse_glm_response.
The parser must return a dictionary structured as {"function": "func_name", "args": {"arg1": "value"}}.
5. Implement Response Factories
Add builder functions in libs/python/agent/cua_agent/responses.py (or a local helper module) that convert parsed arguments into sandbox-compatible payloads.
def make_screenshot_item(region: str) -> dict:
# region format: "x1,y1,x2,y2"
x1, y1, x2, y2 = map(int, region.split(','))
return {
"type": "computer_call",
"action": "screenshot",
"params": {"region": [x1, y1, x2, y2]},
}
Existing factories like make_click_item and make_type_item in responses.py serve as reference implementations.
6. Wire the Execution Logic
Within your agent loop, connect the parser output to the appropriate factory based on the function name.
from ..responses import make_click_item, make_type_item, make_screenshot_item
def build_computer_call(parsed: dict):
fn = parsed["function"]
args = parsed["args"]
if fn == "click":
return make_click_item(**args)
elif fn == "type":
return make_type_item(content=args["content"])
elif fn == "screenshot":
return make_screenshot_item(region=args["region"])
else:
raise ValueError(f"Unsupported action: {fn}")
7. Register the Agent
Expose your loop to the Cua CLI and server using the @register_agent decorator, specifying the agent name and capabilities.
from ..decorators import register_agent
@register_agent(name="my_custom_agent", capabilities=["step"])
async def my_custom_agent_loop(messages, model, **cfg):
# Build prompt with injected action space
prompt = MY_PROMPT_TEMPLATE.format(
action_space=MY_ACTION_SPACE,
instruction=messages[-1]["content"],
screenshots=cfg.get("screenshots", [])
)
# Send to LLM via AsyncAgentConfig
response = await cfg["agent_config"].predict_step(
messages=[{"role": "user", "content": prompt}],
model=model
)
# Parse and convert to computer call
action_dict = parse_action(response["output"])
computer_call = build_computer_call(action_dict)
return {"output": [computer_call], "usage": response.get("usage", [])}
8. Test the Implementation
Invoke your custom agent from the command line to verify the sandbox receives correct JSON payloads:
cua run --agent my_custom_agent --task "Take a screenshot of the top-left corner"
Complete Implementation Example
Here is a condensed example combining the declaration, parsing, and response logic:
# libs/python/agent/cua_agent/loops/my_custom_loop.py
from ..decorators import register_agent
from ..responses import make_click_item, make_type_item
MY_ACTION_SPACE = """
click(start_box='<|box_start|>(x,y)<|box_end|>')
type(content='')
screenshot(region='<|box_start|>(x1,y1,x2,y2)<|box_end|>')
"""
MY_PROMPT_TEMPLATE = """You are a GUI agent. ...
## Action Space
{action_space}
## Task
{instruction}
"""
def parse_action(action_str: str):
# Simplified parser implementation
import re
match = re.match(r'(\w+)\((.*?)\)', action_str.strip())
if not match:
return None
fn_name, args_str = match.groups()
args = {}
for arg in args_str.split(','):
if '=' in arg:
k, v = arg.split('=', 1)
args[k.strip()] = v.strip().strip("'\"")
return {"function": fn_name, "args": args}
def make_screenshot_item(region: str) -> dict:
x1, y1, x2, y2 = map(int, region.split(','))
return {
"type": "computer_call",
"action": "screenshot",
"params": {"region": [x1, y1, x2, y2]},
}
@register_agent(name="my_custom_agent", capabilities=["step"])
async def my_custom_agent_loop(messages, model, **cfg):
prompt = MY_PROMPT_TEMPLATE.format(
action_space=MY_ACTION_SPACE,
instruction=messages[-1]["content"],
)
response = await cfg["agent_config"].predict_step(
messages=[{"role": "user", "content": prompt}],
model=model
)
parsed = parse_action(response["output"])
if parsed["function"] == "screenshot":
item = make_screenshot_item(parsed["args"]["region"])
elif parsed["function"] == "click":
item = make_click_item(**parsed["args"])
else:
item = make_type_item(content=parsed["args"]["content"])
return {"output": [item], "usage": response.get("usage", [])}
Key Files in the Action Space Workflow
Understanding these source files is critical when implementing custom action spaces:
libs/python/agent/cua_agent/loops/uitars.py: Reference implementation containingUITARS_ACTION_SPACE,UITARS_PROMPT_TEMPLATE, and the genericparse_actionfunction.libs/python/agent/cua_agent/loops/glm45v.py: Demonstrates complex action spaces (GLM_ACTION_SPACE) and custom parsing logic viaparse_glm_response.libs/python/agent/cua_agent/responses.py: Contains factory functions (make_click_item,make_type_item, etc.) that constructComputerCallpayloads for the sandbox.libs/python/agent/cua_agent/loops/base.py: Defines theAsyncAgentConfigprotocol that each loop must implement to communicate with the LLM and sandbox.libs/python/agent/cua_agent/decorators.py: Provides the@register_agentdecorator used to expose custom loops to the CLI and server.
Summary
- Action spaces are plain-text strings, not hard-coded logic, residing in agent loop files like
uitars.py. - Three components must stay synchronized: the action-space declaration, the parser (e.g.,
parse_action), and the response factory (e.g.,make_click_item). - New actions require: adding the function signature to the action-space string, implementing a builder in
responses.py, and updating the execution logic in your loop. - Registration is mandatory: Use
@register_agentto make custom agents discoverable by the Cua CLI. - Reference existing implementations: Study
uitars.pyfor standard parsing andglm45v.pyfor custom token formats.
Frequently Asked Questions
How do I add a completely new action type that doesn't exist in the base Cua implementation?
Define the function signature in your custom ACTION_SPACE string, then implement a corresponding factory function in responses.py (or your loop module) that returns a dictionary with "type": "computer_call", the action name, and any required parameters. Update your loop's execution logic to route the parsed function name to this new factory. As long as the sandbox recognizes the action name in the payload, the automation will execute.
Can I remove standard actions like click or type from the action space?
Yes. Since the action space is simply a string injected into the prompt, removing a function declaration from the ACTION_SPACE string prevents the model from calling it. Ensure you also remove or modify the corresponding conditional branches in your execution logic that would otherwise handle those parsed actions, or raise appropriate errors if the model somehow generates them.
What is the difference between parse_action and parse_glm_response?
parse_action (in libs/python/agent/cua_agent/loops/uitars.py) handles standard Python-like function calls with simple key='value' arguments. parse_glm_response (specific to GLM-4.5V agents) processes responses that use custom bounding-box token formats like <|box_start|>(x,y)<|box_end|>. Use parse_action for standard signatures; implement a custom parser only if your action space uses unique tokenization or syntax that the generic regex cannot capture.
Where should I place my custom action space definition files?
Create new loop files under libs/python/agent/cua_agent/loops/ (e.g., my_custom_loop.py) alongside existing implementations like uitars.py and glm45v.py. This location ensures your module can import from ..responses and ..decorators using relative imports, and keeps your custom agent aligned with the project's architecture for registration and discovery.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →