How to Implement Computer Use and GUI Automation Agents: A Complete Guide
To implement computer use and GUI automation agents, create a tool-registered agent that captures screenshots, sends them to an LLM with computer-use capabilities (OpenAI CUA or Anthropic), parses the returned action commands, and executes them through CDP or native OS APIs.
The bojieli/ai-agent-book repository provides production-ready reference implementations for building agents that control graphical user interfaces. This guide walks through the architecture, implementation patterns, and code paths used to automate browsers and desktop environments with large language models.
Core Architecture for GUI Automation Agents
A computer-use agent requires four coordinated components working in sequence: the agent orchestrator, tool registry, LLM interface with computer-use capabilities, and action executor.
| Component | Role | Location in Repository |
|---|---|---|
| Agent core | Orchestrates task execution and maintains action history | agentbook/agent.py |
| Tool registry | Registers invocable actions as JSON-schema tools | browser_use/Tools.py |
| LLM computer-use interface | Sends screenshots, receives structured action commands | OpenAI CUA or Anthropic sampling_loop |
| Action executor | Translates commands into CDP or OS-level interactions | cua.py handlers or computer_use_demo |
The Agent class receives a high-level task (e.g., "search Google for today's weather"), builds a message history, and passes it to the LLM. When the model determines a concrete GUI operation is needed, it calls a registered tool with parameters defined by a Pydantic BaseModel.
Implementation Pattern 1: OpenAI Computer-Use Assistant (CUA)
The OpenAI CUA fallback in chapter8/browser-use-rpa/browser-use/examples/custom-functions/cua.py demonstrates the complete flow for browser automation using the computer-use-preview model.
Step 1: Capture Browser State
The agent requests current visual state from the BrowserSession:
state = await browser_session.get_browser_state_summary()
screenshot = Image.open(BytesIO(base64.b64decode(state.screenshot)))
screenshot = screenshot.resize((
state.page_info.viewport_width,
state.page_info.viewport_height
))
Step 2: Construct Multimodal Prompt
The screenshot is base64-encoded and sent alongside a textual description:
buf = BytesIO()
screenshot.save(buf, format="PNG")
screenshot_b64 = base64.b64encode(buf.getvalue()).decode()
response = await client.responses.create(
model="computer-use-preview",
tools=[{
"type": "computer_use_preview",
"display_width": state.page_info.viewport_width,
"display_height": state.page_info.viewport_height,
"environment": "browser",
}],
input=[{
"role": "user",
"content": [
{"type": "input_text", "text": params.description},
{"type": "input_image", "detail": "auto",
"image_url": f"data:image/png;base64,{screenshot_b64}"},
],
}],
)
Step 3: Parse and Execute Computer Calls
The response contains a computer_call object describing the UI action (click, keypress, scroll, etc.). The handle_model_action function translates this into Chrome DevTools Protocol (CDP) commands:
computer_call = next(
(c for c in response.output if c.type == "computer_call"),
None,
)
if not computer_call:
return ActionResult(error="No computer call returned")
return await handle_model_action(browser_session, computer_call.action)
CDP commands include Input.dispatchMouseEvent, Input.dispatchKeyEvent, and Input.synthesizePinchGesture for touch emulation.
Implementation Pattern 2: Anthropic Native Computer Use
The Claude implementation in chapter9/claude-computer-use-native/run_weather_task.py uses Anthropic's built-in computer_use_demo package, which handles action execution internally.
Bounded Sampling Loop with Trajectory Logging
from computer_use_demo.loop import APIProvider, sampling_loop
TASK = (
"Open Google, search for San Francisco weather today, "
"and report the temperature."
)
async def main():
await sampling_loop(
model="claude-sonnet-4-5-20250929",
provider=APIProvider.ANTHROPIC,
system_prompt_suffix=(
"Perform a read-only GUI search, stop once the weather "
"JSON is visible. Do not sign in or click CAPTCHAs."
),
messages=[{
"role": "user",
"content": [{"type": "text", "text": TASK}]
}],
tool_version="computer_use_20250124",
action_limit=25,
api_key=os.getenv("ANTHROPIC_API_KEY"),
)
Key differences from the OpenAI approach:
- No manual action translation — the
computer_use_demolibrary executes GUI actions directly - Built-in safety boundaries —
action_limitprevents runaway automation - Automatic trajectory recording — every API call, screenshot, and action is logged to
trajectory.json
Implementation Pattern 3: Swappable Action Spaces
The run_multienv_aworldAgent.py file demonstrates configuring the same agent with different action backends via command-line flags.
Selecting Backend at Runtime
if args.action_space == "pyautogui":
# Build command string for pyautogui interpreter
fixed_command = _fix_pyautogui_less_than_bug(action)
elif args.action_space == "claude_computer_use":
# Send action to Claude computer-use tool
computer_call = await claude_client.run(action)
fixed_command = generate_python_from_computer_call(computer_call)
This architecture allows rapid prototyping: test with pyautogui for transparency and debugging, then switch to claude_computer_use for production reliability.
Environment Detection and Setup
Before launching GUI automation, validate the execution environment. The helper in chapter4/execution-tools/extended_tools.py checks for required utilities:
# Validates presence of xvfb (X virtual framebuffer), xwd (X window dump),
# and other GUI dependencies in the container
computer_use_container_image_present = check_container_gui_utils()
This prevents runtime failures when the Docker image lacks display server components.
Complete Working Example: OpenAI CUA Tool Registration
Register the CUA fallback as a tool the LLM can invoke when standard actions fail:
from browser_use import Agent, ChatOpenAI, Tools
from browser_use.agent.views import ActionResult
from browser_use.browser import BrowserSession
from openai import AsyncOpenAI
from pydantic import BaseModel, Field
import base64, os
from io import BytesIO
from PIL import Image
tools = Tools()
class OpenAICUAAction(BaseModel):
description: str = Field(
...,
description="What the agent should achieve next"
)
@tools.registry.action(
"Fallback to OpenAI Computer Use Assistant when normal actions fail",
param_model=OpenAICUAAction,
)
async def openai_cua_fallback(
params: OpenAICUAAction,
browser_session: BrowserSession
):
# State capture, prompt construction, and execution
# as detailed in the previous sections
...
Key Implementation Files
| File | Purpose |
|---|---|
chapter8/browser-use-rpa/browser-use/examples/custom-functions/cua.py |
Full OpenAI CUA implementation with CDP execution |
chapter9/claude-computer-use-native/run_weather_task.py |
Anthropic native computer-use with bounded sampling |
chapter8/gaia-experience/AWorld/examples/osworld/run_multienv_aworldAgent.py |
Swappable action spaces (pyautogui/Claude) |
chapter4/execution-tools/extended_tools.py |
Container environment validation |
chapter8/gaia-experience/AWorld/examples/osworld/aworldAgent/grounding.py |
PyAutoGUI helper functions for OS-level control |
chapter8/browser-use-rpa/browser-use/browser_use/tools.py |
Tool registration infrastructure |
Summary
- Agent orchestration — Use the
Agentclass to manage task state and LLM conversation history - Tool registration — Define actions with Pydantic
BaseModelparameters and the@tools.registry.actiondecorator - OpenAI CUA — Capture screenshots, send to
computer-use-preview, parsecomputer_callresponses, execute via CDP - Anthropic native — Use
sampling_loopwithcomputer_use_20250124tool version for built-in GUI control - Action space flexibility — Configure agents to use
pyautogui,claude_computer_use, oropenai_cuabackends interchangeably - Environment validation — Check for
xvfb,xwd, and container requirements before execution
Frequently Asked Questions
What is the difference between OpenAI CUA and Anthropic computer use?
OpenAI CUA requires manual implementation of the action execution layer—you capture screenshots, send them to the computer-use-preview model, parse the returned computer_call object, and translate it into CDP or OS commands yourself. Anthropic's computer-use tool bundles the execution logic internally through computer_use_demo, so you only need to call sampling_loop with appropriate parameters.
How do I handle action limits and prevent infinite loops?
Both implementations support bounded execution. In the Anthropic example, pass action_limit=25 to sampling_loop. For OpenAI CUA, wrap the agent loop with a counter check and terminate when the threshold is exceeded. The trajectory.json output in Anthropic's implementation provides audit trails for debugging runaway behavior.
Can I use computer-use agents for desktop applications beyond browsers?
Yes. The pyautogui action space in run_multienv_aworldAgent.py operates at the OS level and can interact with any visible window. For browser-specific automation, the CUA implementation with CDP provides more reliable element targeting. Choose based on whether your target application exposes a web interface or requires native GUI interaction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →