How to Implement Computer Use and GUI Automation Agents: A Complete Guide

To implement computer use and GUI automation agents, create a tool-registered agent that captures screenshots, sends them to an LLM with computer-use capabilities (OpenAI CUA or Anthropic), parses the returned action commands, and executes them through CDP or native OS APIs.

The bojieli/ai-agent-book repository provides production-ready reference implementations for building agents that control graphical user interfaces. This guide walks through the architecture, implementation patterns, and code paths used to automate browsers and desktop environments with large language models.

Core Architecture for GUI Automation Agents

A computer-use agent requires four coordinated components working in sequence: the agent orchestrator, tool registry, LLM interface with computer-use capabilities, and action executor.

Component Role Location in Repository
Agent core Orchestrates task execution and maintains action history agentbook/agent.py
Tool registry Registers invocable actions as JSON-schema tools browser_use/Tools.py
LLM computer-use interface Sends screenshots, receives structured action commands OpenAI CUA or Anthropic sampling_loop
Action executor Translates commands into CDP or OS-level interactions cua.py handlers or computer_use_demo

The Agent class receives a high-level task (e.g., "search Google for today's weather"), builds a message history, and passes it to the LLM. When the model determines a concrete GUI operation is needed, it calls a registered tool with parameters defined by a Pydantic BaseModel.

Implementation Pattern 1: OpenAI Computer-Use Assistant (CUA)

The OpenAI CUA fallback in chapter8/browser-use-rpa/browser-use/examples/custom-functions/cua.py demonstrates the complete flow for browser automation using the computer-use-preview model.

Step 1: Capture Browser State

The agent requests current visual state from the BrowserSession:

state = await browser_session.get_browser_state_summary()
screenshot = Image.open(BytesIO(base64.b64decode(state.screenshot)))
screenshot = screenshot.resize((
    state.page_info.viewport_width,
    state.page_info.viewport_height
))

Step 2: Construct Multimodal Prompt

The screenshot is base64-encoded and sent alongside a textual description:

buf = BytesIO()
screenshot.save(buf, format="PNG")
screenshot_b64 = base64.b64encode(buf.getvalue()).decode()

response = await client.responses.create(
    model="computer-use-preview",
    tools=[{
        "type": "computer_use_preview",
        "display_width": state.page_info.viewport_width,
        "display_height": state.page_info.viewport_height,
        "environment": "browser",
    }],
    input=[{
        "role": "user",
        "content": [
            {"type": "input_text", "text": params.description},
            {"type": "input_image", "detail": "auto",
             "image_url": f"data:image/png;base64,{screenshot_b64}"},
        ],
    }],
)

Step 3: Parse and Execute Computer Calls

The response contains a computer_call object describing the UI action (click, keypress, scroll, etc.). The handle_model_action function translates this into Chrome DevTools Protocol (CDP) commands:

computer_call = next(
    (c for c in response.output if c.type == "computer_call"),
    None,
)
if not computer_call:
    return ActionResult(error="No computer call returned")
return await handle_model_action(browser_session, computer_call.action)

CDP commands include Input.dispatchMouseEvent, Input.dispatchKeyEvent, and Input.synthesizePinchGesture for touch emulation.

Implementation Pattern 2: Anthropic Native Computer Use

The Claude implementation in chapter9/claude-computer-use-native/run_weather_task.py uses Anthropic's built-in computer_use_demo package, which handles action execution internally.

Bounded Sampling Loop with Trajectory Logging

from computer_use_demo.loop import APIProvider, sampling_loop

TASK = (
    "Open Google, search for San Francisco weather today, "
    "and report the temperature."
)

async def main():
    await sampling_loop(
        model="claude-sonnet-4-5-20250929",
        provider=APIProvider.ANTHROPIC,
        system_prompt_suffix=(
            "Perform a read-only GUI search, stop once the weather "
            "JSON is visible. Do not sign in or click CAPTCHAs."
        ),
        messages=[{
            "role": "user",
            "content": [{"type": "text", "text": TASK}]
        }],
        tool_version="computer_use_20250124",
        action_limit=25,
        api_key=os.getenv("ANTHROPIC_API_KEY"),
    )

Key differences from the OpenAI approach:

  • No manual action translation — the computer_use_demo library executes GUI actions directly
  • Built-in safety boundaries — action_limit prevents runaway automation
  • Automatic trajectory recording — every API call, screenshot, and action is logged to trajectory.json

Implementation Pattern 3: Swappable Action Spaces

The run_multienv_aworldAgent.py file demonstrates configuring the same agent with different action backends via command-line flags.

Selecting Backend at Runtime

if args.action_space == "pyautogui":
    # Build command string for pyautogui interpreter

    fixed_command = _fix_pyautogui_less_than_bug(action)
elif args.action_space == "claude_computer_use":
    # Send action to Claude computer-use tool

    computer_call = await claude_client.run(action)
    fixed_command = generate_python_from_computer_call(computer_call)

This architecture allows rapid prototyping: test with pyautogui for transparency and debugging, then switch to claude_computer_use for production reliability.

Environment Detection and Setup

Before launching GUI automation, validate the execution environment. The helper in chapter4/execution-tools/extended_tools.py checks for required utilities:


# Validates presence of xvfb (X virtual framebuffer), xwd (X window dump),

# and other GUI dependencies in the container

computer_use_container_image_present = check_container_gui_utils()

This prevents runtime failures when the Docker image lacks display server components.

Complete Working Example: OpenAI CUA Tool Registration

Register the CUA fallback as a tool the LLM can invoke when standard actions fail:

from browser_use import Agent, ChatOpenAI, Tools
from browser_use.agent.views import ActionResult
from browser_use.browser import BrowserSession
from openai import AsyncOpenAI
from pydantic import BaseModel, Field
import base64, os
from io import BytesIO
from PIL import Image

tools = Tools()

class OpenAICUAAction(BaseModel):
    description: str = Field(
        ...,
        description="What the agent should achieve next"
    )

@tools.registry.action(
    "Fallback to OpenAI Computer Use Assistant when normal actions fail",
    param_model=OpenAICUAAction,
)
async def openai_cua_fallback(
    params: OpenAICUAAction,
    browser_session: BrowserSession
):
    # State capture, prompt construction, and execution

    # as detailed in the previous sections

    ...

Key Implementation Files

File Purpose
chapter8/browser-use-rpa/browser-use/examples/custom-functions/cua.py Full OpenAI CUA implementation with CDP execution
chapter9/claude-computer-use-native/run_weather_task.py Anthropic native computer-use with bounded sampling
chapter8/gaia-experience/AWorld/examples/osworld/run_multienv_aworldAgent.py Swappable action spaces (pyautogui/Claude)
chapter4/execution-tools/extended_tools.py Container environment validation
chapter8/gaia-experience/AWorld/examples/osworld/aworldAgent/grounding.py PyAutoGUI helper functions for OS-level control
chapter8/browser-use-rpa/browser-use/browser_use/tools.py Tool registration infrastructure

Summary

  • Agent orchestration — Use the Agent class to manage task state and LLM conversation history
  • Tool registration — Define actions with Pydantic BaseModel parameters and the @tools.registry.action decorator
  • OpenAI CUA — Capture screenshots, send to computer-use-preview, parse computer_call responses, execute via CDP
  • Anthropic native — Use sampling_loop with computer_use_20250124 tool version for built-in GUI control
  • Action space flexibility — Configure agents to use pyautogui, claude_computer_use, or openai_cua backends interchangeably
  • Environment validation — Check for xvfb, xwd, and container requirements before execution

Frequently Asked Questions

What is the difference between OpenAI CUA and Anthropic computer use?

OpenAI CUA requires manual implementation of the action execution layer—you capture screenshots, send them to the computer-use-preview model, parse the returned computer_call object, and translate it into CDP or OS commands yourself. Anthropic's computer-use tool bundles the execution logic internally through computer_use_demo, so you only need to call sampling_loop with appropriate parameters.

How do I handle action limits and prevent infinite loops?

Both implementations support bounded execution. In the Anthropic example, pass action_limit=25 to sampling_loop. For OpenAI CUA, wrap the agent loop with a counter check and terminate when the threshold is exceeded. The trajectory.json output in Anthropic's implementation provides audit trails for debugging runaway behavior.

Can I use computer-use agents for desktop applications beyond browsers?

Yes. The pyautogui action space in run_multienv_aworldAgent.py operates at the OS level and can interact with any visible window. For browser-specific automation, the CUA implementation with CDP provides more reliable element targeting. Choose based on whether your target application exposes a web interface or requires native GUI interaction.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →