Cua Agent Reasoning Loop Architecture: How It Handles Tool Failures

The Cua Agent implements a modular async loop architecture that converts LLM responses into actionable computer commands while gracefully degrading malformed tool calls to generic function calls to prevent execution crashes.

The trycua/cua repository provides an open-source agent framework built around a protocol-driven Cua Agent reasoning loop designed for computer-use automation. Understanding this architecture reveals how the system processes multimodal inputs, manages browser interactions, and maintains robustness when tools fail or return invalid data.

Understanding the Cua Agent Reasoning Loop Architecture

At its core, the Cua Agent follows a modular, async loop architecture where each loop class implements the AsyncAgentConfig protocol defined in libs/python/agent/cua_agent/loops/base.py lines 11-28【^1†L11-L28】. This protocol standardizes how different LLM providers interact with computer environments while allowing provider-specific optimizations.

The AsyncAgentConfig Protocol

Every reasoning loop adheres to the AsyncAgentConfig protocol, which mandates three key methods:

  • predict_step: Processes conversation history and generates the next action
  • predict_click: Provides specialized coordinate prediction capabilities
  • get_capabilities: Returns supported feature flags for the specific implementation

Classes like YutoriN1Config, OpenAIComputerUseConfig, and GenericVlmConfig all implement this protocol, ensuring consistent behavior across OpenAI, Yutori, and generic vision-language models.

The Seven-Step Execution Flow

Each iteration of the Cua Agent reasoning loop follows a precise sequence orchestrated across multiple modules:

  1. Input Preparation: The loop transforms incoming Responses-API items into chat-completion payloads via convert_responses_items_to_completion_messages (yutori.py lines 30-34【^1†L30-L34】).

  2. Screenshot Injection: For browser-focused loops, if the model receives no image, the system automatically captures a screenshot from the computer_handler and injects it as a WebP image (yutori.py lines 49-71【^1†L49-L71】).

  3. Tool List Preparation: Custom function tools pass through unchanged, while native browser tools are omitted because the model already knows its built-in actions (yutori.py lines 85-93【^1†L85-L93】).

  4. LLM Request: The loop dispatches requests via litellm, using acompletion for chat-style models (yutori.py line 111【^1†L111】) or aresponses for the OpenAI Responses API (openai.py line 45【^1†L45】).

  5. Usage Extraction: After the API call, the loop extracts token statistics and cost data, exposing them through the _on_usage callback (yutori.py lines 166-176【^1†L166-L176】; openai.py lines 51-74【^1†L51-L74】).

  6. Response Parsing: The system iterates through tool_calls, converting each to an internal computer_call via _convert_n1_action_to_computer_action (yutori.py lines 66-124【^1†L66-L124】), or preserving them as function_call items for custom tools.

  7. Reasoning Handling: If the model provides a separate reasoning field, the loop creates a ResponseReasoningItem and adds it to the output (yutori.py lines 36-38【^1†L36-L38】).

The loop returns a dictionary with "output" (list of response items) and "usage" (token metrics) at yutori.py line 86【^1†L86】 that the Cua runtime passes to the next iteration.

How the Cua Agent Handles Tool Failures

Robustness against tool execution failures is engineered into multiple layers of the Cua Agent reasoning loop, from conversion fallbacks to explicit status enums.

Graceful Degradation for Invalid Tool Calls

When the LLM returns malformed or unsupported tool calls, the _convert_n1_action_to_computer_action function returns None instead of crashing. This triggers a fallback path (lines 81-84【^1†L81-L84】) where the invalid call converts to a generic function_call item rather than a computer_call.

This design ensures that missing coordinates or unknown action types do not abort the execution loop. Instead, the system surfaces the failure as a standard function call that downstream components can handle appropriately.

Runtime Validation and Early Failure

The architecture employs fail-fast validation for critical dependencies. If the browser loop requires a screenshot but the computer_handler lacks a screenshot method, yutori.py raises a clear runtime error immediately (lines 52-55【^1†L52-L55】) rather than producing an invalid API request. This explicit error signaling makes integration issues obvious during development.

Defensive Error Handling and Retry Logic

Every external interaction wraps in broad-scope try/except blocks:

  • Image conversion failures in _prepare_image_for_n1 return the original image rather than aborting
  • Usage extraction in openai.py handles both dictionary and Pydantic response formats safely (lines 51-64【^1†L51-L64】)

Higher-level loops like ComposedGrounded and Moondream3 implement explicit retry logic, wrapping predict_step in a for _ in range(3): loop (see composed_grounded.py line 275【^1†L275】) that only breaks on successful execution, automatically recovering from transient API failures.

Explicit Status Enums in Tool Schemas

The native computer_use_preview tool schema defines a status property with an enum ["success", "failure"] (lines 30-31【^1†L30-L31】). Downstream components inspecting computer_call items can check this field to determine whether to retry the action, log the error, or terminate the session, providing structured failure semantics rather than relying solely on exceptions.

Practical Implementation Examples

Basic Usage with YutoriN1Config

The following example demonstrates a single reasoning step with the browser-focused loop:

from cua.agent.cua_agent.loops.yutori import YutoriN1Config

loop = YutoriN1Config()

# Simulated prior messages in Responses-API format

messages = [
    {"type": "message", "role": "assistant", "content": "What should I click?"},
]

result = await loop.predict_step(
    messages=messages,
    model="yutori/n1",
    tools=[],
    computer_handler=my_browser_handler,
)

print(result["output"])  # Response items including reasoning, computer_call

print(result["usage"])   # Token and cost information

This loop automatically injects screenshots when no image is present and converts tool calls through the failure-resistant pipeline.

Handling Tool Conversion Failures

When the model returns malformed arguments, the conversion layer prevents crashes:


# Malformed tool call missing required coordinates

bad_tool = {
    "id": "call_1",
    "type": "function", 
    "function": {"name": "left_click", "arguments": "{}"}
}

# Inside predict_step, conversion returns None for invalid actions

computer_action = _convert_n1_action_to_computer_action(
    fn_name="left_click",
    args={},
    screen_width=1280,
    screen_height=800,
)

# Fallback to function_call prevents exception propagation

if computer_action is None:
    output.append(make_function_call_item("left_click", {}, call_id="call_1"))

This fallback path in yutori.py lines 81-84【^1†L81-L84】 guarantees that no uncaught exception propagates to the outer loop.

OpenAI Computer-Use Integration

For OpenAI's native computer-use models:

from cua.agent.cua_agent.loops.openai import OpenAIComputerUseConfig

loop = OpenAIComputerUseConfig()
response = await loop.predict_step(
    messages=[{"role": "user", "content": "Open the settings page"}],
    model="gpt-4o-computer-use-preview",
    tools=[{"type": "computer", "computer": my_computer_handler}],
)

This loop automatically builds the proper tool schema via _prepare_tools_for_openai (openai.py lines 47-77【^1†L47-L77】) and safely extracts usage statistics even when response formats vary.

Summary

  • Protocol-driven design: All loops implement AsyncAgentConfig from base.py, standardizing predict_step, predict_click, and get_capabilities across providers.
  • Modular conversion layer: Shared utilities translate between Responses API and LLM provider formats, with specific logic in yutori.py and openai.py handling provider quirks.
  • Failure resilience: The architecture uses graceful degradation (converting invalid computer calls to function calls), early validation for missing handlers, and explicit status enums to handle tool failures without crashing the execution loop.
  • Defensive programming: Extensive try/except guards around image processing, API calls, and usage extraction ensure that partial failures do not abort entire sessions.

Frequently Asked Questions

What protocol defines the Cua Agent reasoning loop interface?

The AsyncAgentConfig protocol defined in libs/python/agent/cua_agent/loops/base.py lines 11-28【^1†L11-L28】 specifies the required interface. All loop implementations must provide predict_step, predict_click, and get_capabilities methods, ensuring consistent behavior whether using YutoriN1Config, OpenAIComputerUseConfig, or custom VLM loops.

How does the Cua Agent prevent crashes from malformed tool arguments?

When _convert_n1_action_to_computer_action encounters unsupported or malformed actions (such as missing coordinates), it returns None instead of raising an exception. The calling code in yutori.py lines 81-84【^1†L81-L84】 then falls back to emitting a generic function_call item, allowing the loop to continue execution while surfacing the failure in a structured format.

What happens when a screenshot cannot be captured in browser loops?

If the computer_handler does not expose a screenshot method and no image is present in the context, yutori.py raises a runtime error with a clear message at lines 52-55【^1†L52-L55】. This fail-fast approach prevents the creation of invalid API requests and makes integration configuration errors immediately obvious to developers.

Which loops implement retry logic for transient failures?

Higher-level implementations like ComposedGrounded and Moondream3 wrap the entire predict_step call in a retry loop (for _ in range(3)), as seen in composed_grounded.py line 275【^1†L275】. This automatic retry mechanism recovers from transient network errors or temporary API unavailability without requiring manual intervention from the calling code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →