How Cua Handles Screen State Verification and Element Detection for Reliable Automation
Cua combines AI-driven screenshot validation with dual-path element discovery—accessibility tree queries and computer vision OCR—to ensure every automated action hits the correct target.
The trycua/cua repository implements a robust automation framework that guarantees reliability through layered verification mechanisms. By integrating AWS Bedrock-based screen state verification with deterministic element detection strategies, Cua eliminates the brittleness of traditional pixel-matching approaches. This article examines the specific source code implementations that enable reliable cross-platform UI automation.
Screen State Verification with AWS Bedrock
Cua validates that expected UI elements are actually visible before proceeding with critical automation steps. This verification mechanism relies on multimodal AI analysis rather than fragile pixel comparisons.
The Verification Pipeline
After capturing a screenshot via ComputerInterface.screenshot, the system uploads the image to AWS Bedrock Claude-Haiku for analysis. The _verify_screenshot function in libs/python/cua-sandbox-apps/cua_sandbox_apps/pipeline/creator_agent.py constructs a multipart prompt containing the base64-encoded image and specific verification instructions.
The model must respond with a single line starting with PASS or FAIL, followed by explanatory text. This strict formatting allows the orchestration loop to parse results automatically and retry or abort when criteria are not met.
# libs/python/cua-sandbox-apps/cua_sandbox_apps/pipeline/creator_agent.py
def _verify_screenshot(image_path: Path, verification_prompt: str) -> tuple[bool, str]:
# Upload image to Bedrock with prompt
response = bedrock_client.invoke_model(
model_id="anthropic.claude-3-haiku",
body=json.dumps({
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": verification_prompt},
{"type": "image", "source": {"type": "base64", "data": image_base64}}
]
}]
})
)
# Parse PASS/FAIL response
output = json.loads(response['body'].read())['content'][0]['text']
passed = output.strip().startswith("PASS")
return passed, output
The orchestration code in the same file treats PASS as success; otherwise, it returns CRITERIA NOT MET and triggers an automatic retry of the step.
Element Detection Strategies
Cua discovers interactable UI elements through two complementary approaches: native accessibility tree queries and computer vision processing.
Accessibility Tree Queries via find_element
Each platform implements a find_element RPC that traverses the native UI automation tree. The Windows handler in libs/python/computer-server/computer_server/handlers/windows.py demonstrates this pattern:
# libs/python/computer-server/computer_server/handlers/windows.py
async def find_element(self, role=None, title=None, value=None):
# Walk UIAutomation tree matching criteria
element = self.automation.GetRootElement()
condition = {
"role": role,
"title": title,
"value": value
}
found = self._search_tree(element, condition)
return {
"success": True,
"element": {
"role": found.ControlType,
"title": found.Name,
"position": {"x": found.BoundingRectangle.x, "y": found.BoundingRectangle.y},
"size": {"width": found.BoundingRectangle.width, "height": found.BoundingRectangle.height}
}
}
This method returns structured JSON describing the element's role, title, and precise screen coordinates, enabling deterministic mouse interactions without relying on visual heuristics.
Computer Vision and OCR Pipeline
When accessibility trees are unavailable or insufficient, Cua falls back to the vision-based detection system in libs/python/som/som/detect.py. This pipeline combines YOLOv8 icon detection with EasyOCR text recognition to produce structured UIElement objects:
# libs/python/som/som/detect.py
class OmniParser:
def process_image(self, image: Image.Image):
# Icon detection using YOLO
icon_detections = self.detector.detect_icons(
image=image,
box_threshold=0.3,
iou_threshold=0.1
)
elements = [
IconElement(
id=i+1,
bbox=BoundingBox.from_coords(det["bbox"]),
confidence=det["confidence"]
) for i, det in enumerate(icon_detections)
]
# Text detection using EasyOCR
text_detections = self.ocr.detect_text(
image=image,
confidence_threshold=0.5
)
text_elements = [
TextElement(
id=len(elements)+i+1,
bbox=BoundingBox.from_coords(det["bbox"]),
content=det["content"]
) for i, det in enumerate(text_detections)
]
return elements + text_elements
The detector returns typed subclasses—IconElement for UI controls and TextElement for readable strings—each containing bounding boxes, confidence scores, and (for text) extracted content.
Merging Detection Results
Cua applies collision filtering to prevent duplicate detections when OCR text overlaps with icon boundaries. Icons that are overlapped by OCR boxes are pruned, yielding a clean hierarchy of distinct UIElement objects that can be safely used for click coordinates or verification logic.
Integrating Verification and Detection in Automation Scripts
A typical reliable automation workflow combines these mechanisms to handle failures gracefully:
import asyncio
from pathlib import Path
from computer.interface import ComputerInterface
async def reliable_automation_workflow():
interface = ComputerInterface()
# 1. Launch application
await interface.execute("open -a 'MyApp'")
await asyncio.sleep(2)
# 2. Verify screen state via Bedrock
screenshot = await interface.screenshot()
await interface.save_screenshot(screenshot, "launch.png")
passed, reason = await _verify_screenshot(
Path("launch.png"),
"Verify MyApp main window is visible with toolbar"
)
if not passed:
raise RuntimeError(f"Screen verification failed: {reason}")
# 3. Locate element (accessibility first, vision fallback)
result = await interface.find_element(title="Start")
if result["success"]:
target = result["element"]["position"]
else:
# Fallback to vision-based detection
elements = interface.omni_parser.process_image(Image.open("launch.png"))
start_btn = next(e for e in elements if isinstance(e, IconElement) and "start" in str(e.id).lower())
target = {"x": start_btn.bbox.center_x, "y": start_btn.bbox.center_y}
# 4. Execute interaction
await interface.move(target["x"], target["y"])
await interface.click()
# 5. Verify result state
post_click_screenshot = await interface.screenshot()
passed, _ = await _verify_screenshot(
post_click_screenshot,
"Verify dialog box appeared after click"
)
return passed
This pattern ensures that if the accessibility query fails, the vision pipeline steps in, and critical interactions are always followed by screen state verification to confirm the expected UI change.
Summary
- Cua verifies screen states by sending screenshots to AWS Bedrock Claude-Haiku, which returns structured PASS/FAIL responses parsed by
_verify_screenshotincreator_agent.py. - Element detection operates via dual paths: native accessibility tree queries through
find_elementRPCs (such as inwindows.py), and computer vision processing via theOmniParserclass using YOLO and EasyOCR. - Collision filtering prevents duplicate detections by pruning icons that overlap OCR text regions, producing clean
UIElementhierarchies suitable for precise automation. - Integration patterns allow accessibility queries to fail over to vision-based detection, followed by explicit screen verification to confirm UI state changes across heterogeneous environments.
Frequently Asked Questions
How does Cua verify that a UI element is actually visible before clicking?
Cua captures a screenshot and sends it to AWS Bedrock Claude-Haiku through the _verify_screenshot function in creator_agent.py. The model analyzes the image and returns a response starting with PASS if the expected UI is present, or FAIL if not. This verification triggers before critical interactions and after state changes to ensure the automation targets the correct screen state, with the orchestration loop automatically retrying on failure.
What happens if the accessibility tree query fails to find an element?
When find_element returns a failure response (e.g., from libs/python/computer-server/computer_server/handlers/windows.py), Cua automatically falls back to the vision-based detection pipeline. The OmniParser class processes the current screenshot using YOLO icon detection and EasyOCR to locate the element by visual appearance rather than accessibility metadata, ensuring the automation can continue even when native UI trees are unavailable.
Which AI model does Cua use for screen state verification?
Cua uses Claude-3 Haiku via AWS Bedrock for screen verification tasks. This model receives base64-encoded screenshots along with text prompts asking it to verify specific UI conditions, leveraging its multimodal capabilities to understand visual interface states across different operating systems and window managers.
Can Cua's vision-based detection work across different operating systems?
Yes. The computer vision pipeline in libs/python/som/som/detect.py operates purely on image inputs without requiring OS-specific accessibility APIs. While the find_element RPC requires platform-specific implementations for Windows, macOS, Linux, or Android, the YOLO and OCR-based detection works universally across any operating system that can provide screenshot data, making it ideal for heterogeneous or remote sandbox environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →