# How Cua Handles Screen State Verification and Element Detection for Reliable Automation

> Learn how Cua achieves reliable automation by combining AI screenshot validation with dual-path element discovery using accessibility trees and computer vision OCR.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: how-to-guide
- Published: 2026-04-27

---

**Cua combines AI-driven screenshot validation with dual-path element discovery—accessibility tree queries and computer vision OCR—to ensure every automated action hits the correct target.**

The trycua/cua repository implements a robust automation framework that guarantees reliability through layered verification mechanisms. By integrating AWS Bedrock-based screen state verification with deterministic element detection strategies, Cua eliminates the brittleness of traditional pixel-matching approaches. This article examines the specific source code implementations that enable reliable cross-platform UI automation.

## Screen State Verification with AWS Bedrock

Cua validates that expected UI elements are actually visible before proceeding with critical automation steps. This verification mechanism relies on multimodal AI analysis rather than fragile pixel comparisons.

### The Verification Pipeline

After capturing a screenshot via `ComputerInterface.screenshot`, the system uploads the image to AWS Bedrock Claude-Haiku for analysis. The `_verify_screenshot` function in [`libs/python/cua-sandbox-apps/cua_sandbox_apps/pipeline/creator_agent.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-sandbox-apps/cua_sandbox_apps/pipeline/creator_agent.py) constructs a multipart prompt containing the base64-encoded image and specific verification instructions.

The model must respond with a single line starting with **PASS** or **FAIL**, followed by explanatory text. This strict formatting allows the orchestration loop to parse results automatically and retry or abort when criteria are not met.

```python

# libs/python/cua-sandbox-apps/cua_sandbox_apps/pipeline/creator_agent.py

def _verify_screenshot(image_path: Path, verification_prompt: str) -> tuple[bool, str]:
    # Upload image to Bedrock with prompt

    response = bedrock_client.invoke_model(
        model_id="anthropic.claude-3-haiku",
        body=json.dumps({
            "messages": [{
                "role": "user",
                "content": [
                    {"type": "text", "text": verification_prompt},
                    {"type": "image", "source": {"type": "base64", "data": image_base64}}
                ]
            }]
        })
    )
    # Parse PASS/FAIL response

    output = json.loads(response['body'].read())['content'][0]['text']
    passed = output.strip().startswith("PASS")
    return passed, output

```

The orchestration code in the same file treats `PASS` as success; otherwise, it returns *CRITERIA NOT MET* and triggers an automatic retry of the step.

## Element Detection Strategies

Cua discovers interactable UI elements through two complementary approaches: native accessibility tree queries and computer vision processing.

### Accessibility Tree Queries via find_element

Each platform implements a `find_element` RPC that traverses the native UI automation tree. The Windows handler in [`libs/python/computer-server/computer_server/handlers/windows.py`](https://github.com/trycua/cua/blob/main/libs/python/computer-server/computer_server/handlers/windows.py) demonstrates this pattern:

```python

# libs/python/computer-server/computer_server/handlers/windows.py

async def find_element(self, role=None, title=None, value=None):
    # Walk UIAutomation tree matching criteria

    element = self.automation.GetRootElement()
    condition = {
        "role": role,
        "title": title,
        "value": value
    }
    found = self._search_tree(element, condition)
    return {
        "success": True,
        "element": {
            "role": found.ControlType,
            "title": found.Name,
            "position": {"x": found.BoundingRectangle.x, "y": found.BoundingRectangle.y},
            "size": {"width": found.BoundingRectangle.width, "height": found.BoundingRectangle.height}
        }
    }

```

This method returns structured JSON describing the element's role, title, and precise screen coordinates, enabling deterministic mouse interactions without relying on visual heuristics.

### Computer Vision and OCR Pipeline

When accessibility trees are unavailable or insufficient, Cua falls back to the vision-based detection system in [`libs/python/som/som/detect.py`](https://github.com/trycua/cua/blob/main/libs/python/som/som/detect.py). This pipeline combines YOLOv8 icon detection with EasyOCR text recognition to produce structured `UIElement` objects:

```python

# libs/python/som/som/detect.py

class OmniParser:
    def process_image(self, image: Image.Image):
        # Icon detection using YOLO

        icon_detections = self.detector.detect_icons(
            image=image, 
            box_threshold=0.3, 
            iou_threshold=0.1
        )
        elements = [
            IconElement(
                id=i+1, 
                bbox=BoundingBox.from_coords(det["bbox"]), 
                confidence=det["confidence"]
            ) for i, det in enumerate(icon_detections)
        ]
        
        # Text detection using EasyOCR

        text_detections = self.ocr.detect_text(
            image=image, 
            confidence_threshold=0.5
        )
        text_elements = [
            TextElement(
                id=len(elements)+i+1,
                bbox=BoundingBox.from_coords(det["bbox"]),
                content=det["content"]
            ) for i, det in enumerate(text_detections)
        ]
        
        return elements + text_elements

```

The detector returns typed subclasses—**IconElement** for UI controls and **TextElement** for readable strings—each containing bounding boxes, confidence scores, and (for text) extracted content.

### Merging Detection Results

Cua applies collision filtering to prevent duplicate detections when OCR text overlaps with icon boundaries. Icons that are overlapped by OCR boxes are pruned, yielding a clean hierarchy of distinct `UIElement` objects that can be safely used for click coordinates or verification logic.

## Integrating Verification and Detection in Automation Scripts

A typical reliable automation workflow combines these mechanisms to handle failures gracefully:

```python
import asyncio
from pathlib import Path
from computer.interface import ComputerInterface

async def reliable_automation_workflow():
    interface = ComputerInterface()
    
    # 1. Launch application

    await interface.execute("open -a 'MyApp'")
    await asyncio.sleep(2)
    
    # 2. Verify screen state via Bedrock

    screenshot = await interface.screenshot()
    await interface.save_screenshot(screenshot, "launch.png")
    
    passed, reason = await _verify_screenshot(
        Path("launch.png"), 
        "Verify MyApp main window is visible with toolbar"
    )
    if not passed:
        raise RuntimeError(f"Screen verification failed: {reason}")
    
    # 3. Locate element (accessibility first, vision fallback)

    result = await interface.find_element(title="Start")
    if result["success"]:
        target = result["element"]["position"]
    else:
        # Fallback to vision-based detection

        elements = interface.omni_parser.process_image(Image.open("launch.png"))
        start_btn = next(e for e in elements if isinstance(e, IconElement) and "start" in str(e.id).lower())
        target = {"x": start_btn.bbox.center_x, "y": start_btn.bbox.center_y}
    
    # 4. Execute interaction

    await interface.move(target["x"], target["y"])
    await interface.click()
    
    # 5. Verify result state

    post_click_screenshot = await interface.screenshot()
    passed, _ = await _verify_screenshot(
        post_click_screenshot, 
        "Verify dialog box appeared after click"
    )
    return passed

```

This pattern ensures that if the accessibility query fails, the vision pipeline steps in, and critical interactions are always followed by screen state verification to confirm the expected UI change.

## Summary

- **Cua verifies screen states** by sending screenshots to AWS Bedrock Claude-Haiku, which returns structured **PASS/FAIL** responses parsed by `_verify_screenshot` in [`creator_agent.py`](https://github.com/trycua/cua/blob/main/creator_agent.py).
- **Element detection operates via dual paths**: native accessibility tree queries through `find_element` RPCs (such as in [`windows.py`](https://github.com/trycua/cua/blob/main/windows.py)), and computer vision processing via the `OmniParser` class using YOLO and EasyOCR.
- **Collision filtering** prevents duplicate detections by pruning icons that overlap OCR text regions, producing clean `UIElement` hierarchies suitable for precise automation.
- **Integration patterns** allow accessibility queries to fail over to vision-based detection, followed by explicit screen verification to confirm UI state changes across heterogeneous environments.

## Frequently Asked Questions

### How does Cua verify that a UI element is actually visible before clicking?

Cua captures a screenshot and sends it to AWS Bedrock Claude-Haiku through the `_verify_screenshot` function in [`creator_agent.py`](https://github.com/trycua/cua/blob/main/creator_agent.py). The model analyzes the image and returns a response starting with **PASS** if the expected UI is present, or **FAIL** if not. This verification triggers before critical interactions and after state changes to ensure the automation targets the correct screen state, with the orchestration loop automatically retrying on failure.

### What happens if the accessibility tree query fails to find an element?

When `find_element` returns a failure response (e.g., from [`libs/python/computer-server/computer_server/handlers/windows.py`](https://github.com/trycua/cua/blob/main/libs/python/computer-server/computer_server/handlers/windows.py)), Cua automatically falls back to the vision-based detection pipeline. The `OmniParser` class processes the current screenshot using YOLO icon detection and EasyOCR to locate the element by visual appearance rather than accessibility metadata, ensuring the automation can continue even when native UI trees are unavailable.

### Which AI model does Cua use for screen state verification?

Cua uses **Claude-3 Haiku** via AWS Bedrock for screen verification tasks. This model receives base64-encoded screenshots along with text prompts asking it to verify specific UI conditions, leveraging its multimodal capabilities to understand visual interface states across different operating systems and window managers.

### Can Cua's vision-based detection work across different operating systems?

Yes. The computer vision pipeline in [`libs/python/som/som/detect.py`](https://github.com/trycua/cua/blob/main/libs/python/som/som/detect.py) operates purely on image inputs without requiring OS-specific accessibility APIs. While the `find_element` RPC requires platform-specific implementations for Windows, macOS, Linux, or Android, the YOLO and OCR-based detection works universally across any operating system that can provide screenshot data, making it ideal for heterogeneous or remote sandbox environments.