How to Debug Cua Agent Failures Using Telemetry and Tracing: A Complete Guide
Enable the PostHog telemetry client by setting CUA_TELEMETRY_ENABLED=true for aggregate crash analytics, and wrap your Computer instance with ComputerTracing to capture detailed per-session logs, screenshots, and API call sequences to local storage for step-by-step failure analysis.
Debugging autonomous agents requires visibility into both aggregate failure patterns and granular execution details. The Cua framework provides two complementary observability systems implemented in trycua/cua—a PostHog-based telemetry client for anonymous usage analytics and a local tracing system for full session replay. Together, these tools allow you to debug Cua agent failures using telemetry and tracing capabilities that capture everything from high-level crash rates to individual mouse movement recordings.
Understanding Cua's Observability Architecture
Cua ships with two distinct backends that operate independently but complement each other during debugging workflows.
Telemetry (PostHog Analytics)
The telemetry system sends anonymous usage events to a PostHog instance, enabling developers to monitor aggregate crash rates and feature adoption across deployments. Implemented in libs/python/core/cua_core/telemetry/posthog.py, the PostHogTelemetryClient class operates as a lazy-loaded singleton that queues events until the network client is ready.
The singleton pattern ensures duplicate initialization is avoided:
# libs/python/core/cua_core/telemetry/posthog.py
class PostHogTelemetryClient:
_singleton: Optional["PostHogTelemetryClient"] = None
@classmethod
def get_client(cls) -> "PostHogTelemetryClient":
if cls._singleton is None:
cls._singleton = cls()
return cls._singleton
Environment variables control enablement. The method is_telemetry_enabled reads CUA_TELEMETRY_ENABLED (defaulting to true), with backward compatibility for the deprecated CUA_TELEMETRY_DISABLED flag.
Local Session Tracing
For detailed debugging, the ComputerTracing class in libs/python/computer/computer/tracing.py records a complete per-run session including API calls, screenshots, accessibility trees, and custom metadata. Unlike telemetry, tracing writes to local storage (either a directory or zip file) that you can inspect after a failure to understand exactly what the agent did.
The tracer hooks into the Computer interface via ComputerTracingWrapper (located in libs/python/computer/computer/tracing_wrapper.py), which intercepts every async call and forwards it to record_api_call along with arguments and return values.
Configuring Telemetry for Aggregate Debugging
Telemetry is opt-in by default but can be toggled via environment variables or programmatic checks.
Enabling and Disabling Telemetry
Set the environment variable before starting your agent:
# Enable (default behavior)
export CUA_TELEMETRY_ENABLED=true
# Disable for privacy-sensitive runs
export CUA_TELEMETRY_ENABLED=false
Verify the current state and record custom events anywhere in your codebase:
from cua_core.telemetry import record_event, is_telemetry_enabled
if is_telemetry_enabled():
record_event("agent_started", {"agent": "my_agent", "version": "1.2.0"})
The record_event function (lines 48-52 in posthog.py) automatically retrieves the singleton client and queues events if the network connection is not yet available.
Recording Custom Error Events
When catching exceptions, emit structured telemetry to categorize failure types:
from cua_core.telemetry import record_event
try:
await computer.execute_action(risky_operation)
except Exception as e:
record_event(
"agent_error",
{
"error_type": type(e).__name__,
"message": str(e),
"step": "action_execution",
},
)
raise
Enabling Detailed Session Tracing
While telemetry reveals that failures occur, tracing reveals why they occur by preserving the exact sequence of actions and system states.
Basic Tracer Setup
Instantiate ComputerTracing with your Computer implementation and configure output options:
import asyncio
from computer.tracing import ComputerTracing
from my_agent import MyComputer # Your concrete Computer implementation
async def run():
computer = MyComputer()
tracer = ComputerTracing(computer)
# Start tracing – creates traces/<trace_id>/ directory
await tracer.start({
"screenshots": True,
"api_calls": True,
"metadata": True
})
try:
# Your agent logic here...
await computer.move_mouse(100, 200)
await computer.type_text("hello world")
finally:
# Stop and bundle as zip for sharing
trace_path = await tracer.stop({"format": "zip"})
print(f"Trace saved to {trace_path}")
asyncio.run(run())
The start method (lines 51-99 in tracing.py) generates a unique trace ID, creates the directory structure, and writes an initial trace_metadata.json header.
Automatic API Call Recording
When you wrap your computer with TracingComputerWrapper, all method calls are automatically intercepted and logged:
from computer.tracing_wrapper import TracingComputerWrapper
from computer.tracing import ComputerTracing
from my_agent import MyComputer
computer = MyComputer()
tracer = ComputerTracing(computer)
# Wrap the interface to enable automatic recording
wrapped = TracingComputerWrapper(computer.interface, tracer)
# All calls through wrapped.interface are traced automatically
await wrapped.interface.click(50, 50) # Records event to disk
await wrapped.interface.screenshot() # Records PNG + metadata
The wrapper's __getattr__ method (lines 38-49 in tracing_wrapper.py) dynamically creates tracing wrappers for every attribute access on the underlying interface, forwarding arguments and recording results without modifying your agent code.
Adding Custom Metadata
Enrich traces with test case identifiers or debugging context:
await tracer.add_metadata("test_case", "login_flow")
await tracer.add_metadata("user_tier", "enterprise")
The add_metadata method (lines 210-230 in tracing.py) appends key-value pairs to the trace metadata, making it easier to filter and categorize logs when analyzing multiple runs.
Manual Screenshot Capture
Capture specific moments for visual regression analysis:
await tracer._take_screenshot("after_login_button_click")
This writes both a PNG file and an accompanying JSON event descriptor to the trace directory.
Analyzing Trace Outputs
Each trace generates multiple file types in the output directory:
- Event JSON files: Prefixed with
event_<timestamp>_, these includeapi_call,metadata,trace_start,trace_end,screenshot, andaccessibility_treeevents - PNG screenshots: Binary captures associated with screenshot events
- trace_metadata.json: Header file containing session configuration and custom metadata
When calling stop({"format": "zip"}), the entire directory structure is compressed into a single archive suitable for attaching to bug reports or uploading to debugging dashboards.
Complete Debugging Workflow Example
Combine both systems to capture both aggregate metrics and reproducible failure states:
import asyncio
from cua_core.telemetry import record_event, is_telemetry_enabled
from computer.tracing import ComputerTracing
from computer.tracing_wrapper import TracingComputerWrapper
from my_agent import MyComputer
async def debuggable_agent_run():
computer = MyComputer()
tracer = ComputerTracing(computer)
wrapped = TracingComputerWrapper(computer.interface, tracer)
# Record session start in aggregate analytics
if is_telemetry_enabled():
record_event("agent_session_start", {"task": "data_entry"})
await tracer.start({"screenshots": True, "api_calls": True})
try:
await wrapped.interface.navigate("https://example.com")
await wrapped.interface.click(100, 200)
# ... additional agent actions ...
except Exception as e:
# Capture error in telemetry
if is_telemetry_enabled():
record_event("agent_failure", {
"error": type(e).__name__,
"step": "navigation"
})
# Capture final state in trace
await tracer._take_screenshot("error_state")
await tracer.add_metadata("error_message", str(e))
raise
finally:
# Always produce trace, even on failure
trace_path = await tracer.stop({"format": "zip"})
print(f"Debug trace available at: {trace_path}")
asyncio.run(debuggable_agent_run())
Summary
- Telemetry in
cua_core.telemetry.posthogprovides anonymized aggregate analytics via thePostHogTelemetryClientsingleton, controlled by theCUA_TELEMETRY_ENABLEDenvironment variable - Tracing via
ComputerTracingandComputerTracingWrappercreates reproducible local records of every API call, screenshot, and metadata entry inlibs/python/computer/computer/tracing.py - Automatic instrumentation occurs when wrapping Computer interfaces with
TracingComputerWrapper, requiring no code changes to your agent logic - Output formats include both directory structures for live inspection and zip archives for sharing via
stop({"format": "zip"}) - Error correlation is achieved by recording telemetry events for failure categorization while preserving full execution traces for root-cause analysis
Frequently Asked Questions
How do I completely disable telemetry for privacy-sensitive environments?
Set CUA_TELEMETRY_ENABLED=false in your environment before importing the Cua libraries. The is_telemetry_enabled() function in libs/python/core/cua_core/telemetry/posthog.py checks this variable (lines 44-71) and returns False when explicitly disabled, preventing any network connections to PostHog and ensuring no usage data leaves your infrastructure.
What's the difference between telemetry and tracing in Cua?
Telemetry sends lightweight, anonymous events to a centralized PostHog instance for aggregate statistical analysis—useful for identifying which failure types occur most frequently across your user base. Tracing writes detailed, potentially sensitive execution data (screenshots, accessibility trees, exact API calls) to local disk only—essential for reproducing and debugging specific agent failures without exposing user data to external services.
How do I add custom metadata to a trace for better filtering?
Use the add_metadata(key, value) method on your ComputerTracing instance at any point during execution. According to the implementation in libs/python/computer/computer/tracing.py (lines 210-230), these key-value pairs are written to the trace events and can be queried when analyzing the trace_metadata.json file after the run completes.
Can I export traces in a shareable format for bug reports?
Yes. When calling the stop method, pass {"format": "zip"} to compress the entire trace directory into a single archive. The implementation in libs/python/computer/computer/tracing.py (lines 340-365) handles the compression automatically, making it simple to attach execution logs to GitHub issues or share with team members for collaborative debugging.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →