How to Create a Custom Skill for CuaBot to Automate Specific GUI Tasks
To create a custom skill for CuaBot, place a SKILL.md file with YAML front-matter and a trajectory folder containing video and JSON metadata in ~/.cua/skills/<skill-name>/, or use the cua skills record CLI command to automate the capture and LLM captioning process.
CuaBot uses skills—recorded GUI demonstrations enriched with LLM-generated captions—as reusable prompts for agents. Creating a custom skill allows you to teach the bot specific workflows once and reuse them across automation tasks. This guide covers both the automated recording pipeline and manual skill creation based on the trycua/cua source code.
Understanding the CuaBot Skill Structure
A skill is fundamentally a directory under ~/.cua/skills/ containing structured artifacts that combine visual demonstrations with textual instructions.
Required Directory Layout
Every skill follows this exact structure:
~/.cua/skills/<skill-name>/
├─ SKILL.md
└─ trajectory/
├─ <skill-name>.mp4
├─ events.json
└─ trajectory.json
File Specifications
SKILL.md: Markdown file with YAML front-matter definingnameanddescription, followed by an agent prompt and step descriptions.trajectory/*.mp4: The raw video recording of the GUI demonstration.trajectory/events.json: Low-level VNC events captured during recording.trajectory/trajectory.json: LLM-generated captions and metadata for each step.
Recording a Custom Skill via CLI
The fastest way to create a custom skill is using the cua skills record command, which orchestrates a complete pipeline from screen capture to caption generation.
The Recording Pipeline
According to the implementation in libs/python/cua-cli/cua_cli/commands/skills.py, the recording process executes five distinct phases:
-
WebSocket Server Initialization: The
_record_skill_asyncfunction spawns a temporary WebSocket server usingwebsockets.serve(lines 74-78) to receive binary recording data from the sandbox. -
VNC Viewer Launch: The CLI constructs a VNC viewer URL with
autorecord=trueand the WebSocket endpoint (lines 107-118), automatically opening the viewer with recording enabled. -
Binary Data Collection: The WebSocket handler accumulates bytes into a
recording_databuffer (lines 55-60), receiving a JSON header followed by MP4 video payload. -
Processing and Captioning: The
_process_recordingfunction (lines 100-114) splits the JSON header from the video, extracts frames usingffmpeg, and invokes_caption_step(lines 108-166) to generate step descriptions via OpenAI or Anthropic APIs. -
Artifact Generation: Finally, the pipeline writes
trajectory.jsonwith structured captions and rendersSKILL.md(lines 446-491) containing the full agent prompt.
CLI Recording Example
Execute the following to record a new skill against a running sandbox:
cua skills record --sandbox demo-sandbox \
--provider openai \
--model gpt-4o-mini \
--name open-settings \
--description "Open the Settings window and enable Dark Mode"
This command spins up the local WebSocket server, opens the VNC viewer with auto-record enabled, and upon completion stores the processed skill under ~/.cua/skills/open-settings/.
Creating a Custom Skill Manually
For scenarios where you have existing video footage or need precise control over captions, manual skill creation bypasses the automated recording pipeline.
Manual Directory Setup
Create the skill structure without using the recorder:
~/.cua/skills/my-custom-skill/
├─ SKILL.md
└─ trajectory/
├─ my-custom-skill.mp4
├─ events.json # Can be empty dict {}
└─ trajectory.json
SKILL.md Format Requirements
The SKILL.md must begin with YAML front-matter:
---
name: my-custom-skill
description: Automates opening a settings dialog and toggling a checkbox
---
Following the front-matter, include a Steps section describing each action and an Agent Prompt section instructing the bot how to execute the task.
trajectory.json Schema
The trajectory.json file must contain an array of step objects with captions:
{
"events": [],
"trajectory": [
{
"step_idx": 1,
"caption": {
"observation": "Dashboard with stale data",
"think": "Need fresh data",
"action": "click",
"expectation": "Data refreshes"
},
"raw_event": {}
}
],
"metadata": {
"task_description": "Click the Refresh button",
"total_steps": 1,
"created_at": "2024-01-15T10:30:00"
}
}
Programmatic Skill Creation
Generate a skill programmatically using Python:
import json
import pathlib
import datetime
skill_dir = pathlib.Path.home() / ".cua" / "skills" / "my-custom-skill"
(skill_dir / "trajectory").mkdir(parents=True, exist_ok=True)
# Write SKILL.md with front-matter
(skill_dir / "SKILL.md").write_text("""---
name: my-custom-skill
description: Click the Refresh button on the dashboard
---
# my-custom-skill
Click the Refresh button on the dashboard.
## Steps
### Step 1: Click Refresh
**Context:** The dashboard shows stale data.
**Intent:** User wants the latest information.
**Expected Result:** Table updates with fresh rows.
## Agent Prompt
You have been shown a demonstration of how to perform this task:
Click the Refresh button on the dashboard.
...
""")
# Write trajectory.json
trajectory = {
"events": [],
"trajectory": [
{
"step_idx": 1,
"caption": {
"observation": "Dashboard with stale data",
"think": "Need fresh data",
"action": "click",
"expectation": "Data refreshes"
},
"raw_event": {}
}
],
"metadata": {
"task_description": "Click the Refresh button on the dashboard",
"total_steps": 1,
"created_at": datetime.datetime.now().isoformat()
}
}
(skill_dir / "trajectory" / "trajectory.json").write_text(
json.dumps(trajectory, indent=2)
)
Managing Existing Skills
Once created, skills can be listed, inspected, and replayed using the CLI commands implemented in libs/python/cua-cli/cua_cli/commands/skills.py.
Listing Available Skills
View all skills in a human-readable table:
cua skills list
For programmatic integration, output JSON:
cua skills list --json
The cmd_list function (lines 199-229) walks the SKILLS_DIR and extracts metadata via _get_skill_info.
Reading Skill Definitions
Display the full SKILL.md content:
cua skills read open-settings
Retrieve structured data including the trajectory:
cua skills read open-settings --format json
The cmd_read implementation (lines 301-328) parses the markdown front-matter and merges the trajectory.json data.
Replaying Recordings
Open the original MP4 recording in your default browser:
cua skills replay open-settings
This invokes cmd_replay (lines 42-52), which calls webbrowser.open on the local video file.
Summary
- CuaBot skills are reusable GUI demonstrations stored in
~/.cua/skills/<skill-name>/combining video, JSON metadata, and markdown prompts. - Automated recording via
cua skills recordspins up a WebSocket server (_record_skill_async), captures VNC streams, and uses LLM captioning (_caption_step) to generatetrajectory.jsonandSKILL.md. - Manual creation requires placing
SKILL.mdwith YAML front-matter and atrajectory/folder containing MP4 video and properly structuredtrajectory.json. - Skill management commands (
list,read,replay) are implemented inlibs/python/cua-cli/cua_cli/commands/skills.pyand provide both human-readable and JSON outputs.
Frequently Asked Questions
What file format does the SKILL.md front-matter use?
SKILL.md uses YAML front-matter delimited by triple dashes (---). The front-matter must include name and description keys, followed by markdown content containing Steps and Agent Prompt sections.
Can I create a custom skill without using the recording WebSocket server?
Yes. You can manually create the directory structure under ~/.cua/skills/ and populate SKILL.md and trajectory.json yourself. The events.json file can be an empty dictionary {} if you do not have low-level VNC event data.
Which LLM providers does the automated captioning support?
The _caption_step function in libs/python/cua-cli/cua_cli/commands/skills.py (lines 108-166) supports both OpenAI and Anthropic APIs. You specify the provider and model via the --provider and --model flags when running cua skills record.
How does the CLI locate the skills directory?
The CLI uses the SKILLS_DIR constant, which resolves to ~/.cua/skills/ in the user's home directory. All skill commands, including cmd_list and cmd_read, reference this path when enumerating or reading skill definitions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →