How to Create a Custom Skill for CuaBot to Automate Specific GUI Tasks

To create a custom skill for CuaBot, place a SKILL.md file with YAML front-matter and a trajectory folder containing video and JSON metadata in ~/.cua/skills/<skill-name>/, or use the cua skills record CLI command to automate the capture and LLM captioning process.

CuaBot uses skills—recorded GUI demonstrations enriched with LLM-generated captions—as reusable prompts for agents. Creating a custom skill allows you to teach the bot specific workflows once and reuse them across automation tasks. This guide covers both the automated recording pipeline and manual skill creation based on the trycua/cua source code.

Understanding the CuaBot Skill Structure

A skill is fundamentally a directory under ~/.cua/skills/ containing structured artifacts that combine visual demonstrations with textual instructions.

Required Directory Layout

Every skill follows this exact structure:


~/.cua/skills/<skill-name>/
├─ SKILL.md
└─ trajectory/
   ├─ <skill-name>.mp4
   ├─ events.json
   └─ trajectory.json

File Specifications

  • SKILL.md: Markdown file with YAML front-matter defining name and description, followed by an agent prompt and step descriptions.
  • trajectory/*.mp4: The raw video recording of the GUI demonstration.
  • trajectory/events.json: Low-level VNC events captured during recording.
  • trajectory/trajectory.json: LLM-generated captions and metadata for each step.

Recording a Custom Skill via CLI

The fastest way to create a custom skill is using the cua skills record command, which orchestrates a complete pipeline from screen capture to caption generation.

The Recording Pipeline

According to the implementation in libs/python/cua-cli/cua_cli/commands/skills.py, the recording process executes five distinct phases:

  1. WebSocket Server Initialization: The _record_skill_async function spawns a temporary WebSocket server using websockets.serve (lines 74-78) to receive binary recording data from the sandbox.

  2. VNC Viewer Launch: The CLI constructs a VNC viewer URL with autorecord=true and the WebSocket endpoint (lines 107-118), automatically opening the viewer with recording enabled.

  3. Binary Data Collection: The WebSocket handler accumulates bytes into a recording_data buffer (lines 55-60), receiving a JSON header followed by MP4 video payload.

  4. Processing and Captioning: The _process_recording function (lines 100-114) splits the JSON header from the video, extracts frames using ffmpeg, and invokes _caption_step (lines 108-166) to generate step descriptions via OpenAI or Anthropic APIs.

  5. Artifact Generation: Finally, the pipeline writes trajectory.json with structured captions and renders SKILL.md (lines 446-491) containing the full agent prompt.

CLI Recording Example

Execute the following to record a new skill against a running sandbox:

cua skills record --sandbox demo-sandbox \
    --provider openai \
    --model gpt-4o-mini \
    --name open-settings \
    --description "Open the Settings window and enable Dark Mode"

This command spins up the local WebSocket server, opens the VNC viewer with auto-record enabled, and upon completion stores the processed skill under ~/.cua/skills/open-settings/.

Creating a Custom Skill Manually

For scenarios where you have existing video footage or need precise control over captions, manual skill creation bypasses the automated recording pipeline.

Manual Directory Setup

Create the skill structure without using the recorder:


~/.cua/skills/my-custom-skill/
├─ SKILL.md
└─ trajectory/
   ├─ my-custom-skill.mp4
   ├─ events.json          # Can be empty dict {}

   └─ trajectory.json

SKILL.md Format Requirements

The SKILL.md must begin with YAML front-matter:

---
name: my-custom-skill
description: Automates opening a settings dialog and toggling a checkbox
---

Following the front-matter, include a Steps section describing each action and an Agent Prompt section instructing the bot how to execute the task.

trajectory.json Schema

The trajectory.json file must contain an array of step objects with captions:

{
  "events": [],
  "trajectory": [
    {
      "step_idx": 1,
      "caption": {
        "observation": "Dashboard with stale data",
        "think": "Need fresh data",
        "action": "click",
        "expectation": "Data refreshes"
      },
      "raw_event": {}
    }
  ],
  "metadata": {
    "task_description": "Click the Refresh button",
    "total_steps": 1,
    "created_at": "2024-01-15T10:30:00"
  }
}

Programmatic Skill Creation

Generate a skill programmatically using Python:

import json
import pathlib
import datetime

skill_dir = pathlib.Path.home() / ".cua" / "skills" / "my-custom-skill"
(skill_dir / "trajectory").mkdir(parents=True, exist_ok=True)

# Write SKILL.md with front-matter

(skill_dir / "SKILL.md").write_text("""---
name: my-custom-skill
description: Click the Refresh button on the dashboard
---

# my-custom-skill

Click the Refresh button on the dashboard.

## Steps

### Step 1: Click Refresh

**Context:** The dashboard shows stale data.  
**Intent:** User wants the latest information.  
**Expected Result:** Table updates with fresh rows.

## Agent Prompt

You have been shown a demonstration of how to perform this task:
Click the Refresh button on the dashboard.
...
""")

# Write trajectory.json

trajectory = {
    "events": [],
    "trajectory": [
        {
            "step_idx": 1,
            "caption": {
                "observation": "Dashboard with stale data",
                "think": "Need fresh data",
                "action": "click",
                "expectation": "Data refreshes"
            },
            "raw_event": {}
        }
    ],
    "metadata": {
        "task_description": "Click the Refresh button on the dashboard",
        "total_steps": 1,
        "created_at": datetime.datetime.now().isoformat()
    }
}

(skill_dir / "trajectory" / "trajectory.json").write_text(
    json.dumps(trajectory, indent=2)
)

Managing Existing Skills

Once created, skills can be listed, inspected, and replayed using the CLI commands implemented in libs/python/cua-cli/cua_cli/commands/skills.py.

Listing Available Skills

View all skills in a human-readable table:

cua skills list

For programmatic integration, output JSON:

cua skills list --json

The cmd_list function (lines 199-229) walks the SKILLS_DIR and extracts metadata via _get_skill_info.

Reading Skill Definitions

Display the full SKILL.md content:

cua skills read open-settings

Retrieve structured data including the trajectory:

cua skills read open-settings --format json

The cmd_read implementation (lines 301-328) parses the markdown front-matter and merges the trajectory.json data.

Replaying Recordings

Open the original MP4 recording in your default browser:

cua skills replay open-settings

This invokes cmd_replay (lines 42-52), which calls webbrowser.open on the local video file.

Summary

  • CuaBot skills are reusable GUI demonstrations stored in ~/.cua/skills/<skill-name>/ combining video, JSON metadata, and markdown prompts.
  • Automated recording via cua skills record spins up a WebSocket server (_record_skill_async), captures VNC streams, and uses LLM captioning (_caption_step) to generate trajectory.json and SKILL.md.
  • Manual creation requires placing SKILL.md with YAML front-matter and a trajectory/ folder containing MP4 video and properly structured trajectory.json.
  • Skill management commands (list, read, replay) are implemented in libs/python/cua-cli/cua_cli/commands/skills.py and provide both human-readable and JSON outputs.

Frequently Asked Questions

What file format does the SKILL.md front-matter use?

SKILL.md uses YAML front-matter delimited by triple dashes (---). The front-matter must include name and description keys, followed by markdown content containing Steps and Agent Prompt sections.

Can I create a custom skill without using the recording WebSocket server?

Yes. You can manually create the directory structure under ~/.cua/skills/ and populate SKILL.md and trajectory.json yourself. The events.json file can be an empty dictionary {} if you do not have low-level VNC event data.

Which LLM providers does the automated captioning support?

The _caption_step function in libs/python/cua-cli/cua_cli/commands/skills.py (lines 108-166) supports both OpenAI and Anthropic APIs. You specify the provider and model via the --provider and --model flags when running cua skills record.

How does the CLI locate the skills directory?

The CLI uses the SKILLS_DIR constant, which resolves to ~/.cua/skills/ in the user's home directory. All skill commands, including cmd_list and cmd_read, reference this path when enumerating or reading skill definitions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →