# How to Create a Custom Skill for CuaBot to Automate Specific GUI Tasks

> Learn to create a custom skill for CuaBot to automate GUI tasks. Follow our guide or use the CLI to record and train your unique automation workflows.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: how-to-guide
- Published: 2026-04-27

---

**To create a custom skill for CuaBot, place a [`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md) file with YAML front-matter and a `trajectory` folder containing video and JSON metadata in `~/.cua/skills/<skill-name>/`, or use the `cua skills record` CLI command to automate the capture and LLM captioning process.**

CuaBot uses **skills**—recorded GUI demonstrations enriched with LLM-generated captions—as reusable prompts for agents. Creating a custom skill allows you to teach the bot specific workflows once and reuse them across automation tasks. This guide covers both the automated recording pipeline and manual skill creation based on the `trycua/cua` source code.

## Understanding the CuaBot Skill Structure

A skill is fundamentally a directory under `~/.cua/skills/` containing structured artifacts that combine visual demonstrations with textual instructions.

### Required Directory Layout

Every skill follows this exact structure:

```

~/.cua/skills/<skill-name>/
├─ SKILL.md
└─ trajectory/
   ├─ <skill-name>.mp4
   ├─ events.json
   └─ trajectory.json

```

### File Specifications

- **[`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md)**: Markdown file with YAML front-matter defining `name` and `description`, followed by an agent prompt and step descriptions.
- **`trajectory/*.mp4`**: The raw video recording of the GUI demonstration.
- **[`trajectory/events.json`](https://github.com/trycua/cua/blob/main/trajectory/events.json)**: Low-level VNC events captured during recording.
- **[`trajectory/trajectory.json`](https://github.com/trycua/cua/blob/main/trajectory/trajectory.json)**: LLM-generated captions and metadata for each step.

## Recording a Custom Skill via CLI

The fastest way to create a custom skill is using the `cua skills record` command, which orchestrates a complete pipeline from screen capture to caption generation.

### The Recording Pipeline

According to the implementation in [`libs/python/cua-cli/cua_cli/commands/skills.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-cli/cua_cli/commands/skills.py), the recording process executes five distinct phases:

1. **WebSocket Server Initialization**: The `_record_skill_async` function spawns a temporary WebSocket server using `websockets.serve` (lines 74-78) to receive binary recording data from the sandbox.

2. **VNC Viewer Launch**: The CLI constructs a VNC viewer URL with `autorecord=true` and the WebSocket endpoint (lines 107-118), automatically opening the viewer with recording enabled.

3. **Binary Data Collection**: The WebSocket handler accumulates bytes into a `recording_data` buffer (lines 55-60), receiving a JSON header followed by MP4 video payload.

4. **Processing and Captioning**: The `_process_recording` function (lines 100-114) splits the JSON header from the video, extracts frames using `ffmpeg`, and invokes `_caption_step` (lines 108-166) to generate step descriptions via OpenAI or Anthropic APIs.

5. **Artifact Generation**: Finally, the pipeline writes [`trajectory.json`](https://github.com/trycua/cua/blob/main/trajectory.json) with structured captions and renders [`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md) (lines 446-491) containing the full agent prompt.

### CLI Recording Example

Execute the following to record a new skill against a running sandbox:

```bash
cua skills record --sandbox demo-sandbox \
    --provider openai \
    --model gpt-4o-mini \
    --name open-settings \
    --description "Open the Settings window and enable Dark Mode"

```

This command spins up the local WebSocket server, opens the VNC viewer with auto-record enabled, and upon completion stores the processed skill under `~/.cua/skills/open-settings/`.

## Creating a Custom Skill Manually

For scenarios where you have existing video footage or need precise control over captions, manual skill creation bypasses the automated recording pipeline.

### Manual Directory Setup

Create the skill structure without using the recorder:

```

~/.cua/skills/my-custom-skill/
├─ SKILL.md
└─ trajectory/
   ├─ my-custom-skill.mp4
   ├─ events.json          # Can be empty dict {}

   └─ trajectory.json

```

### SKILL.md Format Requirements

The [`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md) must begin with YAML front-matter:

```yaml
---
name: my-custom-skill
description: Automates opening a settings dialog and toggling a checkbox
---

```

Following the front-matter, include a **Steps** section describing each action and an **Agent Prompt** section instructing the bot how to execute the task.

### trajectory.json Schema

The [`trajectory.json`](https://github.com/trycua/cua/blob/main/trajectory.json) file must contain an array of step objects with captions:

```json
{
  "events": [],
  "trajectory": [
    {
      "step_idx": 1,
      "caption": {
        "observation": "Dashboard with stale data",
        "think": "Need fresh data",
        "action": "click",
        "expectation": "Data refreshes"
      },
      "raw_event": {}
    }
  ],
  "metadata": {
    "task_description": "Click the Refresh button",
    "total_steps": 1,
    "created_at": "2024-01-15T10:30:00"
  }
}

```

### Programmatic Skill Creation

Generate a skill programmatically using Python:

```python
import json
import pathlib
import datetime

skill_dir = pathlib.Path.home() / ".cua" / "skills" / "my-custom-skill"
(skill_dir / "trajectory").mkdir(parents=True, exist_ok=True)

# Write SKILL.md with front-matter

(skill_dir / "SKILL.md").write_text("""---
name: my-custom-skill
description: Click the Refresh button on the dashboard
---

# my-custom-skill

Click the Refresh button on the dashboard.

## Steps

### Step 1: Click Refresh

**Context:** The dashboard shows stale data.  
**Intent:** User wants the latest information.  
**Expected Result:** Table updates with fresh rows.

## Agent Prompt

You have been shown a demonstration of how to perform this task:
Click the Refresh button on the dashboard.
...
""")

# Write trajectory.json

trajectory = {
    "events": [],
    "trajectory": [
        {
            "step_idx": 1,
            "caption": {
                "observation": "Dashboard with stale data",
                "think": "Need fresh data",
                "action": "click",
                "expectation": "Data refreshes"
            },
            "raw_event": {}
        }
    ],
    "metadata": {
        "task_description": "Click the Refresh button on the dashboard",
        "total_steps": 1,
        "created_at": datetime.datetime.now().isoformat()
    }
}

(skill_dir / "trajectory" / "trajectory.json").write_text(
    json.dumps(trajectory, indent=2)
)

```

## Managing Existing Skills

Once created, skills can be listed, inspected, and replayed using the CLI commands implemented in [`libs/python/cua-cli/cua_cli/commands/skills.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-cli/cua_cli/commands/skills.py).

### Listing Available Skills

View all skills in a human-readable table:

```bash
cua skills list

```

For programmatic integration, output JSON:

```bash
cua skills list --json

```

The `cmd_list` function (lines 199-229) walks the `SKILLS_DIR` and extracts metadata via `_get_skill_info`.

### Reading Skill Definitions

Display the full [`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md) content:

```bash
cua skills read open-settings

```

Retrieve structured data including the trajectory:

```bash
cua skills read open-settings --format json

```

The `cmd_read` implementation (lines 301-328) parses the markdown front-matter and merges the [`trajectory.json`](https://github.com/trycua/cua/blob/main/trajectory.json) data.

### Replaying Recordings

Open the original MP4 recording in your default browser:

```bash
cua skills replay open-settings

```

This invokes `cmd_replay` (lines 42-52), which calls `webbrowser.open` on the local video file.

## Summary

- **CuaBot skills** are reusable GUI demonstrations stored in `~/.cua/skills/<skill-name>/` combining video, JSON metadata, and markdown prompts.
- **Automated recording** via `cua skills record` spins up a WebSocket server (`_record_skill_async`), captures VNC streams, and uses LLM captioning (`_caption_step`) to generate [`trajectory.json`](https://github.com/trycua/cua/blob/main/trajectory.json) and [`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md).
- **Manual creation** requires placing [`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md) with YAML front-matter and a `trajectory/` folder containing MP4 video and properly structured [`trajectory.json`](https://github.com/trycua/cua/blob/main/trajectory.json).
- **Skill management** commands (`list`, `read`, `replay`) are implemented in [`libs/python/cua-cli/cua_cli/commands/skills.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-cli/cua_cli/commands/skills.py) and provide both human-readable and JSON outputs.

## Frequently Asked Questions

### What file format does the SKILL.md front-matter use?

[`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md) uses **YAML front-matter** delimited by triple dashes (`---`). The front-matter must include `name` and `description` keys, followed by markdown content containing **Steps** and **Agent Prompt** sections.

### Can I create a custom skill without using the recording WebSocket server?

Yes. You can manually create the directory structure under `~/.cua/skills/` and populate [`SKILL.md`](https://github.com/trycua/cua/blob/main/SKILL.md) and [`trajectory.json`](https://github.com/trycua/cua/blob/main/trajectory.json) yourself. The [`events.json`](https://github.com/trycua/cua/blob/main/events.json) file can be an empty dictionary `{}` if you do not have low-level VNC event data.

### Which LLM providers does the automated captioning support?

The `_caption_step` function in [`libs/python/cua-cli/cua_cli/commands/skills.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-cli/cua_cli/commands/skills.py) (lines 108-166) supports both **OpenAI** and **Anthropic** APIs. You specify the provider and model via the `--provider` and `--model` flags when running `cua skills record`.

### How does the CLI locate the skills directory?

The CLI uses the `SKILLS_DIR` constant, which resolves to `~/.cua/skills/` in the user's home directory. All skill commands, including `cmd_list` and `cmd_read`, reference this path when enumerating or reading skill definitions.