How AI Agents Discover and Utilize Skills from the Garden Skills Repository

AI agents discover skills by scanning manifest.json files for metadata and compatibility, then execute workflows defined in SKILL.md by parsing reference guides and running scaffold scripts while respecting checkpoint protocols for user confirmation.

The Garden Skills repository (ConardLi/garden-skills) provides a structured catalog of reusable content-creation skills that AI agents can autonomously discover and execute. Each skill is a self-contained code pack containing declarative metadata and imperative scripts that enable end-to-end workflow automation. Understanding how AI agents interact with these artifacts reveals the architecture behind autonomous skill utilization in modern development environments.

How AI Agents Discover Skills in the Repository

Parsing the Manifest File for Metadata

Agents begin discovery by scanning the repository for manifest.json files using glob patterns like **/manifest.json. Located at paths such as skills/web-video-presentation/manifest.json, this file declares the skill's identity, version, category, and compatible execution environments via the compat array (e.g., Claude, Gemini, Codex). The agent uses this metadata to determine whether it can execute the skill on its current runtime before loading any workflow definitions.

Loading the Skill Definition

After locating a manifest, the agent opens the sibling SKILL.md file to parse the human-readable workflow definition. This markdown-encoded file contains structured sections (such as ## 工作流总览 and ## Phase 1) that drive the agent's internal state machine. The agent stores the step order in memory, identifying specific phases and checkpoint markers that dictate when to pause for user input.

How AI Agents Execute Skill Workflows

Phase 1 – Content Generation

The agent initiates the workflow by parsing user input against style guides stored in references/SCRIPT-STYLE.md and references/OUTLINE-FORMAT.md. These reference files provide the constraints and formatting rules the agent uses to generate script.md and outline.md files. The agent treats these references as lookup tables for content structure, ensuring the generated artifacts match the skill's specifications.

Checkpoint Plan – Theme Selection and User Confirmation

Before proceeding to development, the agent enters the Checkpoint Plan phase. It scans every theme file under themes/*/theme.json (documented in references/THEMES.md) to recommend a matching visual theme. The agent then pauses to confirm five alignment items with the user: script, outline, theme, assets, and development mode. This hard-coded checkpoint ensures human validation before resource-intensive operations begin.

Phase 2 – Scaffold and Chapter Development

Upon user confirmation, the agent executes scripts/scaffold.sh to generate a Vite + React + TypeScript project structure. Each chapter is built according to references/CHAPTER-CRAFT.md, which encodes ten design principles and a "chapter-craft" decision tree. The agent strictly adheres to the rule that narrations.ts serves as the single source of truth for step count, ensuring consistency across generated components.

Checkpoint Audio and Phase 4 Recording

After web page generation, the agent reaches Checkpoint Audio to determine whether to synthesize narration. If confirmed, it follows the pipeline described in references/AUDIO.md and executes provider scripts located in templates/scripts/tts-providers/README.md. Finally, the agent utilizes references/RECORDING.md to capture screen recordings, supporting both automatic (?auto=1) and manual capture modes.

Self-Validation Checkpoints and State Management

The workflow implements hard-node checkpoint markers (e.g., Checkpoint Plan, Checkpoint Audio) that force the agent to emit summaries and await user approval. These checkpoints are defined in specific line ranges of SKILL.md (lines 67-84 and 100-115) and serve as guardrails that prevent autonomous execution from proceeding without validation. The agent monitors its workflow state continuously, transitioning between declarative configuration and imperative script execution only after explicit confirmation.

Practical Code Example: Discovering and Running a Skill

The following bash commands illustrate how an AI-powered tool discovers and executes the web-video-presentation skill:


# 1. Discover all skills via manifest files

manifest_files=$(opencode glob "**/manifest.json")
skill_dir=$(dirname $(echo $manifest_files | grep "web-video-presentation"))

# 2. Load manifest metadata

cat $skill_dir/manifest.json

# 3. Parse workflow phases from skill definition

opencode read "$skill_dir/SKILL.md" | grep "## Phase"

# 4. Execute scaffold script with selected theme

bash $skill_dir/scripts/scaffold.sh ./my-presentation --theme=warm-keynote

# 5. Navigate to project and generate audio (Phase 3)

cd my-presentation
npm run extract-narrations
npm run synthesize-audio

In a production AI agent, these commands are wrapped inside an execution loop that automatically pauses at each checkpoint to solicit user approval.

Summary

  • Discovery mechanism: Agents scan for manifest.json files to identify compatible skills and parse SKILL.md for workflow definitions.
  • Execution flow: Content generation → Checkpoint Plan → Scaffold/Development → Checkpoint Audio → Recording.
  • Key reference files: SCRIPT-STYLE.md, OUTLINE-FORMAT.md, CHAPTER-CRAFT.md, THEMES.md, AUDIO.md, and RECORDING.md provide the constraints and guides the agent follows.
  • Checkpoint protocol: Hard-coded pauses at Checkpoint Plan and Checkpoint Audio ensure user validation before proceeding.
  • Script invocation: The agent calls scripts/scaffold.sh and npm scripts (extract-narrations, synthesize-audio) to perform concrete operations.

Frequently Asked Questions

What file does an AI agent look for first when discovering a skill?

The agent first scans for manifest.json using a glob pattern like **/manifest.json. This file declares the skill's name, version, and compatible execution environments (Claude, Gemini, Codex), allowing the agent to determine if it can run the skill before loading any workflow definitions.

How does the agent know when to pause for user input?

The agent monitors the workflow state for hard-coded checkpoint markers defined in SKILL.md, specifically Checkpoint Plan (after theme selection) and Checkpoint Audio (before narration synthesis). These checkpoints force the agent to emit a summary and wait for explicit user confirmation before proceeding to the next phase.

What is the single source of truth for step count in chapter development?

According to references/CHAPTER-CRAFT.md, the file narrations.ts serves as the single source of truth for step count. The agent enforces this rule during Phase 2 to ensure that the generated chapters maintain consistent pacing and structure throughout the presentation.

Can AI agents automatically execute the entire workflow without stopping?

No, the Garden Skills architecture requires explicit user validation at checkpoints. While the agent can autonomously parse manifests, generate content, and execute scripts like scaffold.sh, it must pause at Checkpoint Plan and Checkpoint Audio to confirm alignment items and audio decisions before continuing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →