How AI Agents Discover and Utilize Skills from the Garden Skills Repository
AI agents discover skills by scanning manifest.json files for metadata and compatibility, then execute workflows defined in SKILL.md by parsing reference guides and running scaffold scripts while respecting checkpoint protocols for user confirmation.
The Garden Skills repository (ConardLi/garden-skills) provides a structured catalog of reusable content-creation skills that AI agents can autonomously discover and execute. Each skill is a self-contained code pack containing declarative metadata and imperative scripts that enable end-to-end workflow automation. Understanding how AI agents interact with these artifacts reveals the architecture behind autonomous skill utilization in modern development environments.
How AI Agents Discover Skills in the Repository
Parsing the Manifest File for Metadata
Agents begin discovery by scanning the repository for manifest.json files using glob patterns like **/manifest.json. Located at paths such as skills/web-video-presentation/manifest.json, this file declares the skill's identity, version, category, and compatible execution environments via the compat array (e.g., Claude, Gemini, Codex). The agent uses this metadata to determine whether it can execute the skill on its current runtime before loading any workflow definitions.
Loading the Skill Definition
After locating a manifest, the agent opens the sibling SKILL.md file to parse the human-readable workflow definition. This markdown-encoded file contains structured sections (such as ## 工作流总览 and ## Phase 1) that drive the agent's internal state machine. The agent stores the step order in memory, identifying specific phases and checkpoint markers that dictate when to pause for user input.
How AI Agents Execute Skill Workflows
Phase 1 – Content Generation
The agent initiates the workflow by parsing user input against style guides stored in references/SCRIPT-STYLE.md and references/OUTLINE-FORMAT.md. These reference files provide the constraints and formatting rules the agent uses to generate script.md and outline.md files. The agent treats these references as lookup tables for content structure, ensuring the generated artifacts match the skill's specifications.
Checkpoint Plan – Theme Selection and User Confirmation
Before proceeding to development, the agent enters the Checkpoint Plan phase. It scans every theme file under themes/*/theme.json (documented in references/THEMES.md) to recommend a matching visual theme. The agent then pauses to confirm five alignment items with the user: script, outline, theme, assets, and development mode. This hard-coded checkpoint ensures human validation before resource-intensive operations begin.
Phase 2 – Scaffold and Chapter Development
Upon user confirmation, the agent executes scripts/scaffold.sh to generate a Vite + React + TypeScript project structure. Each chapter is built according to references/CHAPTER-CRAFT.md, which encodes ten design principles and a "chapter-craft" decision tree. The agent strictly adheres to the rule that narrations.ts serves as the single source of truth for step count, ensuring consistency across generated components.
Checkpoint Audio and Phase 4 Recording
After web page generation, the agent reaches Checkpoint Audio to determine whether to synthesize narration. If confirmed, it follows the pipeline described in references/AUDIO.md and executes provider scripts located in templates/scripts/tts-providers/README.md. Finally, the agent utilizes references/RECORDING.md to capture screen recordings, supporting both automatic (?auto=1) and manual capture modes.
Self-Validation Checkpoints and State Management
The workflow implements hard-node checkpoint markers (e.g., Checkpoint Plan, Checkpoint Audio) that force the agent to emit summaries and await user approval. These checkpoints are defined in specific line ranges of SKILL.md (lines 67-84 and 100-115) and serve as guardrails that prevent autonomous execution from proceeding without validation. The agent monitors its workflow state continuously, transitioning between declarative configuration and imperative script execution only after explicit confirmation.
Practical Code Example: Discovering and Running a Skill
The following bash commands illustrate how an AI-powered tool discovers and executes the web-video-presentation skill:
# 1. Discover all skills via manifest files
manifest_files=$(opencode glob "**/manifest.json")
skill_dir=$(dirname $(echo $manifest_files | grep "web-video-presentation"))
# 2. Load manifest metadata
cat $skill_dir/manifest.json
# 3. Parse workflow phases from skill definition
opencode read "$skill_dir/SKILL.md" | grep "## Phase"
# 4. Execute scaffold script with selected theme
bash $skill_dir/scripts/scaffold.sh ./my-presentation --theme=warm-keynote
# 5. Navigate to project and generate audio (Phase 3)
cd my-presentation
npm run extract-narrations
npm run synthesize-audio
In a production AI agent, these commands are wrapped inside an execution loop that automatically pauses at each checkpoint to solicit user approval.
Summary
- Discovery mechanism: Agents scan for
manifest.jsonfiles to identify compatible skills and parseSKILL.mdfor workflow definitions. - Execution flow: Content generation → Checkpoint Plan → Scaffold/Development → Checkpoint Audio → Recording.
- Key reference files:
SCRIPT-STYLE.md,OUTLINE-FORMAT.md,CHAPTER-CRAFT.md,THEMES.md,AUDIO.md, andRECORDING.mdprovide the constraints and guides the agent follows. - Checkpoint protocol: Hard-coded pauses at Checkpoint Plan and Checkpoint Audio ensure user validation before proceeding.
- Script invocation: The agent calls
scripts/scaffold.shand npm scripts (extract-narrations,synthesize-audio) to perform concrete operations.
Frequently Asked Questions
What file does an AI agent look for first when discovering a skill?
The agent first scans for manifest.json using a glob pattern like **/manifest.json. This file declares the skill's name, version, and compatible execution environments (Claude, Gemini, Codex), allowing the agent to determine if it can run the skill before loading any workflow definitions.
How does the agent know when to pause for user input?
The agent monitors the workflow state for hard-coded checkpoint markers defined in SKILL.md, specifically Checkpoint Plan (after theme selection) and Checkpoint Audio (before narration synthesis). These checkpoints force the agent to emit a summary and wait for explicit user confirmation before proceeding to the next phase.
What is the single source of truth for step count in chapter development?
According to references/CHAPTER-CRAFT.md, the file narrations.ts serves as the single source of truth for step count. The agent enforces this rule during Phase 2 to ensure that the generated chapters maintain consistent pacing and structure throughout the presentation.
Can AI agents automatically execute the entire workflow without stopping?
No, the Garden Skills architecture requires explicit user validation at checkpoints. While the agent can autonomously parse manifests, generate content, and execute scripts like scaffold.sh, it must pause at Checkpoint Plan and Checkpoint Audio to confirm alignment items and audio decisions before continuing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →