What Is the Agent Skills Spec? A Technical Guide to the ConardLi Garden-Skills Contract

The Agent Skills spec is a lightweight, JSON‑based contract that enables AI agents to discover, invoke, and understand reusable executable functionality through standardized manifest files and CLI entry points.

The Agent Skills spec defines the structural contract used throughout the ConardLi/garden-skills repository (also referenced as the Instagit codebase). This specification allows autonomous agents to parse skill capabilities without hard‑coded logic, enabling dynamic plugin architectures where new capabilities can be added by dropping folders into the skills/ directory.

Core File Structure of the Agent Skills Spec

Every skill package in the repository follows a predictable layout that separates metadata, documentation, and executable logic.

manifest.json - Global Package Description

The manifest.json file serves as the machine‑readable contract. It declares the skill’s identity, runtime requirements, and interface constraints. According to the source code, typical fields include name, description, author, version, license, type (usually set to "skill"), and runtime (e.g., "node" or "python").

The manifest also defines the entry point or script path that the Agent will invoke. For example, the gpt‑image‑2 skill stores its manifest at [skills/gpt-image-2/manifest.json](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/manifest.json), which specifies how the Agent should call the underlying generation logic.

SKILL.md - Human-Readable Agent Context

The SKILL.md file provides natural‑language documentation that the Agent loads as prompt context. Unlike the structured JSON manifest, this file contains high‑level descriptions, usage patterns, mode‑selection logic, example prompts, and operational cautions.

In skills/gpt-image-2/SKILL.md, the document outlines when to delegate image generation to local models versus host‑native APIs, allowing the Agent to make informed routing decisions based on context window and latency requirements.

manifest.schema.json - Formal Validation Schema

Located at [skills/manifest.schema.json](https://github.com/ConardLi/garden-skills/blob/main/skills/manifest.schema.json), this optional JSON‑Schema defines strict validation rules for all manifests. It specifies type constraints, required fields, property types, and enum values, ensuring that every skill package adheres to the Agent Skills spec before execution.

scripts/ - Implementation Entry Points

The scripts/ directory contains the actual executables the Agent invokes via CLI. Each script exposes a stable interface that accepts arguments according to the manifest’s parameters definition.

For instance, a skill might expose:

node skills/gpt-image-2/scripts/generate.js --prompt "A futuristic garden" --output ./result.png

Other common entry points include edit.js or language‑specific equivalents like generate.py, depending on the declared runtime.

references/ - Structured Templates

This directory holds Markdown, JSON, or binary assets that the skill reads at runtime. These files act as structured templates or few‑shot examples that influence the skill’s output without requiring code changes.

Key Elements of the Agent Skills Spec

Name, Description, and Versioning

Every skill must expose a concise name (kebab‑case identifier) and a human‑readable description. The version field follows semantic versioning, allowing Agents to select compatible implementations when multiple versions coexist in the skills/ directory.

Runtime Declaration and Entry Points

The runtime field tells the Agent which interpreter to spawn (node, python, bash, etc.). The entry or script field provides the relative path from the skill root to the executable file. This decoupling allows the Agent Skills spec to support polyglot implementations within the same catalogue.

Input Parameters and Output Schemas

The manifest defines a parameters object that describes required inputs (e.g., prompt, image_path, temperature). These definitions enable the Agent to construct valid CLI arguments or environment variables dynamically.

Outputs are typically defined in the outputs section, specifying whether the skill returns a text string, a file path, or a structured JSON payload. This schema drives how the Agent parses and integrates results into the final user response.

Mode Awareness and Operation Logic

Advanced skills expose multiple operation modes. For example, gpt‑image‑2 includes logic for local generation, host‑native delegation, and advisory‑only modes. These modes are enumerated in the manifest via a modes array and documented in SKILL.md, allowing the Agent to select the appropriate execution strategy based on resource availability or user preferences.

How Agents Use the Agent Skills Spec

Discovery Phase

The Agent scans the repository’s skills/ directory at startup, recursively reading each manifest.json to build an in‑memory catalogue. Invalid manifests are rejected based on the rules defined in manifest.schema.json.

Selection and Matching

When processing a user request, the Agent performs semantic matching against the description and name fields of discovered skills. It weighs version constraints and mode availability to select the optimal implementation.

Invocation and Payload Assembly

Using the parameters schema, the Agent marshals user input into CLI arguments or environment variables. It then spawns a subprocess using the declared runtime, targeting the script specified in entry.


# Example invocation constructed by the Agent

node skills/gpt-image-2/scripts/generate.js \
  --prompt "Generate a logo" \
  --mode local \
  --output ./assets/logo.png

Result Integration

The Agent captures stdout and stderr from the subprocess, parsing the output according to the outputs schema defined in the manifest. Structured results are injected back into the conversation context or saved to disk, depending on the skill’s contract.

Real-World Example: The gpt-image-2 Skill

The gpt‑image‑2 skill demonstrates a complete implementation of the Agent Skills spec:

Summary

  • The Agent Skills spec is a JSON‑based contract enabling dynamic discovery and invocation of executable capabilities.
  • Each skill requires manifest.json for machine metadata, SKILL.md for human context, and scripts/ for implementation.
  • The spec supports multiple runtimes (node, python) and operation modes through declarative configuration.
  • Agents parse parameters and outputs schemas to construct CLI calls and integrate results automatically.
  • The ConardLi/garden-skills repository provides a reference implementation used in the Instagit codebase.

Frequently Asked Questions

What file format does the Agent Skills spec use?

The Agent Skills spec uses JSON for machine‑readable metadata (manifest.json) and Markdown for human‑readable documentation (SKILL.md). This dual‑format approach allows both automated Agents and human developers to understand skill capabilities without parsing code.

How does an Agent determine which runtime to execute?

The Agent reads the runtime field from manifest.json, which specifies the interpreter (such as node, python, or bash). It then locates the executable path from the entry or script field and spawns the corresponding subprocess to invoke the skill.

What is the difference between manifest.json and SKILL.md?

manifest.json provides structured, machine‑parseable data including version, parameters, and runtime requirements. SKILL.md contains unstructured natural‑language instructions that guide the Agent’s reasoning process, including usage examples and mode‑selection logic.

Can a single skill support multiple operation modes?

Yes. The spec supports a modes array in the manifest (e.g., ["local", "remote", "advisory"]) that defines distinct operational behaviors. The Agent uses SKILL.md context to decide which mode to invoke based on environmental constraints or user intent.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →