What Is the Agent Skills Spec? A Technical Guide to the ConardLi Garden-Skills Contract
The Agent Skills spec is a lightweight, JSON‑based contract that enables AI agents to discover, invoke, and understand reusable executable functionality through standardized manifest files and CLI entry points.
The Agent Skills spec defines the structural contract used throughout the ConardLi/garden-skills repository (also referenced as the Instagit codebase). This specification allows autonomous agents to parse skill capabilities without hard‑coded logic, enabling dynamic plugin architectures where new capabilities can be added by dropping folders into the skills/ directory.
Core File Structure of the Agent Skills Spec
Every skill package in the repository follows a predictable layout that separates metadata, documentation, and executable logic.
manifest.json - Global Package Description
The manifest.json file serves as the machine‑readable contract. It declares the skill’s identity, runtime requirements, and interface constraints. According to the source code, typical fields include name, description, author, version, license, type (usually set to "skill"), and runtime (e.g., "node" or "python").
The manifest also defines the entry point or script path that the Agent will invoke. For example, the gpt‑image‑2 skill stores its manifest at [skills/gpt-image-2/manifest.json](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/manifest.json), which specifies how the Agent should call the underlying generation logic.
SKILL.md - Human-Readable Agent Context
The SKILL.md file provides natural‑language documentation that the Agent loads as prompt context. Unlike the structured JSON manifest, this file contains high‑level descriptions, usage patterns, mode‑selection logic, example prompts, and operational cautions.
In skills/gpt-image-2/SKILL.md, the document outlines when to delegate image generation to local models versus host‑native APIs, allowing the Agent to make informed routing decisions based on context window and latency requirements.
manifest.schema.json - Formal Validation Schema
Located at [skills/manifest.schema.json](https://github.com/ConardLi/garden-skills/blob/main/skills/manifest.schema.json), this optional JSON‑Schema defines strict validation rules for all manifests. It specifies type constraints, required fields, property types, and enum values, ensuring that every skill package adheres to the Agent Skills spec before execution.
scripts/ - Implementation Entry Points
The scripts/ directory contains the actual executables the Agent invokes via CLI. Each script exposes a stable interface that accepts arguments according to the manifest’s parameters definition.
For instance, a skill might expose:
node skills/gpt-image-2/scripts/generate.js --prompt "A futuristic garden" --output ./result.png
Other common entry points include edit.js or language‑specific equivalents like generate.py, depending on the declared runtime.
references/ - Structured Templates
This directory holds Markdown, JSON, or binary assets that the skill reads at runtime. These files act as structured templates or few‑shot examples that influence the skill’s output without requiring code changes.
Key Elements of the Agent Skills Spec
Name, Description, and Versioning
Every skill must expose a concise name (kebab‑case identifier) and a human‑readable description. The version field follows semantic versioning, allowing Agents to select compatible implementations when multiple versions coexist in the skills/ directory.
Runtime Declaration and Entry Points
The runtime field tells the Agent which interpreter to spawn (node, python, bash, etc.). The entry or script field provides the relative path from the skill root to the executable file. This decoupling allows the Agent Skills spec to support polyglot implementations within the same catalogue.
Input Parameters and Output Schemas
The manifest defines a parameters object that describes required inputs (e.g., prompt, image_path, temperature). These definitions enable the Agent to construct valid CLI arguments or environment variables dynamically.
Outputs are typically defined in the outputs section, specifying whether the skill returns a text string, a file path, or a structured JSON payload. This schema drives how the Agent parses and integrates results into the final user response.
Mode Awareness and Operation Logic
Advanced skills expose multiple operation modes. For example, gpt‑image‑2 includes logic for local generation, host‑native delegation, and advisory‑only modes. These modes are enumerated in the manifest via a modes array and documented in SKILL.md, allowing the Agent to select the appropriate execution strategy based on resource availability or user preferences.
How Agents Use the Agent Skills Spec
Discovery Phase
The Agent scans the repository’s skills/ directory at startup, recursively reading each manifest.json to build an in‑memory catalogue. Invalid manifests are rejected based on the rules defined in manifest.schema.json.
Selection and Matching
When processing a user request, the Agent performs semantic matching against the description and name fields of discovered skills. It weighs version constraints and mode availability to select the optimal implementation.
Invocation and Payload Assembly
Using the parameters schema, the Agent marshals user input into CLI arguments or environment variables. It then spawns a subprocess using the declared runtime, targeting the script specified in entry.
# Example invocation constructed by the Agent
node skills/gpt-image-2/scripts/generate.js \
--prompt "Generate a logo" \
--mode local \
--output ./assets/logo.png
Result Integration
The Agent captures stdout and stderr from the subprocess, parsing the output according to the outputs schema defined in the manifest. Structured results are injected back into the conversation context or saved to disk, depending on the skill’s contract.
Real-World Example: The gpt-image-2 Skill
The gpt‑image‑2 skill demonstrates a complete implementation of the Agent Skills spec:
- Manifest: [
skills/gpt-image-2/manifest.json](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/manifest.json) declares the runtime asnode, the entry point asscripts/generate.js, and supported modes. - Documentation: [
skills/gpt-image-2/SKILL.md](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/SKILL.md) provides natural‑language guidance on when to use each mode. - Schema: Validated against [
skills/manifest.schema.json](https://github.com/ConardLi/garden-skills/blob/main/skills/manifest.schema.json) to ensure contract compliance.
Summary
- The Agent Skills spec is a JSON‑based contract enabling dynamic discovery and invocation of executable capabilities.
- Each skill requires
manifest.jsonfor machine metadata,SKILL.mdfor human context, andscripts/for implementation. - The spec supports multiple runtimes (
node,python) and operation modes through declarative configuration. - Agents parse
parametersandoutputsschemas to construct CLI calls and integrate results automatically. - The
ConardLi/garden-skillsrepository provides a reference implementation used in the Instagit codebase.
Frequently Asked Questions
What file format does the Agent Skills spec use?
The Agent Skills spec uses JSON for machine‑readable metadata (manifest.json) and Markdown for human‑readable documentation (SKILL.md). This dual‑format approach allows both automated Agents and human developers to understand skill capabilities without parsing code.
How does an Agent determine which runtime to execute?
The Agent reads the runtime field from manifest.json, which specifies the interpreter (such as node, python, or bash). It then locates the executable path from the entry or script field and spawns the corresponding subprocess to invoke the skill.
What is the difference between manifest.json and SKILL.md?
manifest.json provides structured, machine‑parseable data including version, parameters, and runtime requirements. SKILL.md contains unstructured natural‑language instructions that guide the Agent’s reasoning process, including usage examples and mode‑selection logic.
Can a single skill support multiple operation modes?
Yes. The spec supports a modes array in the manifest (e.g., ["local", "remote", "advisory"]) that defines distinct operational behaviors. The Agent uses SKILL.md context to decide which mode to invoke based on environmental constraints or user intent.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →