# What Is the Agent Skills Spec? A Technical Guide to the ConardLi Garden-Skills Contract

> Understand the Agent Skills spec a lightweight JSON based contract allowing AI agents to discover and invoke reusable functionality. Learn more in this technical guide.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: deep-dive
- Published: 2026-08-29

---

**The Agent Skills spec is a lightweight, JSON‑based contract that enables AI agents to discover, invoke, and understand reusable executable functionality through standardized manifest files and CLI entry points.**

The **Agent Skills spec** defines the structural contract used throughout the `ConardLi/garden-skills` repository (also referenced as the *Instagit* codebase). This specification allows autonomous agents to parse skill capabilities without hard‑coded logic, enabling dynamic plugin architectures where new capabilities can be added by dropping folders into the `skills/` directory.

## Core File Structure of the Agent Skills Spec

Every skill package in the repository follows a predictable layout that separates metadata, documentation, and executable logic.

### manifest.json - Global Package Description

The [`manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/manifest.json) file serves as the machine‑readable contract. It declares the skill’s identity, runtime requirements, and interface constraints. According to the source code, typical fields include `name`, `description`, `author`, `version`, `license`, `type` (usually set to `"skill"`), and `runtime` (e.g., `"node"` or `"python"`).

The manifest also defines the `entry` point or `script` path that the Agent will invoke. For example, the **gpt‑image‑2** skill stores its manifest at [[`skills/gpt-image-2/manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/manifest.json)](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/manifest.json), which specifies how the Agent should call the underlying generation logic.

### SKILL.md - Human-Readable Agent Context

The [`SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/SKILL.md) file provides natural‑language documentation that the Agent loads as prompt context. Unlike the structured JSON manifest, this file contains high‑level descriptions, usage patterns, mode‑selection logic, example prompts, and operational cautions.

In [`skills/gpt-image-2/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/SKILL.md), the document outlines when to delegate image generation to local models versus host‑native APIs, allowing the Agent to make informed routing decisions based on context window and latency requirements.

### manifest.schema.json - Formal Validation Schema

Located at [[`skills/manifest.schema.json`](https://github.com/ConardLi/garden-skills/blob/main/skills/manifest.schema.json)](https://github.com/ConardLi/garden-skills/blob/main/skills/manifest.schema.json), this optional JSON‑Schema defines strict validation rules for all manifests. It specifies `type` constraints, required fields, property types, and enum values, ensuring that every skill package adheres to the **Agent Skills spec** before execution.

### scripts/ - Implementation Entry Points

The `scripts/` directory contains the actual executables the Agent invokes via CLI. Each script exposes a stable interface that accepts arguments according to the manifest’s `parameters` definition.

For instance, a skill might expose:

```bash
node skills/gpt-image-2/scripts/generate.js --prompt "A futuristic garden" --output ./result.png

```

Other common entry points include [`edit.js`](https://github.com/ConardLi/garden-skills/blob/main/edit.js) or language‑specific equivalents like [`generate.py`](https://github.com/ConardLi/garden-skills/blob/main/generate.py), depending on the declared `runtime`.

### references/ - Structured Templates

This directory holds Markdown, JSON, or binary assets that the skill reads at runtime. These files act as structured templates or few‑shot examples that influence the skill’s output without requiring code changes.

## Key Elements of the Agent Skills Spec

### Name, Description, and Versioning

Every skill must expose a concise `name` (kebab‑case identifier) and a human‑readable `description`. The `version` field follows semantic versioning, allowing Agents to select compatible implementations when multiple versions coexist in the `skills/` directory.

### Runtime Declaration and Entry Points

The `runtime` field tells the Agent which interpreter to spawn (`node`, `python`, `bash`, etc.). The `entry` or `script` field provides the relative path from the skill root to the executable file. This decoupling allows the **Agent Skills spec** to support polyglot implementations within the same catalogue.

### Input Parameters and Output Schemas

The manifest defines a `parameters` object that describes required inputs (e.g., `prompt`, `image_path`, `temperature`). These definitions enable the Agent to construct valid CLI arguments or environment variables dynamically.

Outputs are typically defined in the `outputs` section, specifying whether the skill returns a text string, a file path, or a structured JSON payload. This schema drives how the Agent parses and integrates results into the final user response.

### Mode Awareness and Operation Logic

Advanced skills expose multiple operation modes. For example, **gpt‑image‑2** includes logic for local generation, host‑native delegation, and advisory‑only modes. These modes are enumerated in the manifest via a `modes` array and documented in [`SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/SKILL.md), allowing the Agent to select the appropriate execution strategy based on resource availability or user preferences.

## How Agents Use the Agent Skills Spec

### Discovery Phase

The Agent scans the repository’s `skills/` directory at startup, recursively reading each [`manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/manifest.json) to build an in‑memory catalogue. Invalid manifests are rejected based on the rules defined in [`manifest.schema.json`](https://github.com/ConardLi/garden-skills/blob/main/manifest.schema.json).

### Selection and Matching

When processing a user request, the Agent performs semantic matching against the `description` and `name` fields of discovered skills. It weighs version constraints and mode availability to select the optimal implementation.

### Invocation and Payload Assembly

Using the `parameters` schema, the Agent marshals user input into CLI arguments or environment variables. It then spawns a subprocess using the declared `runtime`, targeting the script specified in `entry`.

```bash

# Example invocation constructed by the Agent

node skills/gpt-image-2/scripts/generate.js \
  --prompt "Generate a logo" \
  --mode local \
  --output ./assets/logo.png

```

### Result Integration

The Agent captures stdout and stderr from the subprocess, parsing the output according to the `outputs` schema defined in the manifest. Structured results are injected back into the conversation context or saved to disk, depending on the skill’s contract.

## Real-World Example: The gpt-image-2 Skill

The **gpt‑image‑2** skill demonstrates a complete implementation of the **Agent Skills spec**:

- **Manifest**: [[`skills/gpt-image-2/manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/manifest.json)](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/manifest.json) declares the runtime as `node`, the entry point as [`scripts/generate.js`](https://github.com/ConardLi/garden-skills/blob/main/scripts/generate.js), and supported modes.
- **Documentation**: [[`skills/gpt-image-2/SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/SKILL.md)](https://github.com/ConardLi/garden-skills/blob/main/skills/gpt-image-2/SKILL.md) provides natural‑language guidance on when to use each mode.
- **Schema**: Validated against [[`skills/manifest.schema.json`](https://github.com/ConardLi/garden-skills/blob/main/skills/manifest.schema.json)](https://github.com/ConardLi/garden-skills/blob/main/skills/manifest.schema.json) to ensure contract compliance.

## Summary

- The **Agent Skills spec** is a JSON‑based contract enabling dynamic discovery and invocation of executable capabilities.
- Each skill requires [`manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/manifest.json) for machine metadata, [`SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/SKILL.md) for human context, and `scripts/` for implementation.
- The spec supports multiple runtimes (`node`, `python`) and operation modes through declarative configuration.
- Agents parse `parameters` and `outputs` schemas to construct CLI calls and integrate results automatically.
- The `ConardLi/garden-skills` repository provides a reference implementation used in the Instagit codebase.

## Frequently Asked Questions

### What file format does the Agent Skills spec use?

The **Agent Skills spec** uses JSON for machine‑readable metadata ([`manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/manifest.json)) and Markdown for human‑readable documentation ([`SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/SKILL.md)). This dual‑format approach allows both automated Agents and human developers to understand skill capabilities without parsing code.

### How does an Agent determine which runtime to execute?

The Agent reads the `runtime` field from [`manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/manifest.json), which specifies the interpreter (such as `node`, `python`, or `bash`). It then locates the executable path from the `entry` or `script` field and spawns the corresponding subprocess to invoke the skill.

### What is the difference between manifest.json and SKILL.md?

[`manifest.json`](https://github.com/ConardLi/garden-skills/blob/main/manifest.json) provides structured, machine‑parseable data including version, parameters, and runtime requirements. [`SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/SKILL.md) contains unstructured natural‑language instructions that guide the Agent’s reasoning process, including usage examples and mode‑selection logic.

### Can a single skill support multiple operation modes?

Yes. The spec supports a `modes` array in the manifest (e.g., `["local", "remote", "advisory"]`) that defines distinct operational behaviors. The Agent uses [`SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/SKILL.md) context to decide which mode to invoke based on environmental constraints or user intent.