# How the TDD Approach Applies to Writing Plugin Skills: The RED-GREEN-REFACTOR Cycle Explained

> Master TDD for plugin skills using the RED-GREEN-REFACTOR cycle. Learn how to write failing markdown, implement minimal SKILL.md files, and refine your code for continuous verification.

- Repository: [OpenAI/plugins](https://github.com/openai/plugins)
- Tags: best-practices
- Published: 2026-09-10

---

**The TDD approach applies to plugin skills through a strict RED-GREEN-REFACTOR cycle where developers first write failing markdown scenarios without the skill, then implement minimal SKILL.md files to make them pass, and finally refine the implementation while maintaining continuous verification.**

The OpenAI Plugins repository treats agent skills as first-class testable artifacts, embedding Test-Driven Development principles directly into the skill authoring workflow. This methodology ensures that every skill addresses a proven failure mode before being shipped, creating a safety-first development environment for AI agent capabilities. According to the source code in `openai/plugins`, the `superpowers` package contains canonical guidance that transforms traditional unit testing patterns into agent-skill interaction tests.

## The RED-GREEN-REFACTOR Cycle for Skills

The TDD approach for plugin skills follows the same rigorous cycle used in software development, but adapts it for agent behavior validation through markdown-based scenarios.

### RED Phase – Establishing the Baseline Failure

During the **RED** phase, you write a baseline scenario **without** the skill and execute it against the agent. The agent is expected to fail by producing a rationalization or error, which exposes the exact behavior the skill must prevent. This "pressure-test" step validates that the agent cannot perform the target task unassisted.

In [`plugins/superpowers/skills/writing-skills/testing-skills-with-subagents.md`](https://github.com/openai/plugins/blob/main/plugins/superpowers/skills/writing-skills/testing-skills-with-subagents.md), this phase is described as creating a scenario block that documents the expected failure message. By capturing the specific rationalization the agent provides—such as *"I don't know how to add numbers yet"*—you create a reproducible baseline that defines the skill's scope.

### GREEN Phase – Implementing the Minimal Skill

During the **GREEN** phase, you draft the [`SKILL.md`](https://github.com/openai/plugins/blob/main/SKILL.md) file that directly addresses the observed failure from the RED phase. The implementation must be minimal, containing only enough logic to transform the failing scenario into a success.

The canonical guide in [`plugins/superpowers/skills/test-driven-development/SKILL.md`](https://github.com/openai/plugins/blob/main/plugins/superpowers/skills/test-driven-development/SKILL.md) formalizes this process, prescribing a specific markdown structure that includes:
- A description block explaining the capability
- Implementation guidelines (often pseudo-code)
- Explicit acceptance criteria referencing the RED phase scenario

### REFACTOR Phase – Verifying and Optimizing

During the **REFACTOR** phase, you rerun the scenario **with** the new skill to verify compliance. If the agent still fails, you refine the skill until the test passes, then remove any redundant clauses or ambiguous language. This ensures the skill remains focused and maintainable.

As documented in the testing guidelines, this phase includes continuous integration checks where each skill commit executes its associated scenario files, catching regressions at the agent-skill interaction level rather than the code level.

## Architectural Benefits of TDD for Plugin Development

Applying the TDD approach to skill authoring creates four distinct architectural advantages:

- **Safety-First Development** – Skills are only shipped after a proven failure mode has been demonstrated and remedied, preventing brittle implementations.
- **Documentation-Driven Tests** – The same markdown describing the skill encodes its test cases, ensuring documentation and verification remain synchronized.
- **Explicit Failure Modes** – The RED phase records the exact rationalization the agent gave, providing future maintainers with the context for *why* a skill exists.
- **Composable Sub-Agents** – When skills depend on other skills, the TDD pattern applies recursively, enabling a tree of verified capabilities where each layer is pressure-tested.

## Practical Example – Building the `add_numbers` Skill

The following example demonstrates the complete TDD workflow using a fictional arithmetic skill, illustrating how markdown scenarios drive the implementation.

### Step 1 – RED Phase (Failing Scenario)

Create a test scenario file that expects failure when the skill is absent:

```markdown

# Scenario: Add two numbers

## Input

{
  "operation": "add",
  "a": 3,
  "b": 5
}

## Expected Failure (RED)

The agent should respond with an error like:
> "I don't know how to add numbers yet."

```

Executing this against the agent confirms the baseline failure, establishing the need for the skill.

### Step 2 – GREEN Phase (SKILL.md Implementation)

Create the minimal skill at [`plugins/example/skills/add_numbers/SKILL.md`](https://github.com/openai/plugins/blob/main/plugins/example/skills/add_numbers/SKILL.md):

```markdown

# Skill: add_numbers

## Description

Enables the agent to perform simple arithmetic addition.

## Implementation (pseudo-code)

```python
def add_numbers(a, b):
    return a + b

```

## Acceptance Criteria (GREEN)

When the above RED scenario is rerun, the agent should now reply:
> "The result is 8."

```

### Step 3 – REFACTOR Phase (Verification)

Rerun the scenario to confirm the agent returns *"The result is 8."* If the response includes extraneous content, adjust the skill constraints to enforce minimal, consistent output. Finally, remove any placeholder text in the skill description to maintain clarity.

## Key Source Files for TDD Skill Development

The OpenAI Plugins repository provides three critical resources for implementing this workflow:

- **[`plugins/superpowers/skills/test-driven-development/SKILL.md`](https://github.com/openai/plugins/blob/main/plugins/superpowers/skills/test-driven-development/SKILL.md)** – The canonical TDD guide defining required markdown structure, naming conventions, and success criteria formats.
- **[`plugins/superpowers/skills/writing-skills/testing-skills-with-subagents.md`](https://github.com/openai/plugins/blob/main/plugins/superpowers/skills/writing-skills/testing-skills-with-subagents.md)** – Detailed documentation of the RED-GREEN-REFACTOR process and checklists for skill testing with subagents.
- **`plugins/example/skills/add_numbers/`** – A concrete demonstration directory showing a complete skill written using TDD principles.

## Summary

- The TDD approach for plugin skills uses a **RED-GREEN-REFACTOR** cycle adapted for agent behavior validation through markdown scenarios.
- **RED phase**: Document baseline failures by running scenarios without the skill to capture exact agent rationalizations.
- **GREEN phase**: Implement minimal [`SKILL.md`](https://github.com/openai/plugins/blob/main/SKILL.md) files that address only the observed failure mode.
- **REFACTOR phase**: Verify compliance through continuous scenario execution and remove redundant clauses.
- The methodology is codified in [`testing-skills-with-subagents.md`](https://github.com/openai/plugins/blob/main/testing-skills-with-subagents.md) and [`test-driven-development/SKILL.md`](https://github.com/openai/plugins/blob/main/test-driven-development/SKILL.md) within the OpenAI Plugins repository.

## Frequently Asked Questions

### What makes the TDD approach different for plugin skills versus traditional code?

Traditional TDD tests functions and classes, while the TDD approach for plugin skills tests **agent-skill interactions** through markdown scenarios. Instead of asserting on return values, you validate that the agent's rationalization changes from a specific failure message to a correct response, as recorded in the RED and GREEN phase documentation.

### How do you handle dependencies between skills in the TDD workflow?

When a skill depends on other skills, the TDD pattern applies recursively. You pressure-test the dependent skill by first demonstrating that the sub-agent fails without the dependency (RED), then introducing the dependency to achieve success (GREEN). This creates a **tree of verified capabilities** where each composable layer is validated before integration.

### Where should scenario files be stored in the repository?

Scenario files should be stored alongside the [`SKILL.md`](https://github.com/openai/plugins/blob/main/SKILL.md) file in the skill directory, typically under paths like `plugins/example/skills/[skill_name]/`. This collocation enables continuous integration systems to execute the markdown scenarios on every commit, ensuring that changes to the skill do not break existing agent behaviors.

### What is the "pressure-test" step in skill development?

The pressure-test step refers to the RED phase validation where you run the scenario **without** the skill implemented to confirm the agent genuinely cannot perform the task. This exposes the exact failure mode—whether an error, hallucination, or refusal—that the skill must remediate, preventing false positives in the GREEN phase.