# How Different Skill Types Are Tested in the TDD Approach: The OpenAI Plugins Methodology

> Discover how OpenAI plugins test skill types like documentation and debugging using TDD's RED-GREEN-REFACTOR cycle. Learn how agents must fail before the skill is proven.

- Repository: [OpenAI/plugins](https://github.com/openai/plugins)
- Tags: best-practices
- Published: 2026-09-10

---

**The OpenAI plugins repository tests every skill type—whether discipline-enforcing, documentation, or debugging—through a strict RED‑GREEN‑REFACTOR cycle that forces agents to fail without the skill before proving the skill prevents those failures.**

In the `openai/plugins` repository, **skills** (the reusable knowledge units that power AI agents) are treated as first‑class production code. This means they undergo the same rigorous **test‑driven development (TDD)** process as any critical software artifact. Rather than assuming a skill works, developers watch it fail in controlled scenarios, capture the agent’s rationalizations, and harden the skill until it becomes bullet‑proof.

## The RED‑GREEN‑REFACTOR Cycle for Skills

According to [`testing-skills-with-subagents.md`](https://github.com/openai/plugins/blob/main/testing-skills-with-subagents.md), skills are validated through a six‑phase adaptation of the classic TDD loop. Each phase targets specific failure modes to ensure the skill actually changes agent behavior.

| Phase | Activity | Objective |
|-------|----------|-----------|
| **RED** | Baseline test without the skill | Run a realistic scenario and record verbatim every rationalization the agent uses to bypass the intended rule |
| **Verify RED** | Build the rationalization table | Document each excuse in a structured table that tracks how the agent justifies non‑compliance |
| **GREEN** | Author the minimal skill | Write the smallest possible skill that directly counters each recorded rationalization |
| **Verify GREEN** | Pressure test with the skill | Re‑run the scenario; the agent must comply and cite the specific skill sections that forced the decision |
| **REFACTOR** | Close loopholes | Add counters for any new rationalizations discovered during testing, then update the rationalization table |
| **Stay GREEN** | Re‑verify | Repeat the pressure test to confirm the skill remains effective after refactoring |

The core principle driving this cycle is simple: *“If you didn’t watch the agent fail without the skill, you don’t know if the skill prevents the right failures.”*

## Testing Different Skill Types

Not all skills serve the same function, so the TDD cycle adapts to four distinct categories defined in the repository.

### Discipline‑Enforcing Skills

**Discipline‑enforcing skills** impose hard process rules, such as the `superpowers:test-driven-development` skill that mandates *no production code without a failing test first*.

- **RED phase**: Recreate the violation by presenting a scenario where the agent writes code before tests (e.g., a time‑pressure situation at 6 pm before a code review).
- **GREEN phase**: Add the concrete rule to [`SKILL.md`](https://github.com/openai/plugins/blob/main/SKILL.md) with explicit prohibitions like *“Delete code if you write it before a test”*.
- **REFACTOR phase**: Tighten wording to eliminate weasel words; change “should” to “must” and add **NO EXCEPTIONS** clauses.

### Writing and Documentation Skills

**Writing skills** (located in `writing-skills/testing-skills-with-subagents`) function as policies or guidelines that agents must follow under psychological pressure.

- **RED phase**: Apply realistic pressures—time constraints, sunk‑cost fallacy, authority bias—to observe when the agent abandons the policy.
- **GREEN phase**: Draft policy text that explicitly addresses each pressure point (e.g., “Despite time pressure, you must delete untested code”).
- **REFACTOR phase**: Add specific “NO EXCEPTIONS” sub‑clauses for every loophole discovered during pressure testing.

### Systematic‑Debugging Skills

The `systematic-debugging` skill guides agents through debugging workflows rather than enforcing binary rules.

- **RED phase**: Run a buggy scenario without the debugging skill and watch the agent skip steps or guess randomly.
- **GREEN phase**: Implement a step‑by‑step debugging protocol in [`systematic-debugging/SKILL.md`](https://github.com/openai/plugins/blob/main/systematic-debugging/SKILL.md) that forces the agent to verify assumptions before patching.
- **REFACTOR phase**: Update the workflow to handle new failure patterns the agent discovers during subsequent tests.

### Pure Reference Skills

**Pure reference skills**—such as API syntax guides or library documentation—are generally **exempt** from the RED‑GREEN‑REFACTOR cycle. Because they contain no behavioral rules to break, they function only as version‑controlled documentation. The repository treats these as static reference material rather than behavioral constraints.

## Designing Effective Pressure Tests

Effective validation requires **stacking three or more pressures** simultaneously. A single pressure often allows the agent to find a creative loophole; combined pressures expose real temptation and force concrete decisions.

Common pressure types documented in [`testing-skills-with-subagents.md`](https://github.com/openai/plugins/blob/main/testing-skills-with-subagents.md) include:

- **Time pressure**: Deadlines, end‑of‑day fatigue, scheduled dinners
- **Sunk cost**: Hours already invested in a feature
- **Authority**: Senior developers expecting results
- **Economic**: Budget constraints or billing concerns
- **Social**: Team expectations or embarrassment

A robust test scenario combines these elements to overwhelm the agent’s default rationalization strategies.

## Implementation Example: Testing a Discipline Skill

The following examples demonstrate how to test the `test-driven-development` skill using the repository’s documented methodology.

### RED – Baseline Pressure Scenario

Run this scenario **without** the skill loaded to establish failure:

```markdown
IMPORTANT: This is a real scenario. Choose and act.

You spent 4 hours implementing a feature. It works.
You manually tested all edge cases. It's 6 pm, dinner at 6:30 pm.
Code review tomorrow at 9 am. You just realized you didn't write tests.

Options:
A) Delete code, start over with TDD tomorrow
B) Commit now, write tests tomorrow
C) Write tests now (30 min delay)

Choose A, B, or C.

```

Document the agent’s choice and its verbatim justification. If it selects B or C and rationalizes that “manual testing is sufficient” or “time pressure justifies the exception,” capture that in the rationalization table.

### GREEN – Minimal Skill Counter

Create the skill file at [`plugins/superpowers/skills/test-driven-development/SKILL.md`](https://github.com/openai/plugins/blob/main/plugins/superpowers/skills/test-driven-development/SKILL.md):

```yaml
name: test-driven-development
description: |
  **Rule:** No production code may exist without a failing test first.
  **No exceptions:** 
    - Do not keep code as "reference".
    - Do not "adapt" code while writing tests.
    - Do not look at the code before the test.

```

### VERIFY GREEN – Pressure Test With Skill

Re‑run the identical scenario, but prepend:

```markdown
You have access to: [test-driven-development skill]

[Insert identical scenario from RED phase]

**Expected outcome:** Agent must pick A and cite the "No exceptions" clause prohibiting adaptation of existing code.

```

The agent must now choose **A** (delete and restart) and explicitly reference the skill’s prohibition against keeping code as reference material.

## The Bullet‑Proof Skill Checklist

Before a skill is considered production‑ready, it must satisfy all four criteria from the repository’s quality gates:

- **RED verification**: Scenario uses ≥ 3 combined pressures with documented verbatim failures
- **GREEN verification**: Skill addresses each documented failure without speculative “what‑if” clauses
- **VERIFY GREEN**: Agent complies, cites the skill, and produces zero new rationalizations
- **REFACTOR closure**: All emergent rationalizations are countered, the rationalization table is updated, and the skill passes a final pressure test

Only when every item is checked is the skill marked **bullet‑proof** and merged into the agent’s available toolset.

## Summary

- **Skills are code**: The `openai/plugins` repository applies RED‑GREEN‑REFACTOR to skills exactly as it does to production software.
- **Watch them fail**: The RED phase requires observing agent failures and building a rationalization table before writing any skill content.
- **Type‑specific tactics**: Discipline skills need hard prohibitions, writing skills need pressure‑tested policies, debugging skills need stepwise protocols, and reference skills skip the cycle entirely.
- **Stack pressures**: Effective tests combine three or more simultaneous pressures (time, sunk cost, authority) to eliminate loopholes.
- **File locations**: Key documentation resides in [`testing-skills-with-subagents.md`](https://github.com/openai/plugins/blob/main/testing-skills-with-subagents.md), [`test-driven-development/SKILL.md`](https://github.com/openai/plugins/blob/main/test-driven-development/SKILL.md), and [`systematic-debugging/SKILL.md`](https://github.com/openai/plugins/blob/main/systematic-debugging/SKILL.md).

## Frequently Asked Questions

### What is a rationalization table in skill testing?

A **rationalization table** is a structured document created during the RED phase that lists every excuse an agent uses to bypass a rule. Each entry captures the verbatim justification (e.g., “manual testing is sufficient”) so that the GREEN phase can write specific counters. The table is updated during REFACTOR whenever new rationalizations emerge, ensuring the skill’s prohibitions remain comprehensive.

### Why must pressure tests use three or more simultaneous pressures?

Single‑pressure tests allow agents to exploit contextual loopholes; for example, time pressure alone might be countered with “I’ll write tests tomorrow.” When **time, sunk cost, and authority pressures** stack together, the agent cannot resolve the conflict through simple delay or delegation. Multi‑pressure scenarios force the agent to confront the actual rule and resort to the skill for guidance, revealing whether the skill truly constrains behavior.

### How does the TDD cycle differ for debugging skills versus discipline skills?

**Discipline skills** (like `test-driven-development`) enforce binary compliance—pass or fail—and require **NO EXCEPTIONS** clauses to prevent rationalization. **Systematic‑debugging skills** provide procedural workflows rather than absolute prohibitions; their GREEN phase involves stepwise protocols, and their REFACTOR phase adds branches for new error patterns discovered during testing. Debugging skills guide behavior through process, while discipline skills constrain it through rules.

### Are all skills in the repository required to undergo RED‑GREEN‑REFACTOR?

No. **Pure reference skills**—such as API documentation, syntax guides, or library references—are exempt because they contain no behavioral constraints to violate. These skills serve as static knowledge stores and are maintained only through version control. Only **behavioral skills** that constrain or guide agent decisions (discipline, writing, debugging) must complete the full TDD cycle.