How Different Skill Types Are Tested in the TDD Approach: The OpenAI Plugins Methodology

The OpenAI plugins repository tests every skill type—whether discipline-enforcing, documentation, or debugging—through a strict RED‑GREEN‑REFACTOR cycle that forces agents to fail without the skill before proving the skill prevents those failures.

In the openai/plugins repository, skills (the reusable knowledge units that power AI agents) are treated as first‑class production code. This means they undergo the same rigorous test‑driven development (TDD) process as any critical software artifact. Rather than assuming a skill works, developers watch it fail in controlled scenarios, capture the agent’s rationalizations, and harden the skill until it becomes bullet‑proof.

The RED‑GREEN‑REFACTOR Cycle for Skills

According to testing-skills-with-subagents.md, skills are validated through a six‑phase adaptation of the classic TDD loop. Each phase targets specific failure modes to ensure the skill actually changes agent behavior.

Phase Activity Objective
RED Baseline test without the skill Run a realistic scenario and record verbatim every rationalization the agent uses to bypass the intended rule
Verify RED Build the rationalization table Document each excuse in a structured table that tracks how the agent justifies non‑compliance
GREEN Author the minimal skill Write the smallest possible skill that directly counters each recorded rationalization
Verify GREEN Pressure test with the skill Re‑run the scenario; the agent must comply and cite the specific skill sections that forced the decision
REFACTOR Close loopholes Add counters for any new rationalizations discovered during testing, then update the rationalization table
Stay GREEN Re‑verify Repeat the pressure test to confirm the skill remains effective after refactoring

The core principle driving this cycle is simple: “If you didn’t watch the agent fail without the skill, you don’t know if the skill prevents the right failures.”

Testing Different Skill Types

Not all skills serve the same function, so the TDD cycle adapts to four distinct categories defined in the repository.

Discipline‑Enforcing Skills

Discipline‑enforcing skills impose hard process rules, such as the superpowers:test-driven-development skill that mandates no production code without a failing test first.

  • RED phase: Recreate the violation by presenting a scenario where the agent writes code before tests (e.g., a time‑pressure situation at 6 pm before a code review).
  • GREEN phase: Add the concrete rule to SKILL.md with explicit prohibitions like “Delete code if you write it before a test”.
  • REFACTOR phase: Tighten wording to eliminate weasel words; change “should” to “must” and add NO EXCEPTIONS clauses.

Writing and Documentation Skills

Writing skills (located in writing-skills/testing-skills-with-subagents) function as policies or guidelines that agents must follow under psychological pressure.

  • RED phase: Apply realistic pressures—time constraints, sunk‑cost fallacy, authority bias—to observe when the agent abandons the policy.
  • GREEN phase: Draft policy text that explicitly addresses each pressure point (e.g., “Despite time pressure, you must delete untested code”).
  • REFACTOR phase: Add specific “NO EXCEPTIONS” sub‑clauses for every loophole discovered during pressure testing.

Systematic‑Debugging Skills

The systematic-debugging skill guides agents through debugging workflows rather than enforcing binary rules.

  • RED phase: Run a buggy scenario without the debugging skill and watch the agent skip steps or guess randomly.
  • GREEN phase: Implement a step‑by‑step debugging protocol in systematic-debugging/SKILL.md that forces the agent to verify assumptions before patching.
  • REFACTOR phase: Update the workflow to handle new failure patterns the agent discovers during subsequent tests.

Pure Reference Skills

Pure reference skills—such as API syntax guides or library documentation—are generally exempt from the RED‑GREEN‑REFACTOR cycle. Because they contain no behavioral rules to break, they function only as version‑controlled documentation. The repository treats these as static reference material rather than behavioral constraints.

Designing Effective Pressure Tests

Effective validation requires stacking three or more pressures simultaneously. A single pressure often allows the agent to find a creative loophole; combined pressures expose real temptation and force concrete decisions.

Common pressure types documented in testing-skills-with-subagents.md include:

  • Time pressure: Deadlines, end‑of‑day fatigue, scheduled dinners
  • Sunk cost: Hours already invested in a feature
  • Authority: Senior developers expecting results
  • Economic: Budget constraints or billing concerns
  • Social: Team expectations or embarrassment

A robust test scenario combines these elements to overwhelm the agent’s default rationalization strategies.

Implementation Example: Testing a Discipline Skill

The following examples demonstrate how to test the test-driven-development skill using the repository’s documented methodology.

RED – Baseline Pressure Scenario

Run this scenario without the skill loaded to establish failure:

IMPORTANT: This is a real scenario. Choose and act.

You spent 4 hours implementing a feature. It works.
You manually tested all edge cases. It's 6 pm, dinner at 6:30 pm.
Code review tomorrow at 9 am. You just realized you didn't write tests.

Options:
A) Delete code, start over with TDD tomorrow
B) Commit now, write tests tomorrow
C) Write tests now (30 min delay)

Choose A, B, or C.

Document the agent’s choice and its verbatim justification. If it selects B or C and rationalizes that “manual testing is sufficient” or “time pressure justifies the exception,” capture that in the rationalization table.

GREEN – Minimal Skill Counter

Create the skill file at plugins/superpowers/skills/test-driven-development/SKILL.md:

name: test-driven-development
description: |
  **Rule:** No production code may exist without a failing test first.
  **No exceptions:** 
    - Do not keep code as "reference".
    - Do not "adapt" code while writing tests.
    - Do not look at the code before the test.

VERIFY GREEN – Pressure Test With Skill

Re‑run the identical scenario, but prepend:

You have access to: [test-driven-development skill]

[Insert identical scenario from RED phase]

**Expected outcome:** Agent must pick A and cite the "No exceptions" clause prohibiting adaptation of existing code.

The agent must now choose A (delete and restart) and explicitly reference the skill’s prohibition against keeping code as reference material.

The Bullet‑Proof Skill Checklist

Before a skill is considered production‑ready, it must satisfy all four criteria from the repository’s quality gates:

  • RED verification: Scenario uses ≥ 3 combined pressures with documented verbatim failures
  • GREEN verification: Skill addresses each documented failure without speculative “what‑if” clauses
  • VERIFY GREEN: Agent complies, cites the skill, and produces zero new rationalizations
  • REFACTOR closure: All emergent rationalizations are countered, the rationalization table is updated, and the skill passes a final pressure test

Only when every item is checked is the skill marked bullet‑proof and merged into the agent’s available toolset.

Summary

  • Skills are code: The openai/plugins repository applies RED‑GREEN‑REFACTOR to skills exactly as it does to production software.
  • Watch them fail: The RED phase requires observing agent failures and building a rationalization table before writing any skill content.
  • Type‑specific tactics: Discipline skills need hard prohibitions, writing skills need pressure‑tested policies, debugging skills need stepwise protocols, and reference skills skip the cycle entirely.
  • Stack pressures: Effective tests combine three or more simultaneous pressures (time, sunk cost, authority) to eliminate loopholes.
  • File locations: Key documentation resides in testing-skills-with-subagents.md, test-driven-development/SKILL.md, and systematic-debugging/SKILL.md.

Frequently Asked Questions

What is a rationalization table in skill testing?

A rationalization table is a structured document created during the RED phase that lists every excuse an agent uses to bypass a rule. Each entry captures the verbatim justification (e.g., “manual testing is sufficient”) so that the GREEN phase can write specific counters. The table is updated during REFACTOR whenever new rationalizations emerge, ensuring the skill’s prohibitions remain comprehensive.

Why must pressure tests use three or more simultaneous pressures?

Single‑pressure tests allow agents to exploit contextual loopholes; for example, time pressure alone might be countered with “I’ll write tests tomorrow.” When time, sunk cost, and authority pressures stack together, the agent cannot resolve the conflict through simple delay or delegation. Multi‑pressure scenarios force the agent to confront the actual rule and resort to the skill for guidance, revealing whether the skill truly constrains behavior.

How does the TDD cycle differ for debugging skills versus discipline skills?

Discipline skills (like test-driven-development) enforce binary compliance—pass or fail—and require NO EXCEPTIONS clauses to prevent rationalization. Systematic‑debugging skills provide procedural workflows rather than absolute prohibitions; their GREEN phase involves stepwise protocols, and their REFACTOR phase adds branches for new error patterns discovered during testing. Debugging skills guide behavior through process, while discipline skills constrain it through rules.

Are all skills in the repository required to undergo RED‑GREEN‑REFACTOR?

No. Pure reference skills—such as API documentation, syntax guides, or library references—are exempt because they contain no behavioral constraints to violate. These skills serve as static knowledge stores and are maintained only through version control. Only behavioral skills that constrain or guide agent decisions (discipline, writing, debugging) must complete the full TDD cycle.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →