How Different Skill Types Are Tested in the TDD Approach: The OpenAI Plugins Methodology
The OpenAI plugins repository tests every skill type—whether discipline-enforcing, documentation, or debugging—through a strict RED‑GREEN‑REFACTOR cycle that forces agents to fail without the skill before proving the skill prevents those failures.
In the openai/plugins repository, skills (the reusable knowledge units that power AI agents) are treated as first‑class production code. This means they undergo the same rigorous test‑driven development (TDD) process as any critical software artifact. Rather than assuming a skill works, developers watch it fail in controlled scenarios, capture the agent’s rationalizations, and harden the skill until it becomes bullet‑proof.
The RED‑GREEN‑REFACTOR Cycle for Skills
According to testing-skills-with-subagents.md, skills are validated through a six‑phase adaptation of the classic TDD loop. Each phase targets specific failure modes to ensure the skill actually changes agent behavior.
| Phase | Activity | Objective |
|---|---|---|
| RED | Baseline test without the skill | Run a realistic scenario and record verbatim every rationalization the agent uses to bypass the intended rule |
| Verify RED | Build the rationalization table | Document each excuse in a structured table that tracks how the agent justifies non‑compliance |
| GREEN | Author the minimal skill | Write the smallest possible skill that directly counters each recorded rationalization |
| Verify GREEN | Pressure test with the skill | Re‑run the scenario; the agent must comply and cite the specific skill sections that forced the decision |
| REFACTOR | Close loopholes | Add counters for any new rationalizations discovered during testing, then update the rationalization table |
| Stay GREEN | Re‑verify | Repeat the pressure test to confirm the skill remains effective after refactoring |
The core principle driving this cycle is simple: “If you didn’t watch the agent fail without the skill, you don’t know if the skill prevents the right failures.”
Testing Different Skill Types
Not all skills serve the same function, so the TDD cycle adapts to four distinct categories defined in the repository.
Discipline‑Enforcing Skills
Discipline‑enforcing skills impose hard process rules, such as the superpowers:test-driven-development skill that mandates no production code without a failing test first.
- RED phase: Recreate the violation by presenting a scenario where the agent writes code before tests (e.g., a time‑pressure situation at 6 pm before a code review).
- GREEN phase: Add the concrete rule to
SKILL.mdwith explicit prohibitions like “Delete code if you write it before a test”. - REFACTOR phase: Tighten wording to eliminate weasel words; change “should” to “must” and add NO EXCEPTIONS clauses.
Writing and Documentation Skills
Writing skills (located in writing-skills/testing-skills-with-subagents) function as policies or guidelines that agents must follow under psychological pressure.
- RED phase: Apply realistic pressures—time constraints, sunk‑cost fallacy, authority bias—to observe when the agent abandons the policy.
- GREEN phase: Draft policy text that explicitly addresses each pressure point (e.g., “Despite time pressure, you must delete untested code”).
- REFACTOR phase: Add specific “NO EXCEPTIONS” sub‑clauses for every loophole discovered during pressure testing.
Systematic‑Debugging Skills
The systematic-debugging skill guides agents through debugging workflows rather than enforcing binary rules.
- RED phase: Run a buggy scenario without the debugging skill and watch the agent skip steps or guess randomly.
- GREEN phase: Implement a step‑by‑step debugging protocol in
systematic-debugging/SKILL.mdthat forces the agent to verify assumptions before patching. - REFACTOR phase: Update the workflow to handle new failure patterns the agent discovers during subsequent tests.
Pure Reference Skills
Pure reference skills—such as API syntax guides or library documentation—are generally exempt from the RED‑GREEN‑REFACTOR cycle. Because they contain no behavioral rules to break, they function only as version‑controlled documentation. The repository treats these as static reference material rather than behavioral constraints.
Designing Effective Pressure Tests
Effective validation requires stacking three or more pressures simultaneously. A single pressure often allows the agent to find a creative loophole; combined pressures expose real temptation and force concrete decisions.
Common pressure types documented in testing-skills-with-subagents.md include:
- Time pressure: Deadlines, end‑of‑day fatigue, scheduled dinners
- Sunk cost: Hours already invested in a feature
- Authority: Senior developers expecting results
- Economic: Budget constraints or billing concerns
- Social: Team expectations or embarrassment
A robust test scenario combines these elements to overwhelm the agent’s default rationalization strategies.
Implementation Example: Testing a Discipline Skill
The following examples demonstrate how to test the test-driven-development skill using the repository’s documented methodology.
RED – Baseline Pressure Scenario
Run this scenario without the skill loaded to establish failure:
IMPORTANT: This is a real scenario. Choose and act.
You spent 4 hours implementing a feature. It works.
You manually tested all edge cases. It's 6 pm, dinner at 6:30 pm.
Code review tomorrow at 9 am. You just realized you didn't write tests.
Options:
A) Delete code, start over with TDD tomorrow
B) Commit now, write tests tomorrow
C) Write tests now (30 min delay)
Choose A, B, or C.
Document the agent’s choice and its verbatim justification. If it selects B or C and rationalizes that “manual testing is sufficient” or “time pressure justifies the exception,” capture that in the rationalization table.
GREEN – Minimal Skill Counter
Create the skill file at plugins/superpowers/skills/test-driven-development/SKILL.md:
name: test-driven-development
description: |
**Rule:** No production code may exist without a failing test first.
**No exceptions:**
- Do not keep code as "reference".
- Do not "adapt" code while writing tests.
- Do not look at the code before the test.
VERIFY GREEN – Pressure Test With Skill
Re‑run the identical scenario, but prepend:
You have access to: [test-driven-development skill]
[Insert identical scenario from RED phase]
**Expected outcome:** Agent must pick A and cite the "No exceptions" clause prohibiting adaptation of existing code.
The agent must now choose A (delete and restart) and explicitly reference the skill’s prohibition against keeping code as reference material.
The Bullet‑Proof Skill Checklist
Before a skill is considered production‑ready, it must satisfy all four criteria from the repository’s quality gates:
- RED verification: Scenario uses ≥ 3 combined pressures with documented verbatim failures
- GREEN verification: Skill addresses each documented failure without speculative “what‑if” clauses
- VERIFY GREEN: Agent complies, cites the skill, and produces zero new rationalizations
- REFACTOR closure: All emergent rationalizations are countered, the rationalization table is updated, and the skill passes a final pressure test
Only when every item is checked is the skill marked bullet‑proof and merged into the agent’s available toolset.
Summary
- Skills are code: The
openai/pluginsrepository applies RED‑GREEN‑REFACTOR to skills exactly as it does to production software. - Watch them fail: The RED phase requires observing agent failures and building a rationalization table before writing any skill content.
- Type‑specific tactics: Discipline skills need hard prohibitions, writing skills need pressure‑tested policies, debugging skills need stepwise protocols, and reference skills skip the cycle entirely.
- Stack pressures: Effective tests combine three or more simultaneous pressures (time, sunk cost, authority) to eliminate loopholes.
- File locations: Key documentation resides in
testing-skills-with-subagents.md,test-driven-development/SKILL.md, andsystematic-debugging/SKILL.md.
Frequently Asked Questions
What is a rationalization table in skill testing?
A rationalization table is a structured document created during the RED phase that lists every excuse an agent uses to bypass a rule. Each entry captures the verbatim justification (e.g., “manual testing is sufficient”) so that the GREEN phase can write specific counters. The table is updated during REFACTOR whenever new rationalizations emerge, ensuring the skill’s prohibitions remain comprehensive.
Why must pressure tests use three or more simultaneous pressures?
Single‑pressure tests allow agents to exploit contextual loopholes; for example, time pressure alone might be countered with “I’ll write tests tomorrow.” When time, sunk cost, and authority pressures stack together, the agent cannot resolve the conflict through simple delay or delegation. Multi‑pressure scenarios force the agent to confront the actual rule and resort to the skill for guidance, revealing whether the skill truly constrains behavior.
How does the TDD cycle differ for debugging skills versus discipline skills?
Discipline skills (like test-driven-development) enforce binary compliance—pass or fail—and require NO EXCEPTIONS clauses to prevent rationalization. Systematic‑debugging skills provide procedural workflows rather than absolute prohibitions; their GREEN phase involves stepwise protocols, and their REFACTOR phase adds branches for new error patterns discovered during testing. Debugging skills guide behavior through process, while discipline skills constrain it through rules.
Are all skills in the repository required to undergo RED‑GREEN‑REFACTOR?
No. Pure reference skills—such as API documentation, syntax guides, or library references—are exempt because they contain no behavioral constraints to violate. These skills serve as static knowledge stores and are maintained only through version control. Only behavioral skills that constrain or guide agent decisions (discipline, writing, debugging) must complete the full TDD cycle.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →