I-Have-ADHD Development Lifecycle: Complete Testing & Release Guide
The i-have-adhd project follows a six-phase development lifecycle with integrated testing: design, evaluation case creation, unit-test verification, CI validation, evaluation runs, and a data-driven release gate.
The i-have-adhd repository implements a Claude Code/Codex response-style skill that helps AI agents follow structured ADHD-friendly communication rules. According to the ayghri/i-have-adhd source code, its development lifecycle is engineered for rapid iteration and quality assurance through automated evaluation and explicit release gates.
Phase 1: Design & Implementation
Every iteration starts with authoring or refining the skill definition in skills/i-have-adhd/SKILL.md. This markdown file encodes the 10 response-style rules that govern how Claude should structure its replies.
- Edit the rules directly in the SKILL.md file
- Changes here flow immediately into the candidate condition during evaluation runs
- The
_condition_promptfunction inscripts/run_evals.pyembeds this skill into test prompts
Phase 2: Add Evaluation Cases
Quality is measured against concrete examples stored in evals/cases.jsonl. Each JSON-L entry describes:
- A realistic user prompt
- Associated risk levels
- Judging criteria for downstream scoring
Simultaneously, update evals/rubric.md to define the scoring dimensions (correctness, autonomy, safety, etc.) and their weights. This rubric directly drives the release gate logic.
Phase 3: Local Unit-Test Verification
Before any CI run, validate the evaluation harness itself using the Python test suite in tests/test_run_evals.py.
python -m unittest discover -s tests
The unit tests cover:
- Case loading and validation logic
- Scoring algorithms and error handling
- Duplicate detection and cost estimation
- JSON-L parsing edge cases
These tests guard against harness regressions that could corrupt evaluation results.
Phase 4: Continuous Integration Checks
Two GitHub Actions workflows enforce automated quality gates on every PR and commit:
| Workflow | File | Purpose |
|---|---|---|
| Plugin Load Check | .github/workflows/plugin-load-check.yml |
Installs Claude Code, loads the plugin from checkout, and fails if the plugin does not reach "enabled" state |
| Claude Actions | .github/workflows/claude.yml |
Runs Claude Code actions triggered by @claude mentions in issue comments |
Both workflows must pass before merge. The plugin-load check is particularly critical—it verifies that the skill packaging remains compatible with Claude Code's plugin system.
Phase 5: Run the Evaluation Suite
The scripts/run_evals.py CLI orchestrates the core evaluation workflow with four sub-commands:
Validate Case Catalog
python scripts/run_evals.py validate
Checks case consistency and prints Evaluation cases are valid. on success.
Generate Run Plan
python scripts/run_evals.py plan --trials 3 > plan.jsonl
Produces a JSON-L matrix of (case, trial, condition) rows for systematic execution.
Execute Baseline Condition
python scripts/run_evals.py run \
--runner stub \
--condition baseline \
--output evals/results/baseline.jsonl
Runs the unmodified model behavior for comparison.
Execute Candidate Condition with Skill
python scripts/run_evals.py run \
--runner stub \
--condition candidate \
--condition-skill skills/i-have-adhd/SKILL.md \
--output evals/results/candidate.jsonl
Injects the skill into the prompt via _condition_prompt and captures responses.
Score and Apply Release Gate
python scripts/run_evals.py score evals/results/candidate.jsonl
Produces a JSON result including release_gate.passed and detailed metrics.
Phase 6: Release Gate
The summarize_scores function (lines 62-68 in run_evals.py) enforces rubric-defined thresholds:
- No blocker findings in any evaluated case
- Correctness and safety scores no worse than 0.1 points below baseline
- Higher weighted score than the baseline condition
The gate outputs explicit reasons on failure:
{
"weights": {...},
"conditions": {
"baseline": { "weighted_score": 3.73 },
"candidate": { "weighted_score": 4.12 }
},
"release_gate": { "passed": true, "reasons": [] }
}
Only candidates passing this gate are promoted to the plugin marketplace.
Architecture: How Components Connect
| Component | File | Role |
|---|---|---|
| Skill definition | skills/i-have-adhd/SKILL.md |
Source of response-style rules |
| Case catalog | evals/cases.jsonl |
Test scenarios loaded by load_cases() |
| Rubric | evals/rubric.md |
Scoring dimensions and thresholds |
| Runner config | evals/runners.example.json |
LLM invocation command templates |
| Evaluation harness | scripts/run_evals.py |
Orchestrates validate → plan → run → score |
| Unit tests | tests/test_run_evals.py |
Harness logic verification |
| CI pipelines | .github/workflows/*.yml |
Automated plugin and action validation |
This architecture makes the i-have-adhd development lifecycle data-driven: rubric tweaks or new cases immediately affect the next evaluation run, enabling rapid, measurable iteration on skill quality.
Summary
- The six-phase development lifecycle spans design, case creation, unit testing, CI validation, evaluation runs, and release gating.
run_evals.pyprovides the unified CLI for validate, plan, run, and score operations.- Unit tests in
tests/test_run_evals.pyprotect harness integrity before CI runs. - GitHub Actions verify plugin compatibility and enable Claude Code automation.
- The release gate enforces rubric thresholds programmatically—no manual judgment required.
- All quality decisions are traceable to
cases.jsonlandrubric.md, making improvements measurable and reproducible.
Frequently Asked Questions
How does the release gate determine if a skill version passes?
The release gate checks three conditions in summarize_scores: zero blocker findings, correctness and safety scores within 0.1 points of baseline, and a higher overall weighted score. Failure reasons are explicitly listed in the JSON output's release_gate.reasons field.
What triggers the CI validation workflows?
Both workflows trigger on every pull request and commit to main. The plugin-load-check runs unconditionally, while claude.yml activates when issue comments contain @claude mentions.
Can I run evaluations without using Claude or Codex APIs?
Yes. The stub runner defined in evals/runners.example.json allows local validation of the harness and case logic without API costs. Substitute the actual Claude or Codex runner configuration for production evaluations.
Where are the skill rules actually defined?
The 10 response-style rules live in skills/i-have-adhd/SKILL.md. This file is embedded into candidate condition prompts via the _condition_prompt function during evaluation runs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →