How to Test and Validate the Behavior of i-have-adhd: A Complete Guide
Run the two bundled test suites—tests/test_run_evals.py for evaluation logic and tests/test_always_on_hooks.py for cross-platform hook behavior—then use the CLI commands in scripts/run_evals.py to validate the case catalog and verify release gates.
The i-have-adhd repository by ayghri provides a structured evaluation framework for Claude-style coding assistants. To ensure the skill performs reliably across different scenarios and platforms, the codebase includes comprehensive testing utilities that validate everything from JSON case definitions to shell hook injections. Understanding how to test and validate the behavior of i-have-adhd ensures that any modifications to the skill, evaluation weights, or hook logic maintain the project's quality standards.
Architecture of the i-have-adhd Test Framework
The testing infrastructure splits into two distinct domains: the evaluation pipeline that validates LLM performance against defined cases, and the always-on hook system that manages runtime skill injection.
Evaluation Pipeline Components
The evaluation system relies on a JSON-Lines case catalog and several Python utilities in scripts/run_evals.py:
load_cases– Readsevals/cases.jsonlinto memory for processingvalidate_cases(lines 59-76) – Verifies that every case contains required fields, enforces unique IDs, and confirms risk values are constrained tolow,medium, orhighrun_evaluations(lines 9-42) – Orchestrates LLM execution with budget constraints, retry logic, and unmetered mode supportsummarize_scores(lines 30-68) – Aggregates manually-judged results, applies weighted metrics via theWEIGHTSmapping, and computes therelease_gatestatus (lines 55-62)
The Always-On Hook System
The always-on hook provides cross-platform banner injection through runtime-specific scripts (hooks/always-on.sh, always-on.mjs, always-on.ps1). When the hidden flag file .i-have-adhd-always exists in the configuration directory, the hook strips YAML front-matter from skills/i-have-adhd/SKILL.md and outputs only the skill body.
Running the Automated Test Suites
The repository includes two primary test files that cover distinct validation surfaces.
Evaluation Logic Validation
The tests/test_run_evals.py suite verifies case loading, weighting algorithms, release-gate logic, and error handling. To execute these tests:
python -m unittest tests/test_run_evals.py
This validates that validate_cases correctly identifies malformed entries and that summarize_scores properly pairs baseline and candidate rows before applying weighted calculations.
Cross-Platform Hook Testing
The tests/test_always_on_hooks.py suite requires Node, POSIX shell, and PowerShell on your PATH. It verifies three critical behaviors:
- Silent by default – The hook produces no output when the
.i-have-adhd-alwaysflag is absent - Front-matter stripping – Correctly removes YAML delimiters (
---) even with trailing whitespace - Unclosed fence handling – Preserves content when front-matter delimiters are unclosed
Run these tests with:
python -m unittest tests/test_always_on_hooks.py
Manual Validation Workflow
Beyond automated tests, the CLI provides granular control over the validation process.
Validating the Case Catalog
Before executing any evaluations, verify the integrity of evals/cases.jsonl:
python scripts/run_evals.py validate
Success prints "Evaluation cases are valid." Failures emit ERROR: lines indicating missing fields, duplicate IDs, or invalid risk classifications.
Executing Evaluation Runs
To test the evaluation pipeline with the stub runner:
python scripts/run_evals.py run \
--runner stub \
--condition baseline \
--output /tmp/baseline.out.jsonl \
--budget-usd 5
This command respects the USD budget, implements retry logic for failures, and writes one JSON line per completed trial to the specified output file.
Scoring and Release Gate Verification
After manually judging the output file for correctness and autonomy, compute the final scores:
python scripts/run_evals.py score /tmp/baseline.out.jsonl
The resulting JSON includes a release_gate object indicating whether the candidate passes based on weighted score regressions and blocker flags.
Programmatic Testing Examples
Validating Cases in Python
Import the evaluation utilities directly to test catalog integrity within your own scripts:
from scripts.run_evals import load_cases, validate_cases
cases = load_cases()
errors = validate_cases(cases)
assert not errors, f"Catalog validation failed: {errors}"
print("✅ All evaluation cases are valid")
Testing Score Summarization
Verify the scoring logic programmatically by loading judged result files:
import json
from pathlib import Path
from scripts.run_evals import summarize_scores, read_jsonl
score_rows = read_jsonl(Path("sample_scores.jsonl"))
summary = summarize_scores(score_rows)
print(json.dumps(summary, indent=2))
# Output includes release_gate.passed boolean
Local Hook Testing (POSIX)
Test the always-on hook behavior directly without the full test suite:
export CLAUDE_CONFIG_DIR=$(mktemp -d)
touch "$CLAUDE_CONFIG_DIR/.i-have-adhd-always"
sh hooks/always-on.sh
# Outputs SKILL.md content without YAML front-matter
Remove the flag file to verify the hook remains silent when users have not opted in.
Summary
- The evaluation framework in
scripts/run_evals.pyprovidesvalidate_casesfor catalog integrity andsummarize_scoresfor weighted performance analysis - Two test suites cover distinct surfaces:
tests/test_run_evals.pyfor evaluation logic andtests/test_always_on_hooks.pyfor cross-platform hook behavior - Release gate logic (lines 55-62 of
summarize_scores) automatically determines if a candidate passes based on weighted metrics and blocker flags - The always-on hook requires the hidden file
.i-have-adhd-alwaysto activate and strips YAML front-matter fromskills/i-have-adhd/SKILL.md - Validation commands include
validate,run, andscoresubcommands to check the case catalog, execute trials within budget constraints, and judge results
Frequently Asked Questions
What files contain the core test logic for i-have-adhd?
The primary test implementations reside in tests/test_run_evals.py and tests/test_always_on_hooks.py. The first validates the evaluation pipeline including case loading and scoring algorithms, while the second verifies that the always-on hook behaves correctly across Node, POSIX shell, and PowerShell environments.
How do I validate the evaluation case catalog before running tests?
Run python scripts/run_evals.py validate to execute the validate_cases function against evals/cases.jsonl. This checks for required fields, unique identifiers, and valid risk classifications (low, medium, or high), printing specific error messages for any malformed entries.
What is the release gate in i-have-adhd testing?
The release gate is a boolean flag computed by summarize_scores in scripts/run_evals.py (lines 55-62). It evaluates whether weighted scores meet baseline thresholds and checks for blocker flags. The gate fails if performance regresses beyond tolerance limits or if critical errors are detected in the evaluation pairs.
How does the always-on hook detect user opt-in?
The hook searches for a hidden flag file named .i-have-adhd-always in the CLAUDE_CONFIG_DIR directory. When present, the hook reads skills/i-have-adhd/SKILL.md, strips the YAML front-matter delimited by ---, and injects the skill body into the assistant context. Without this flag, the hook produces no output.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →