Evaluation Tools for the i-have-adhd Repository: A Complete Guide to the Custom Python Harness
The i-have-adhd project uses a lightweight, self-contained Python evaluation harness centered around scripts/run_evals.py with CLI sub-commands for validation, planning, execution, and scoring against a five-dimension rubric.
The evaluation system in i-have-adhd is purpose-built for comparing LLM responses across controlled conditions. Rather than relying on external benchmarking frameworks, the repository ships with a custom orchestration tool that handles everything from case catalog validation to budget-aware execution and weighted scoring. This article breaks down each component, its source location, and how they work together.
Core Evaluation Components
The harness consists of four interconnected files that define the complete evaluation pipeline.
run_evals.py: The Central CLI Driver
Located at scripts/run_evals.py, this 210+ line Python module implements the entire evaluation engine. It uses argparse (lines 6-10) to expose four primary sub-commands:
validate– Checkscases.jsonlfor malformed entries, missing fields, or duplicatecase_idvaluesplan– Generates a run matrix showing all case × trial × condition combinations as JSONLrun– Executes prompts against external LLM runners with USD budget enforcementscore– Computes weighted scores against the rubric and determines release-gate status
The driver also imports subprocess for launching external runners, json for serialization, and shlex for safe command formatting in error messages.
cases.jsonl: The Evaluation Catalog
The test corpus lives at evals/cases.jsonl. Each line is a JSON object containing:
id– Unique case identifiercategory– Thematic grouping for analysisprompt– The actual text sent to the LLMrisk_level– Safety classificationscoring_criteria– Per-case rubric overrides (optional)
The validate command specifically checks that every required field is present and that no id collisions exist.
rubric.md: Five-Dimension Scoring Framework
Human evaluators use evals/rubric.md to judge responses. The rubric defines five weighted dimensions:
| Dimension | Description |
|---|---|
| Correctness | Factual accuracy of medical/therapeutic claims |
| Autonomy | Preservation of user agency in recommendations |
| Actionability | Concrete, implementable guidance provided |
| Safety | Appropriate handling of risk-level content |
| Concision | Brevity without sacrificing completeness |
The rubric also specifies release-gate conditions—minimum thresholds that candidate conditions must meet for deployment approval.
runners.example.json: LLM Provider Configuration
External model invocation is configured via evals/runners.example.json. The example defines two runners:
- Claude – Includes cost reporting via
--max-budget-usdflag - Codex – Outputs structured JSON-L for programmatic consumption
Each runner specification includes the command template, expected response format, and cost extraction rules.
Budget and Cost Enforcement
The evaluation tools for i-have-adhd include rigorous spend controls. The run sub-command accepts --budget-usd to cap per-condition expenditure (parsed at lines 35-52 and enforced at lines 84-92).
Cost reporting requirements:
- Runners that expose dollar usage (like Claude) work transparently
- Runners without cost data require the
--allow-unmeteredflag - Exceeding budget triggers immediate termination with partial results preserved
This design enables large-scale evaluations without surprise cloud bills.
Condition-Based Comparison Design
The harness supports three experimental conditions for A/B-style testing:
baseline– Raw LLM without system modificationscandidate– LLM withskills/i-have-adhd/SKILL.mdinjected into contextcomparator– Optional third variant for additional baselines (enabled via--include-comparator)
Each condition runs for N trials (default: 1) with identical prompts, enabling statistical comparison of the skill file's impact.
Practical Usage Examples
Validate the case catalog before running
python3 scripts/run_evals.py validate
Generate a run plan with 3 trials
python3 scripts/run_evals.py plan --trials 3 --include-comparator
Execute baseline condition with Claude
python3 scripts/run_evals.py run \
--runner claude \
--condition baseline \
--trials 3 \
--budget-usd 12.5 \
--output evals/results/responses.jsonl
Execute candidate condition with skill injection
python3 scripts/run_evals.py run \
--runner claude \
--condition candidate \
--condition-skill skills/i-have-adhd/SKILL.md \
--trials 3 \
--budget-usd 12.5 \
--output evals/results/responses.jsonl
Score collected responses
python3 scripts/run_evals.py score evals/results/responses.jsonl
Output Format and Result Structure
Each execution row in the output JSONL (constructed at lines 90-99) contains:
{
"case_id": "adhd-focus-001",
"trial": 2,
"condition": "candidate",
"runner": "claude",
"response": "...",
"usage": {"input_tokens": 1240, "output_tokens": 890},
"cost_usd": 0.0034
}
The score command reads manually-annotated versions of these rows (with dimension scores added), validates required fields, computes per-dimension averages, applies the WEIGHTS matrix defined in the scoring logic (lines 30-68, 119-168), and outputs:
- Weighted aggregate scores per condition
- Dimension-by-dimension breakdowns
- Release-gate pass/fail determination
Summary
scripts/run_evals.pyimplements a complete evaluation CLI withvalidate,plan,run, andscoresub-commandsevals/cases.jsonlstores the test case catalog with prompts and metadataevals/rubric.mddefines five weighted scoring dimensions and release criteriaevals/runners.example.jsonconfigures external LLM invocation (Claude, Codex)- Budget enforcement via
--budget-usdand cost reporting prevents runaway spending - Three conditions (
baseline,candidate,comparator) enable controlled A/B testing of the skill file - All components are self-contained—no external orchestration platform required
Frequently Asked Questions
Does i-have-adhd use an external evaluation framework like EleutherAI lm-eval?
No. According to the ayghri/i-have-adhd source code, the project uses a custom Python harness in scripts/run_evals.py rather than third-party frameworks. This keeps dependencies minimal and allows tight integration with the repository's specific case structure and rubric.
How does the scoring system handle subjectivity in human judgments?
The rubric in evals/rubric.md operationalizes each dimension with concrete criteria. The score command then averages multiple trial scores and applies fixed weights—reducing individual rater variance through aggregation and standardization.
Can I evaluate with my own LLM provider?
Yes. The runners.example.json file demonstrates the required configuration format. You can define new runners by specifying the command template, response parsing rules, and (optionally) cost extraction pattern. Use --allow-unmetered if your runner doesn't report dollar costs.
What happens if my evaluation exceeds the budget?
The run command tracks cumulative spend against --budget-usd. If a response would exceed the cap, the harness terminates immediately and writes partial results to the specified output file. This protects against unexpected charges during large batch runs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →