How to Run Built-In Evaluations and Create Custom Eval Problems for ADHD
To run built-in evaluations in the ADHD framework, execute npm run evals for the full suite or npm run evals:quick for a sanity check; to create custom problems, define a new JSON object in bench/problems.json with id, description, and prompt fields, then run with --problem <your-id>.
The ADHD repository (UditAkhourii/adhd) ships with a self-contained evaluation suite that benchmarks the framework against engineering-style problems using a skeptical-staff-engineer judge. Whether you are validating changes to the core engine or measuring performance against custom scenarios, you can run built-in evaluations and create custom eval problems for ADHD through simple npm commands and JSON configuration.
Running the Built-In Evaluation Suite
The evaluation suite includes approximately six engineering problems that require roughly ten LLM calls each to complete. All commands execute locally on your machine without requiring external CI integration.
Full Benchmark Run
To execute the complete evaluation suite against all shipped problems:
npm run evals
This command generates two artifacts: a human-readable report in EVALS.md and a machine-readable transcript in bench/results.json.
Quick Sanity Check
For rapid validation during development, run only the first two problems:
npm run evals:quick
Targeting Specific Problems
To evaluate a single problem by its identifier (for example, the LRU cache test):
npm run evals -- --problem lru-100ms
The double dash (--) separates npm arguments from the script's own flags.
Understanding the Output Files
After any run, examine these generated files:
- EVALS.md – Aggregated scores and per-problem judgments in markdown format.
- bench/results.json – The complete LLM-to-LLM exchange transcript for programmatic analysis.
To update the published benchmark figures in the repository, run the suite locally and commit the regenerated EVALS.md file.
Creating Custom Evaluation Problems
Custom problems require no code changes—only a four-line edit to a JSON manifest. The evaluation engine in bench/run-evals.ts automatically detects new entries and executes the generation pass followed by the critic pass.
The Problem Manifest Structure
Each problem lives in bench/problems.json and follows this schema:
{
"id": "unique-identifier",
"description": "Brief natural-language statement of the task",
"prompt": "Full prompt text sent to the LLM",
"metadata": {
"tags": ["performance", "cache"],
"difficulty": "medium"
}
}
Required fields are id, description, and prompt. The metadata object is optional.
Step-by-Step: Adding a New Problem
- Open
bench/problems.jsonin your editor. - Append a new entry with a unique
id, cleardescription, and completeprompt. - Save the file.
Testing Your Custom Problem
Run your new problem in isolation to verify configuration:
npm run evals -- --problem <your-id>
The judge LLM scores the output alongside built-in problems using the dimensions defined in bench/judge.ts.
How the Evaluation Engine Works
According to the UditAkhourii/adhd source code, three core modules orchestrate the benchmarking process:
bench/run-evals.ts– Loadsproblems.json, iterates over selected problems, invokes the generation LLM viasrc/llm.ts, and triggers the judge pass.bench/judge.ts– Implements the "skeptical-staff-engineer" system prompt that scores outputs on breadth, novelty, trap detection, actionability, and builder usefulness.bench/baseline.ts– Provides the single-shot baseline LLM for head-to-head comparison against the ADHD framework's iterative approach.
These scripts interface with the core engine (src/engine.ts) which handles frame selection and result aggregation.
Summary
- Execute
npm run evalsto benchmark the full suite ornpm run evals:quickfor a truncated sanity check. - Target individual problems using
-- --problem <id>syntax. - Add custom evaluations by editing
bench/problems.jsonwith standard JSON objects containingid,description, andprompt. - Review human-readable results in
EVALS.mdand machine-readable data inbench/results.json. - The judge evaluates solutions on five dimensions: breadth, novelty, trap detection, actionability, and builder usefulness.
Frequently Asked Questions
Where does ADHD store evaluation results?
The framework writes human-readable verdicts to EVALS.md in the repository root and machine-readable transcripts to bench/results.json. These files are overwritten on each run.
Can I integrate ADHD evaluations into CI/CD pipelines?
Currently, the suite runs entirely on your local machine with no native CI integration. To publish updated benchmarks, run the suite locally and commit the regenerated EVALS.md file to version control.
What criteria does the judge use to score solutions?
As implemented in bench/judge.ts, the LLM-as-judge evaluates outputs across five dimensions: breadth of analysis, novelty of insights, trap detection, actionability of recommendations, and overall builder usefulness.
How do I compare ADHD against a simple baseline?
The evaluation suite automatically runs against the baseline defined in bench/baseline.ts, which provides a single-shot LLM implementation for head-to-head comparison with ADHD's iterative frame-selection approach.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →