How to Run Built-In Evaluations and Create Custom Eval Problems for ADHD

To run built-in evaluations in the ADHD framework, execute npm run evals for the full suite or npm run evals:quick for a sanity check; to create custom problems, define a new JSON object in bench/problems.json with id, description, and prompt fields, then run with --problem <your-id>.

The ADHD repository (UditAkhourii/adhd) ships with a self-contained evaluation suite that benchmarks the framework against engineering-style problems using a skeptical-staff-engineer judge. Whether you are validating changes to the core engine or measuring performance against custom scenarios, you can run built-in evaluations and create custom eval problems for ADHD through simple npm commands and JSON configuration.

Running the Built-In Evaluation Suite

The evaluation suite includes approximately six engineering problems that require roughly ten LLM calls each to complete. All commands execute locally on your machine without requiring external CI integration.

Full Benchmark Run

To execute the complete evaluation suite against all shipped problems:

npm run evals

This command generates two artifacts: a human-readable report in EVALS.md and a machine-readable transcript in bench/results.json.

Quick Sanity Check

For rapid validation during development, run only the first two problems:

npm run evals:quick

Targeting Specific Problems

To evaluate a single problem by its identifier (for example, the LRU cache test):

npm run evals -- --problem lru-100ms

The double dash (--) separates npm arguments from the script's own flags.

Understanding the Output Files

After any run, examine these generated files:

  • EVALS.md – Aggregated scores and per-problem judgments in markdown format.
  • bench/results.json – The complete LLM-to-LLM exchange transcript for programmatic analysis.

To update the published benchmark figures in the repository, run the suite locally and commit the regenerated EVALS.md file.

Creating Custom Evaluation Problems

Custom problems require no code changes—only a four-line edit to a JSON manifest. The evaluation engine in bench/run-evals.ts automatically detects new entries and executes the generation pass followed by the critic pass.

The Problem Manifest Structure

Each problem lives in bench/problems.json and follows this schema:

{
  "id": "unique-identifier",
  "description": "Brief natural-language statement of the task",
  "prompt": "Full prompt text sent to the LLM",
  "metadata": {
    "tags": ["performance", "cache"],
    "difficulty": "medium"
  }
}

Required fields are id, description, and prompt. The metadata object is optional.

Step-by-Step: Adding a New Problem

  1. Open bench/problems.json in your editor.
  2. Append a new entry with a unique id, clear description, and complete prompt.
  3. Save the file.

Testing Your Custom Problem

Run your new problem in isolation to verify configuration:

npm run evals -- --problem <your-id>

The judge LLM scores the output alongside built-in problems using the dimensions defined in bench/judge.ts.

How the Evaluation Engine Works

According to the UditAkhourii/adhd source code, three core modules orchestrate the benchmarking process:

  • bench/run-evals.ts – Loads problems.json, iterates over selected problems, invokes the generation LLM via src/llm.ts, and triggers the judge pass.
  • bench/judge.ts – Implements the "skeptical-staff-engineer" system prompt that scores outputs on breadth, novelty, trap detection, actionability, and builder usefulness.
  • bench/baseline.ts – Provides the single-shot baseline LLM for head-to-head comparison against the ADHD framework's iterative approach.

These scripts interface with the core engine (src/engine.ts) which handles frame selection and result aggregation.

Summary

  • Execute npm run evals to benchmark the full suite or npm run evals:quick for a truncated sanity check.
  • Target individual problems using -- --problem <id> syntax.
  • Add custom evaluations by editing bench/problems.json with standard JSON objects containing id, description, and prompt.
  • Review human-readable results in EVALS.md and machine-readable data in bench/results.json.
  • The judge evaluates solutions on five dimensions: breadth, novelty, trap detection, actionability, and builder usefulness.

Frequently Asked Questions

Where does ADHD store evaluation results?

The framework writes human-readable verdicts to EVALS.md in the repository root and machine-readable transcripts to bench/results.json. These files are overwritten on each run.

Can I integrate ADHD evaluations into CI/CD pipelines?

Currently, the suite runs entirely on your local machine with no native CI integration. To publish updated benchmarks, run the suite locally and commit the regenerated EVALS.md file to version control.

What criteria does the judge use to score solutions?

As implemented in bench/judge.ts, the LLM-as-judge evaluates outputs across five dimensions: breadth of analysis, novelty of insights, trap detection, actionability of recommendations, and overall builder usefulness.

How do I compare ADHD against a simple baseline?

The evaluation suite automatically runs against the baseline defined in bench/baseline.ts, which provides a single-shot LLM implementation for head-to-head comparison with ADHD's iterative frame-selection approach.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →