# How to Run Built-In Evaluations and Create Custom Eval Problems for ADHD

> Learn how to run built-in ADHD evaluations using npm commands or create custom eval problems by editing the problems JSON file. Get started with ADHD evaluations today.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: how-to-guide
- Published: 2026-08-01

---

**To run built-in evaluations in the ADHD framework, execute `npm run evals` for the full suite or `npm run evals:quick` for a sanity check; to create custom problems, define a new JSON object in [`bench/problems.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/problems.json) with `id`, `description`, and `prompt` fields, then run with `--problem <your-id>`.**

The ADHD repository (UditAkhourii/adhd) ships with a self-contained evaluation suite that benchmarks the framework against engineering-style problems using a skeptical-staff-engineer judge. Whether you are validating changes to the core engine or measuring performance against custom scenarios, you can run built-in evaluations and create custom eval problems for ADHD through simple npm commands and JSON configuration.

## Running the Built-In Evaluation Suite

The evaluation suite includes approximately six engineering problems that require roughly ten LLM calls each to complete. All commands execute locally on your machine without requiring external CI integration.

### Full Benchmark Run

To execute the complete evaluation suite against all shipped problems:

```bash
npm run evals

```

This command generates two artifacts: a human-readable report in [`EVALS.md`](https://github.com/UditAkhourii/adhd/blob/main/EVALS.md) and a machine-readable transcript in [`bench/results.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/results.json).

### Quick Sanity Check

For rapid validation during development, run only the first two problems:

```bash
npm run evals:quick

```

### Targeting Specific Problems

To evaluate a single problem by its identifier (for example, the LRU cache test):

```bash
npm run evals -- --problem lru-100ms

```

The double dash (`--`) separates npm arguments from the script's own flags.

### Understanding the Output Files

After any run, examine these generated files:

- **EVALS.md** – Aggregated scores and per-problem judgments in markdown format.
- **bench/results.json** – The complete LLM-to-LLM exchange transcript for programmatic analysis.

To update the published benchmark figures in the repository, run the suite locally and commit the regenerated [`EVALS.md`](https://github.com/UditAkhourii/adhd/blob/main/EVALS.md) file.

## Creating Custom Evaluation Problems

Custom problems require no code changes—only a four-line edit to a JSON manifest. The evaluation engine in [`bench/run-evals.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/run-evals.ts) automatically detects new entries and executes the generation pass followed by the critic pass.

### The Problem Manifest Structure

Each problem lives in [`bench/problems.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/problems.json) and follows this schema:

```json
{
  "id": "unique-identifier",
  "description": "Brief natural-language statement of the task",
  "prompt": "Full prompt text sent to the LLM",
  "metadata": {
    "tags": ["performance", "cache"],
    "difficulty": "medium"
  }
}

```

Required fields are `id`, `description`, and `prompt`. The `metadata` object is optional.

### Step-by-Step: Adding a New Problem

1. Open [`bench/problems.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/problems.json) in your editor.
2. Append a new entry with a unique `id`, clear `description`, and complete `prompt`.
3. Save the file.

### Testing Your Custom Problem

Run your new problem in isolation to verify configuration:

```bash
npm run evals -- --problem <your-id>

```

The judge LLM scores the output alongside built-in problems using the dimensions defined in [`bench/judge.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/judge.ts).

## How the Evaluation Engine Works

According to the UditAkhourii/adhd source code, three core modules orchestrate the benchmarking process:

- **[`bench/run-evals.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/run-evals.ts)** – Loads [`problems.json`](https://github.com/UditAkhourii/adhd/blob/main/problems.json), iterates over selected problems, invokes the generation LLM via [`src/llm.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/llm.ts), and triggers the judge pass.
- **[`bench/judge.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/judge.ts)** – Implements the "skeptical-staff-engineer" system prompt that scores outputs on **breadth**, **novelty**, **trap detection**, **actionability**, and **builder usefulness**.
- **[`bench/baseline.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/baseline.ts)** – Provides the single-shot baseline LLM for head-to-head comparison against the ADHD framework's iterative approach.

These scripts interface with the core engine ([`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts)) which handles frame selection and result aggregation.

## Summary

- Execute `npm run evals` to benchmark the full suite or `npm run evals:quick` for a truncated sanity check.
- Target individual problems using `-- --problem <id>` syntax.
- Add custom evaluations by editing [`bench/problems.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/problems.json) with standard JSON objects containing `id`, `description`, and `prompt`.
- Review human-readable results in [`EVALS.md`](https://github.com/UditAkhourii/adhd/blob/main/EVALS.md) and machine-readable data in [`bench/results.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/results.json).
- The judge evaluates solutions on five dimensions: breadth, novelty, trap detection, actionability, and builder usefulness.

## Frequently Asked Questions

### Where does ADHD store evaluation results?

The framework writes human-readable verdicts to [`EVALS.md`](https://github.com/UditAkhourii/adhd/blob/main/EVALS.md) in the repository root and machine-readable transcripts to [`bench/results.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/results.json). These files are overwritten on each run.

### Can I integrate ADHD evaluations into CI/CD pipelines?

Currently, the suite runs entirely on your local machine with no native CI integration. To publish updated benchmarks, run the suite locally and commit the regenerated [`EVALS.md`](https://github.com/UditAkhourii/adhd/blob/main/EVALS.md) file to version control.

### What criteria does the judge use to score solutions?

As implemented in [`bench/judge.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/judge.ts), the LLM-as-judge evaluates outputs across five dimensions: **breadth** of analysis, **novelty** of insights, **trap detection**, **actionability** of recommendations, and overall **builder usefulness**.

### How do I compare ADHD against a simple baseline?

The evaluation suite automatically runs against the baseline defined in [`bench/baseline.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/baseline.ts), which provides a single-shot LLM implementation for head-to-head comparison with ADHD's iterative frame-selection approach.