# ADHD Evaluation Methodology: How the Engine Measures Performance Gains Against Single-Shot LLMs

> Discover ADHD's A/B benchmark for evaluating performance gains. Learn how this reproducible methodology measures improvements against single-shot LLMs on complex engineering tasks.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: performance
- Published: 2026-08-01

---

**ADHD uses a reproducible A/B benchmark that compares its multi-frame ideation engine against a single-shot baseline on six curated engineering problems, measuring improvements through randomized blind evaluation.**

The evaluation methodology in the UditAkhourii/adhd repository establishes a rigorous, end-to-end testing framework that quantifies how the ADHD engine’s structured ideation process outperforms standard single-shot prompting. This systematic approach ensures that performance claims are backed by reproducible experiments across diverse software engineering domains.

## Benchmark Design and Problem Curation

The foundation of ADHD’s evaluation rests on a curated problem set defined in [`bench/problems.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/problems.json). This JSON file specifies six distinct open-ended engineering challenges spanning multiple domains, including **LRU cache design**, **CLI hang handling**, **rate-limiting implementations**, **debugging scenarios**, **monolith splitting strategies**, and **naming conventions for feature-flag services**.

## Comparative Generation Protocol

### ADHD Engine Configuration

The ADHD engine executes with specific hyperparameters orchestrated through [`bench/run-evals.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/run-evals.ts). The system invokes the `run()` function with a configuration object specifying `framesPerRun: 5`, `ideasPerFrame: 6`, and `topK: 3`, enabling the engine to generate five ideation frames, produce six ideas per frame, and recursively deepen the top three candidates.

```typescript
// From bench/run-evals.ts
run({ 
  framesPerRun: 5, 
  ideasPerFrame: 6, 
  topK: 3 
})

```

### Single-Shot Baseline

The baseline comparison uses a single LLM call governed by the `BASELINE_SYSTEM` constant. This prompt instructs the model to "Ideate on this engineering problem" without the multi-frame refinement process, creating a direct performance contrast against ADHD’s iterative approach.

## Bias Mitigation and Measurement

To ensure objective assessment, the evaluation methodology implements randomized presentation order. For each problem, the system randomly assigns which output appears as option "A" versus option "B" using `swapped = Math.random()`, preventing positional bias in human or automated evaluation.

## Summary

- The ADHD evaluation methodology relies on [`bench/problems.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/problems.json) to define six diverse engineering problems ranging from cache design to service naming.
- The engine runs with **5 frames**, **6 ideas per frame**, and **topK = 3** via the `run()` function in [`bench/run-evals.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/run-evals.ts).
- A **single-shot baseline** using the `BASELINE_SYSTEM` prompt provides the control group for comparison.
- **Randomized A/B ordering** (`swapped = Math.random()`) eliminates presentation bias during evaluation.
- This end-to-end benchmark provides reproducible metrics for measuring ADHD’s iterative ideation improvements over standard LLM inference.

## Frequently Asked Questions

### What specific engineering problems does ADHD use for evaluation?

The benchmark includes six curated challenges: LRU cache design, CLI hang handling, rate-limiting implementation, debugging exercises, monolith splitting architecture, and naming a feature-flag service. These scenarios test the engine’s ability to handle diverse software engineering tasks.

### How does the ADHD engine configuration differ from the baseline during testing?

The ADHD engine executes a multi-frame process defined in [`bench/run-evals.ts`](https://github.com/UditAkhourii/adhd/blob/main/bench/run-evals.ts) with parameters `framesPerRun: 5`, `ideasPerFrame: 6`, and `topK: 3`, while the baseline makes a single LLM call using the concise `BASELINE_SYSTEM` prompt without iterative refinement.

### Why does ADHD randomize the A/B presentation order?

The code assigns `swapped = Math.random()` to randomize which output (ADHD or baseline) appears as option A or B, mitigating positional bias that could influence evaluators during blind assessment of solution quality.

### Where is the evaluation problem set defined in the repository?

The problem definitions reside in [`bench/problems.json`](https://github.com/UditAkhourii/adhd/blob/main/bench/problems.json), which serves as the canonical source for the six engineering challenges used across all benchmark runs.