ADHD Evaluation Methodology: How the Engine Measures Performance Gains Against Single-Shot LLMs

ADHD uses a reproducible A/B benchmark that compares its multi-frame ideation engine against a single-shot baseline on six curated engineering problems, measuring improvements through randomized blind evaluation.

The evaluation methodology in the UditAkhourii/adhd repository establishes a rigorous, end-to-end testing framework that quantifies how the ADHD engine’s structured ideation process outperforms standard single-shot prompting. This systematic approach ensures that performance claims are backed by reproducible experiments across diverse software engineering domains.

Benchmark Design and Problem Curation

The foundation of ADHD’s evaluation rests on a curated problem set defined in bench/problems.json. This JSON file specifies six distinct open-ended engineering challenges spanning multiple domains, including LRU cache design, CLI hang handling, rate-limiting implementations, debugging scenarios, monolith splitting strategies, and naming conventions for feature-flag services.

Comparative Generation Protocol

ADHD Engine Configuration

The ADHD engine executes with specific hyperparameters orchestrated through bench/run-evals.ts. The system invokes the run() function with a configuration object specifying framesPerRun: 5, ideasPerFrame: 6, and topK: 3, enabling the engine to generate five ideation frames, produce six ideas per frame, and recursively deepen the top three candidates.

// From bench/run-evals.ts
run({ 
  framesPerRun: 5, 
  ideasPerFrame: 6, 
  topK: 3 
})

Single-Shot Baseline

The baseline comparison uses a single LLM call governed by the BASELINE_SYSTEM constant. This prompt instructs the model to "Ideate on this engineering problem" without the multi-frame refinement process, creating a direct performance contrast against ADHD’s iterative approach.

Bias Mitigation and Measurement

To ensure objective assessment, the evaluation methodology implements randomized presentation order. For each problem, the system randomly assigns which output appears as option "A" versus option "B" using swapped = Math.random(), preventing positional bias in human or automated evaluation.

Summary

  • The ADHD evaluation methodology relies on bench/problems.json to define six diverse engineering problems ranging from cache design to service naming.
  • The engine runs with 5 frames, 6 ideas per frame, and topK = 3 via the run() function in bench/run-evals.ts.
  • A single-shot baseline using the BASELINE_SYSTEM prompt provides the control group for comparison.
  • Randomized A/B ordering (swapped = Math.random()) eliminates presentation bias during evaluation.
  • This end-to-end benchmark provides reproducible metrics for measuring ADHD’s iterative ideation improvements over standard LLM inference.

Frequently Asked Questions

What specific engineering problems does ADHD use for evaluation?

The benchmark includes six curated challenges: LRU cache design, CLI hang handling, rate-limiting implementation, debugging exercises, monolith splitting architecture, and naming a feature-flag service. These scenarios test the engine’s ability to handle diverse software engineering tasks.

How does the ADHD engine configuration differ from the baseline during testing?

The ADHD engine executes a multi-frame process defined in bench/run-evals.ts with parameters framesPerRun: 5, ideasPerFrame: 6, and topK: 3, while the baseline makes a single LLM call using the concise BASELINE_SYSTEM prompt without iterative refinement.

Why does ADHD randomize the A/B presentation order?

The code assigns swapped = Math.random() to randomize which output (ADHD or baseline) appears as option A or B, mitigating positional bias that could influence evaluators during blind assessment of solution quality.

Where is the evaluation problem set defined in the repository?

The problem definitions reside in bench/problems.json, which serves as the canonical source for the six engineering challenges used across all benchmark runs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →