What Is the Role of the Critic in the ADHD Architecture? A Deep-Dive into the Evaluation Engine

The critic in ADHD is a separate LLM evaluation pass that scores, clusters, and prunes ideas after the divergence phase, then drives focused deepening for the most promising candidates.

The ADHD (Adaptive Divergence & Heuristic Discipline) system is a structured approach to creative problem-solving using large language models. Unlike single-pass generation, ADHD separates ideation from evaluation through a distinct critic component. This article examines the critic's architectural role, implementation details, and practical configuration using the UditAkhourii/adhd codebase.

The Critic's Position in the ADHD Pipeline

The ADHD system operates in four distinct phases. The critic owns the second phase entirely:

Phase Actor Function
1 — Diverge Generator Spawns N independent branches under different cognitive frames
2 — Score + Cluster Critic Evaluates ideas on novelty, viability, and fit; detects traps; clusters by underlying angle
3 — Deepen Generator Expands top-K non-trap ideas with sketches, risks, and sub-ideas
4 — Select Combined Surfaces the final recommendation with traps and provocations

In src/engine.ts, this separation is enforced programmatically. The generator runs first with one system prompt, then the critic runs with entirely different prompts (SCORE_SYSTEM and CLUSTER_SYSTEM)—often on a different model entirely.

Why the Critic Must Be Separate

The critic's isolation from the generator is not incidental. It is a load-bearing design choice with three concrete benefits.

Decorrelation of Errors

When --critic-model is supplied, the critic runs on a different model family than the generator. This prevents systematic errors from infecting both generation and evaluation. As implemented in src/engine.ts, the criticModel parameter overrides only the scoring and clustering passes while leaving divergence and deepening untouched.

Unbiased Trap Detection

The generator cannot "self-censor" when the critic is external. The critic explicitly labels ideas that appear attractive but carry hidden drawbacks, providing concrete, actionable warnings rather than vague risk labels. This trap detection occurs in the scoring phase where the critic outputs structured JSON parsed into a Score map.

Structured Pruning

The critic's JSON output is parsed into typed structures defined in src/types.ts. These Score objects and Cluster arrays drive the "focus" stage, transforming raw idea floods into curated shortlists.

Implementing the Critic: Code Examples

Running ADHD with a Custom Critic Model

import { run } from "./src/engine.js";

await run({
  problem: "Design a low-latency cache for a distributed service",
  framesPerRun: 5,
  ideasPerFrame: 6,
  topK: 3,
  model: "claude-3-opus-20240229",          // generator
  criticModel: "claude-3-sonnet-20240229", // critic (score + cluster)
});

The criticModel parameter controls which LLM handles evaluation, while model continues to power generation.

Inspecting Critic Output Programmatically

const result = await run({ problem, model, criticModel });

// Score objects attached to each idea
console.log("Scores:", result.branches
  .flatMap(b => b.ideas)
  .map(i => i.score));

// Angle-based cluster groupings
console.log("Clusters:", result.clusters);

// Detected traps with explanations
console.log("Traps:", result.traps.map(t => ({
  text: t.text,
  trap: t.score?.trap
})));

After the critic runs, every Idea in result.branches receives a populated score field, and result.clusters contains the angle-level groupings used for downstream selection.

CLI Configuration

npx adhd run \
  --problem "Refactor the monolith into microservices" \
  --model claude-3-opus-20240229 \
  --critic-model claude-3-sonnet-20240229

The --critic-model flag mirrors the programmatic API, enabling quick experimentation with different generator/critic pairings.

Core Implementation Files

File Purpose Key Lines
src/engine.ts Orchestrates the two-phase loop, creates critic pass, wires criticModel Lines 27-30, 66-71, 71-89, 80-86
src/types.ts Type definitions for Score, Cluster, Idea structures consumed by critic Full file
src/llm.ts Low-level Claude Agent SDK wrapper used by both generator and critic Full file
src/frames.ts Cognitive frames feeding the generator; untouched by critic Full file

The critic's behavior is concentrated in src/engine.ts. Lines 71-89 handle the critic's instantiation with opposite system prompts, while lines 80-86 implement trap detection logic. Lines 27-30 manage model selection decorrelation when criticModel differs from model.

How the Critic Differs from Chain-of-Thought and Tree-of-Thought

The documentation/vs-cot-and-tot.md file explains why ADHD's separate critic outperforms integrated evaluation. CoT mixes reasoning with generation, risking confabulation. ToT uses self-evaluation from the same model, permitting self-censorship. The ADHD critic's hard separation guarantees that evaluation criteria are applied consistently without contaminating the ideation process.

Summary

  • The critic is a dedicated LLM pass that runs after divergence, not during or before
  • It performs three functions: scoring (novelty, viability, fit), clustering (by underlying angle), and trap detection (hidden drawbacks)
  • Separation from the generator enables error decorrelation, unbiased evaluation, and structured pruning
  • Configure via --critic-model CLI flag or criticModel option in run() from src/engine.ts
  • Implementation lives primarily in src/engine.ts with type definitions in src/types.ts

Frequently Asked Questions

Can the critic run on the same model as the generator?

Yes, but this defeats a key design benefit. When criticModel is omitted, the critic defaults to the same model as the generator. The system still functions, but loses error decorrelation—the same systematic biases can affect both generation and evaluation.

What exact criteria does the critic use to score ideas?

According to src/engine.ts, the critic evaluates on novelty, viability, and fit to the original problem. It also tags traps—attractive ideas with hidden costs. These criteria are embedded in the SCORE_SYSTEM prompt used during the critic's LLM call.

How does trap detection work in practice?

The critic outputs a structured Score object with a trap field. When populated, this field contains a concrete explanation of why an idea that seems promising actually carries hidden risks. These traps are preserved through to final output so users see both the recommendation and its dangers.

Is the critic's clustering automatic or configurable?

The clustering is automatic based on underlying angles detected in the generated ideas. The CLUSTER_SYSTEM prompt drives this grouping, which produces the Cluster array stored in result.clusters. The topK parameter then determines how many clusters receive deepening treatment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →