How Breadth, Novelty, and Trap Detection Are Measured in bench/judge.ts

The bench/judge.ts module implements an LLM-as-judge evaluator that scores design ideas on breadth, novelty, and trap detection using a 0-10 numeric rubric embedded in a system prompt, returning structured JSON verdicts for pairwise comparison.

The evaluation of open-ended design creativity requires quantitative metrics that capture both diversity and risk awareness. In the UditAkhourii/adhd repository, the bench/judge.ts file implements an LLM-as-judge pattern to measure breadth, novelty, and trap detection by comparing two generated outputs against explicit scoring criteria. This approach enables automated benchmarking of design ideas through structured prompt engineering rather than heuristic code-based evaluation.

The Verdict Structure and Scoring Range

Numeric Ratings from 0 to 10

Each dimension receives independent scores for outputs A and B, stored in the Verdict type as objects containing numeric ratings and explanatory text. According to lines 16-19 of bench/judge.ts, the type definition captures { a: number; b: number; reason: string } for each metric, ensuring every quantitative score includes qualitative justification.

Prompt-Based Rubric Definition

Rather than implementing algorithmic checks, the system encodes evaluation criteria directly into the LLM's system prompt. The rubric defines:

  • Breadth: "range of structurally DISTINCT angles. 10 minor variations of one idea = low breadth" (lines 31-33)
  • Novelty: "how many ideas are non-obvious-but-viable. The obvious textbook answer is NOT novel" (lines 33-34)
  • Trap detection: "does it name ideas that look good but are traps, with reasons?" (lines 34-35)

The LLM receives strict instructions to "Score on substance only" and "Output JSON only" (lines 29-40), constraining responses to the specified schema without conversational filler.

Execution Flow and API Implementation

The Judge Function

The judge function orchestrates the evaluation by constructing a user prompt containing the problem description, both candidate outputs, and explicit instructions to "Score both on the rubric" (lines 47-60). This prompt engineering approach treats the LLM as a trained evaluator rather than a generative model.

LLM Integration and Response Parsing

After building the prompts, bench/judge.ts invokes callLLM from src/llm.js with the rubric-based system prompt and the constructed user prompt (lines 71-73). The raw LLM response passes through parseJSON to extract the structured Verdict object, converting natural language evaluations into machine-readable scores. This separation of concerns allows the judging logic to remain declarative while the transport layer handles API communication.

Practical Implementation Example

To evaluate design ideas programmatically, import the judge function and provide the problem context along with two competing outputs:

import { judge } from "./bench/judge";

// Define the design challenge and candidate solutions
const problem = "Design a task-management tool for remote teams.";
const outputA = "A Kanban board with custom columns, real-time sync, and emoji tags.";
const outputB = "A calendar-centric planner that auto-generates daily stand-ups, integrates with Slack, and suggests blockers.";

// Execute the evaluation
const verdict = await judge(problem, outputA, outputB);

// Access dimensional scores
console.log("Breadth scores:", verdict.breadth);
console.log("Novelty scores:", verdict.novelty);
console.log("Trap detection:", verdict.trap_detection);

The returned verdict object follows the Verdict type structure:

{
  "breadth": { 
    "a": 6, 
    "b": 8, 
    "reason": "B explores a different paradigm (calendar) vs A's incremental board features." 
  },
  "novelty": { 
    "a": 4, 
    "b": 9, 
    "reason": "B introduces auto-generated stand-ups, a less common approach in existing tools." 
  },
  "trap_detection": { 
    "a": 3, 
    "b": 7, 
    "reason": "B explicitly warns about over-automation traps that could alienate users." 
  }
}

Summary

  • LLM-as-judge architecture: The bench/judge.ts module delegates evaluation to an LLM using carefully engineered prompts rather than deterministic algorithms.
  • Standardized 0-10 scale: Breadth, novelty, and trap detection receive numeric scores from 0 to 10 for both outputs A and B, accompanied by textual reasoning.
  • Explicit rubric encoding: Scoring criteria are embedded in the system prompt at lines 31-35, defining structurally distinct angles for breadth, non-obvious viability for novelty, and explicit risk identification for trap detection.
  • Structured output: The Verdict type enforces JSON responses containing dual scores and explanations, parsed via parseJSON after LLM invocation through callLLM.

Frequently Asked Questions

What is the scoring range for breadth, novelty, and trap detection in bench/judge.ts?

Each dimension receives a numeric rating from 0 to 10 for both outputs A and B. The Verdict type defined at lines 16-19 stores these as objects with a, b, and reason properties, ensuring quantitative scores always include qualitative justification.

How does the judge determine what constitutes "high breadth" versus "low breadth"?

The LLM evaluates breadth based on the rubric's definition of "structurally DISTINCT angles" (lines 31-33). Ten minor variations of a single concept would score low despite high volume, while diverse architectural approaches receive higher ratings regardless of implementation detail count.

Can I modify the rubric criteria without changing the code logic?

Yes, the rubric exists as a string literal within the system prompt construction (lines 29-40). Modifying these prompt strings changes the evaluation criteria without requiring changes to the judge function's TypeScript implementation or the Verdict parsing logic.

What dependencies does bench/judge.ts require for LLM communication?

The module imports callLLM and parseJSON from src/llm.js to handle API communication and response normalization. These utilities abstract the underlying LLM provider, allowing the judging logic to remain agnostic to specific model implementations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →