How Clustering in ADHD Identifies Underlying Angles Instead of Surface Keywords
ADHD's clustering phase groups generated ideas by their underlying design angle rather than literal keywords, using a specialized LLM prompt to expose the structural shape of the idea space.
Clustering in ADHD to identify underlying angles versus keywords occurs after the divergent generation and scoring phases in the UditAkhourii/adhd repository. This step transforms a flat list of ideas into a strategic map of distinct architectural approaches. By directing the LLM to ignore surface vocabulary and focus on structural similarities, the system ensures subsequent pruning steps operate on genuinely distinct strategic directions.
The Philosophy: Angles Over Keywords
Why Surface Keywords Fail
Traditional keyword-based clustering groups ideas that share similar vocabulary, which often misses the deeper strategic intent. Two ideas might use completely different terminology yet propose the identical architectural move, while ideas with overlapping words might represent fundamentally different approaches. ADHD avoids this trap by explicitly instructing the model to look beyond literal text matches.
The Underlying Angle Approach
The system prompt for clustering, defined in CLUSTER_SYSTEM within src/engine.ts, explicitly directs the LLM:
"You group ideas into 3‑6 clusters by their UNDERLYING ANGLE (not by surface keywords). Cluster labels name the angle, e.g. 'remove‑the‑server plays', 'push‑work‑to‑client plays', 'cache‑shaped plays'."
This instruction forces the model to identify common motivations or architectural moves—such as eliminating a server or moving work to the client—regardless of the exact wording used in each idea.
Implementation: How the Clustering Engine Works
The System Prompt Strategy
The clusterIdeas() function in src/engine.ts constructs a prompt that provides only the raw text of each idea without pre-categorized metadata. This forces the LLM to infer higher-level relationships based on structural similarities rather than provided tags.
The prompt format lists each idea as id :: text:
const userPrompt = `PROBLEM:
${problem}
IDEAS:
${ideas.map((i) => `${i.id} :: ${i.text}`).join("\n")}
Output JSON: [{"label":"...","ideaIds":["...","..."]}]`;
The clusterIdeas Function
The clustering invocation accepts three critical parameters:
problem– the original problem statement against which all ideas are judgedallIdeas– the flat list of generated leaf ideas, each containing anidandtextfieldcritic– the model instance used for scoring and clustering (which may differ from the generator model)
The function signature effectively operates as:
const clusters = await clusterIdeas(problem, allIdeas, critic);
Parsing and Validation
The LLM must return a JSON array conforming to ClusterSchema, defined in src/types.ts as [{ label: string, ideaIds: string[] }]. If parsing fails, the system gracefully returns an empty list rather than crashing. This validation ensures that only properly structured cluster data propagates to downstream rendering steps.
From Clusters to Output
Storing Cluster Labels
After the LLM returns valid clusters, each idea is stamped with its corresponding angle label for later use. The implementation in src/engine.ts iterates through the cluster assignments:
for (const c of clusters) for (const id of c.ideaIds) {
const idea = allIdeas.find((x) => x.id === id);
if (idea) idea.cluster = c.label;
}
This mutation attaches the underlying angle directly to each idea object, making the categorization available for filtering and display operations.
Rendering by Angle
The render.ts module consumes these cluster labels to produce human-readable output grouped by strategic direction rather than chronological generation order. The renderWideSet() function uses the cluster field to organize ideas:
export function renderWideSet(ideas: Idea[]) {
const byCluster = new Map<string, Idea[]>();
for (const i of ideas) {
const key = i.cluster ?? "(unclustered)";
if (!byCluster.has(key)) byCluster.set(key, []);
byCluster.get(key)!.push(i);
}
for (const [cluster, items] of byCluster) {
console.log(`## ${cluster}`);
for (const it of items) console.log(`- ${it.text}`);
}
}
This produces output where ideas appear under headings like ## remove-the-server plays or ## cache-shaped plays, making the architectural landscape immediately scannable.
Practical Usage
You can invoke the clustering step directly by calling the main run() function from src/engine.js. The following example demonstrates how to access cluster metadata after generation:
import { run } from "./src/engine.js";
async function demo() {
const result = await run({
problem: "How can we reduce latency in a web service?",
framesPerRun: 4,
ideasPerFrame: 5,
topK: 3,
concurrency: 2,
});
console.log("Clusters:");
for (const c of result.clusters) {
console.log(`- ${c.label}: ${c.ideaIds.length} ideas`);
}
}
demo();
Summary
- Angle-based grouping: ADHD clustering explicitly ignores surface keywords in favor of underlying design angles through the
CLUSTER_SYSTEMprompt insrc/engine.ts. - Structured output: The
clusterIdeas()function parses LLM responses againstClusterSchemato produce validated{ label, ideaIds }objects. - Persistent labeling: Each idea receives a
clusterproperty storing its angle label for downstream processing. - Hierarchical rendering: The
render.tsmodule uses these labels to display ideas grouped by strategic direction rather than generation order.
Frequently Asked Questions
What is the difference between keyword clustering and angle clustering in ADHD?
Keyword clustering groups ideas that share similar vocabulary, which often results in redundant strategic categories. Angle clustering, as implemented in src/engine.ts, groups ideas by their underlying architectural motivation—such as "push work to client" versus "optimize server response"—regardless of whether the ideas use different words to describe the same concept.
How does the LLM know to cluster by underlying angles?
The system injects explicit instructions via the CLUSTER_SYSTEM constant, which tells the model: "You group ideas into 3‑6 clusters by their UNDERLYING ANGLE (not by surface keywords)." This prompt engineering technique forces the LLM to analyze structural similarities and architectural trade-offs rather than performing simple text similarity matching.
What happens if the LLM returns malformed cluster JSON?
If the LLM response cannot be parsed against the ClusterSchema defined in src/types.ts, the clusterIdeas() function returns an empty list. This defensive programming approach ensures that downstream rendering and shortlisting steps continue to function, treating the ideas as unclustered rather than crashing the pipeline.
How are clustered ideas displayed in the final output?
The renderWideSet() function in src/render.ts maps each idea to its cluster label (or "(unclustered)" if none exists), groups them into a Map keyed by cluster name, and prints them under markdown headings corresponding to each underlying angle. This produces a structured view where distinct strategic directions appear as separate sections.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →