How the Clustering Phase in ADHD Surfaces the Shape of the Idea Space
The clustering phase in ADHD groups generated ideas by their underlying conceptual angles, transforming a flat list into a structured map that reveals the density, diversity, and macro-structure of the solution space.
The ADHD engine (available at UditAkhourii/adhd) implements a Tree-of-Thought pipeline that diverges, scores, and clusters ideas to solve complex problems. While the divergence phase generates raw candidates and the scoring phase ranks them by quality, the clustering phase in ADHD performs the critical work of surfacing the geometry of the idea space. This step transforms an unstructured list into a navigable conceptual map by grouping ideas according to their fundamental strategic angles rather than surface-level keywords.
Why Clustering Matters After Divergence
The ADHD engine begins its pipeline with a divergent fan-out that produces a large set of raw ideas (implemented in src/engine.ts lines 5‑9). At this stage, ideas exist as a flat list without higher-order structure. By grouping ideas based on their underlying angle rather than surface keywords, the clustering step creates a high-level map of the conceptual landscape.
Each cluster label (for example, “peer‑to‑peer sync” or “CRDT‑based document model”) acts as a coordinate that summarizes a family of related ideas. This allows the system—and the user—to see how many distinct directions exist and where ideas concentrate, providing the necessary context for the subsequent “focus / deepen” phase to select promising paths without getting lost in individual leaf nodes.
Implementation of the Clustering Phase
The clusterIdeas Function
The core clustering logic resides in the clusterIdeas function within src/engine.ts (lines 31‑38, 91‑95, and 122‑130). This function sends the full set of generated ideas to the LLM alongside a dedicated system prompt (CLUSTER_SYSTEM).
The prompt explicitly instructs the model to “group ideas into 3‑6 clusters by their UNDERLYING ANGLE (not by surface keywords)” and to return a JSON array of objects containing label and ideaIds fields. This design forces the LLM to analyze conceptual strategy rather than perform simple keyword matching.
Schema Validation and Type Safety
The LLM’s response is parsed using ClusterSchema, defined at lines 44‑46 in src/engine.ts. This Zod schema validates that the response conforms to the expected structure before processing.
The type definitions for the cluster data model reside in src/types.ts (lines 38‑41), where the Cluster interface specifies the shape of each cluster object. Rigorous typing ensures that downstream code can safely access cluster.label and cluster.ideaIds without runtime ambiguity.
Annotating Ideas with Cluster Labels
After extracting the clusters, the engine annotates each Idea object with its corresponding cluster label (implemented at lines 73‑76 in src/engine.ts). This annotation attaches metadata directly to the idea nodes, allowing the rendering layer to display the shape of the space alongside individual scores. Each idea now carries a pointer to its conceptual family, enabling filtered views and cluster-aware navigation.
Surfacing the Shape: Interpreting Cluster Output
The clustering phase exposes three dimensions of the idea space’s macro-structure:
- Number of clusters – Indicates how many distinct conceptual families the LLM discovered. A higher count suggests a diverse solution space, while fewer clusters reveal concentrated thinking.
- Cluster sizes – Highlight dense regions where many ideas share a common angle versus sparse, potentially novel regions. Large clusters signal conventional approaches; small or singleton clusters may indicate unique, unexplored directions.
- Cluster labels – Provide human-readable descriptors that make the abstract shape concrete. These labels serve as navigational markers during the pruning and deepening stages.
Code Example: Accessing Cluster Structure
The following TypeScript example demonstrates how to invoke the ADHD engine and inspect the resulting clusters:
import { run } from "./engine.js";
(async () => {
const result = await run({
problem: "Build a collaborative text‑editor that works offline",
framesPerRun: 5,
ideasPerFrame: 6,
topK: 3,
});
// 👉 Clusters describe the shape of the idea space
console.log("Clusters:");
for (const c of result.clusters) {
console.log(`- ${c.label} (${c.ideaIds.length} ideas)`);
}
// Each idea now knows its cluster label
console.log("\nIdeas with cluster tags:");
for (const i of result.branches.flatMap(b => b.ideas)) {
console.log(`${i.text} → ${i.cluster}`);
}
})();
Running this snippet surfaces the shape of the solution space through cluster metadata:
Clusters:
- “peer‑to‑peer sync” (4 ideas)
- “CRDT‑based document model” (5 ideas)
- “local‑first UI abstractions” (3 ideas)
Ideas with cluster tags:
...
Summary
- The clustering phase in ADHD transforms a flat list of generated ideas into a structured map by grouping them according to underlying conceptual angles.
- The
clusterIdeasfunction insrc/engine.ts(lines 31‑38, 91‑95, 122‑130) orchestrates the clustering by prompting the LLM withCLUSTER_SYSTEMand validating output againstClusterSchema(lines 44‑46). - Cluster metadata is attached to individual ideas (lines 73‑76), enabling downstream rendering of the idea space’s geometry.
- The resulting clusters reveal the number of distinct families, density of solutions, and human-readable coordinates that guide the focus and deepen phases.
Frequently Asked Questions
How does the clustering phase differ from simple keyword grouping?
The CLUSTER_SYSTEM prompt explicitly instructs the LLM to group by “UNDERLYING ANGLE (not by surface keywords)”. This forces the model to analyze the conceptual strategy or philosophical approach of each idea rather than performing lexical similarity matching. Consequently, two ideas using different vocabulary but sharing the same strategic approach will cluster together.
Where are the TypeScript interfaces for clusters defined?
The Cluster interface is defined in src/types.ts at lines 38‑41, while the runtime validation schema ClusterSchema resides in src/engine.ts at lines 44‑46. These definitions ensure type safety when the engine processes the LLM’s JSON response containing cluster labels and idea ID arrays.
How many clusters does the ADHD engine typically generate?
According to the system prompt implementation in src/engine.ts, the engine requests 3‑6 clusters from the LLM. The exact count depends on the diversity of ideas generated during the divergence phase; the model dynamically selects the appropriate number within this range to best represent the conceptual landscape.
How are cluster labels attached to individual ideas?
After parsing the LLM response through ClusterSchema, the engine iterates through the returned clusters and annotates each corresponding Idea object with its cluster label. This occurs at lines 73‑76 in src/engine.ts, ensuring every idea carries metadata linking it to its conceptual family for downstream filtering and display.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →