# How the Clustering Phase in ADHD Surfaces the Shape of the Idea Space

> Explore how the clustering phase in ADHD organizes ideas, revealing the structured map of your solution space. Understand idea density, diversity, and macro-structure.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: deep-dive
- Published: 2026-07-30

---

**The clustering phase in ADHD groups generated ideas by their underlying conceptual angles, transforming a flat list into a structured map that reveals the density, diversity, and macro-structure of the solution space.**

The **ADHD** engine (available at UditAkhourii/adhd) implements a Tree-of-Thought pipeline that diverges, scores, and clusters ideas to solve complex problems. While the divergence phase generates raw candidates and the scoring phase ranks them by quality, the **clustering phase in ADHD** performs the critical work of surfacing the geometry of the idea space. This step transforms an unstructured list into a navigable conceptual map by grouping ideas according to their fundamental strategic angles rather than surface-level keywords.

## Why Clustering Matters After Divergence

The ADHD engine begins its pipeline with a divergent fan-out that produces a large set of raw ideas (implemented in **[`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts)** lines 5‑9). At this stage, ideas exist as a flat list without higher-order structure. By grouping ideas based on their **underlying angle** rather than surface keywords, the clustering step creates a high-level map of the conceptual landscape.

Each cluster label (for example, “peer‑to‑peer sync” or “CRDT‑based document model”) acts as a coordinate that summarizes a family of related ideas. This allows the system—and the user—to see *how many distinct directions* exist and *where ideas concentrate*, providing the necessary context for the subsequent “focus / deepen” phase to select promising paths without getting lost in individual leaf nodes.

## Implementation of the Clustering Phase

### The clusterIdeas Function

The core clustering logic resides in the `clusterIdeas` function within **[`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts)** (lines 31‑38, 91‑95, and 122‑130). This function sends the full set of generated ideas to the LLM alongside a dedicated system prompt (`CLUSTER_SYSTEM`).

The prompt explicitly instructs the model to *“group ideas into 3‑6 clusters by their UNDERLYING ANGLE (not by surface keywords)”* and to return a JSON array of objects containing `label` and `ideaIds` fields. This design forces the LLM to analyze conceptual strategy rather than perform simple keyword matching.

### Schema Validation and Type Safety

The LLM’s response is parsed using `ClusterSchema`, defined at **lines 44‑46** in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts). This Zod schema validates that the response conforms to the expected structure before processing.

The type definitions for the cluster data model reside in **[`src/types.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/types.ts)** (lines 38‑41), where the `Cluster` interface specifies the shape of each cluster object. Rigorous typing ensures that downstream code can safely access `cluster.label` and `cluster.ideaIds` without runtime ambiguity.

### Annotating Ideas with Cluster Labels

After extracting the clusters, the engine annotates each `Idea` object with its corresponding cluster label (implemented at **lines 73‑76** in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts)). This annotation attaches metadata directly to the idea nodes, allowing the rendering layer to display the shape of the space alongside individual scores. Each idea now carries a pointer to its conceptual family, enabling filtered views and cluster-aware navigation.

## Surfacing the Shape: Interpreting Cluster Output

The clustering phase exposes three dimensions of the idea space’s macro-structure:

- **Number of clusters** – Indicates how many distinct conceptual families the LLM discovered. A higher count suggests a diverse solution space, while fewer clusters reveal concentrated thinking.
- **Cluster sizes** – Highlight dense regions where many ideas share a common angle versus sparse, potentially novel regions. Large clusters signal conventional approaches; small or singleton clusters may indicate unique, unexplored directions.
- **Cluster labels** – Provide human-readable descriptors that make the abstract shape concrete. These labels serve as navigational markers during the pruning and deepening stages.

## Code Example: Accessing Cluster Structure

The following TypeScript example demonstrates how to invoke the ADHD engine and inspect the resulting clusters:

```ts
import { run } from "./engine.js";

(async () => {
  const result = await run({
    problem: "Build a collaborative text‑editor that works offline",
    framesPerRun: 5,
    ideasPerFrame: 6,
    topK: 3,
  });

  // 👉 Clusters describe the shape of the idea space
  console.log("Clusters:");
  for (const c of result.clusters) {
    console.log(`- ${c.label} (${c.ideaIds.length} ideas)`);
  }

  // Each idea now knows its cluster label
  console.log("\nIdeas with cluster tags:");
  for (const i of result.branches.flatMap(b => b.ideas)) {
    console.log(`${i.text} → ${i.cluster}`);
  }
})();

```

Running this snippet surfaces the shape of the solution space through cluster metadata:

```

Clusters:
- “peer‑to‑peer sync” (4 ideas)
- “CRDT‑based document model” (5 ideas)
- “local‑first UI abstractions” (3 ideas)

Ideas with cluster tags:
...

```

## Summary

- The **clustering phase in ADHD** transforms a flat list of generated ideas into a structured map by grouping them according to underlying conceptual angles.
- The `clusterIdeas` function in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) (lines 31‑38, 91‑95, 122‑130) orchestrates the clustering by prompting the LLM with `CLUSTER_SYSTEM` and validating output against `ClusterSchema` (lines 44‑46).
- Cluster metadata is attached to individual ideas (lines 73‑76), enabling downstream rendering of the idea space’s geometry.
- The resulting clusters reveal the **number of distinct families**, **density of solutions**, and **human-readable coordinates** that guide the focus and deepen phases.

## Frequently Asked Questions

### How does the clustering phase differ from simple keyword grouping?

The `CLUSTER_SYSTEM` prompt explicitly instructs the LLM to group by **“UNDERLYING ANGLE (not by surface keywords)”**. This forces the model to analyze the conceptual strategy or philosophical approach of each idea rather than performing lexical similarity matching. Consequently, two ideas using different vocabulary but sharing the same strategic approach will cluster together.

### Where are the TypeScript interfaces for clusters defined?

The `Cluster` interface is defined in **[`src/types.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/types.ts)** at lines 38‑41, while the runtime validation schema `ClusterSchema` resides in **[`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts)** at lines 44‑46. These definitions ensure type safety when the engine processes the LLM’s JSON response containing cluster labels and idea ID arrays.

### How many clusters does the ADHD engine typically generate?

According to the system prompt implementation in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts), the engine requests **3‑6 clusters** from the LLM. The exact count depends on the diversity of ideas generated during the divergence phase; the model dynamically selects the appropriate number within this range to best represent the conceptual landscape.

### How are cluster labels attached to individual ideas?

After parsing the LLM response through `ClusterSchema`, the engine iterates through the returned clusters and annotates each corresponding `Idea` object with its cluster label. This occurs at **lines 73‑76** in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts), ensuring every idea carries metadata linking it to its conceptual family for downstream filtering and display.