# How ADHD Balances Ideation Breadth and Token Cost in AI Agents

> Discover how ADHD agents balance idea breadth and token cost by generating richer ideas exponentially while accepting linear costs. Learn more about this AI tradeoff.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: deep-dive
- Published: 2026-07-30

---

**ADHD handles the breadth-cost tradeoff by spawning parallel divergent thought branches and filtering them through a critic phase, accepting linearly scaling token costs for exponentially richer idea generation.**

The [UditAkhourii/adhd](https://github.com/UditAkhourii/adhd) repository implements a structured reasoning framework that deliberately trades computational expense for creative breadth. Rather than generating a single answer, the system runs multiple isolated cognitive frames in parallel, then scores, clusters, and deepens the most promising results.

## The Two-Phase Architecture

ADHD operates on a **diverge-then-focus** loop implemented in [[`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts)](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts). This architecture separates wide exploration from critical evaluation, ensuring that cost is incurred only after breadth has been established.

### Parallel Divergence Frames

During the divergence phase, the engine spawns **N isolated frames** (default `N = 5`), each operating as an independent Claude-Code session with full context. Each frame explores the problem from a distinct cognitive angle—biology, logistics, game theory, or systems design—without cross-contamination.

According to the evaluation data in the [README](https://github.com/UditAkhourii/adhd/blob/main/README.md#results), this parallelization increases breadth scores from 4.8 to 9.0 (approximately 1.9× improvement) and novelty by 2×–3× compared to single-shot generation.

### The Critic and Deepening Phase

After divergence, a **critic phase** scores the generated ideas and performs clustering to identify the **top-K** candidates (default `K = 3`). These survivors undergo a deepening pass where the system expands rough concepts into detailed execution plans, including trap detection and non-obvious implications.

## Understanding the Cost Model

The cost structure is not merely "number of calls multiplied by average token count." Instead, it follows a substrate-heavy formula documented in [[`documentation/when-to-use.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md)](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md#cost--speed):

```typescript
cost ≈ N × (base_context + branch_output)   // divergence phase
     + critic_context                       // scoring & clustering
     + K × deepen_context                    // focus passes

```

### The Hidden Cost of Base Context

The dominant expense is the **base context** that each branch must reload. In the Claude-Code integration, every divergence branch carries the full session substrate (approximately 26,000 tokens by default). With `N = 5` branches, the substrate alone consumes ~130,000 tokens before any novel content is generated.

This explains why the intuitive "5–10× cost" estimate often understates the actual expense for integrated agent workflows versus the standalone library.

## Configuring the Tradeoff

ADHD exposes parameters to tune the breadth-cost ratio based on problem stakes and budget constraints.

### Adjusting N and K

- **Reduce `framesPerRun`** (N) when exploring familiar problem spaces or when token budgets are tight. Setting `N = 3` cuts substrate costs by 40%.
- **Lower `topK`** (K) to minimize deepening passes. Setting `K = 1` focuses resources on the single most promising idea rather than developing three alternatives.

Use the CLI flags to adjust these dynamically:

```bash

# Reduce cost by lowering breadth and depth

adhd "Design a caching strategy" --frames 3 --top 1

```

### Standalone vs Claude-Code Integration

The [**standalone library**](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md) (`import { run } from "adhd-agent"`) maintains a minimal context substrate, keeping costs closer to the theoretical N× multiplier. Conversely, the Claude-Code skill version ([[`skills/adhd/SKILL.md`](https://github.com/UditAkhourii/adhd/blob/main/skills/adhd/SKILL.md)](https://github.com/UditAkhourii/adhd/blob/main/skills/adhd/SKILL.md)) incurs higher base context costs but benefits from rich session history.

## When to Accept Higher Costs

The framework is designed for **high-stakes decisions** where a cheap but wrong answer is more expensive than a thorough exploration. According to the [usage guidelines](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md#use-it-for), deploy ADHD for architecture decisions, API design, and complex debugging—scenarios where trap detection and non-obvious alternatives justify the token spend.

Avoid the breadth-cost tradeoff for low-stakes lookups, syntax questions, or real-time autocomplete where single-shot latency and cost are paramount.

## Implementation Example

Configure the breadth-depth tradeoff programmatically in [[`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts)](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts):

```typescript
import { run, renderText } from "adhd-agent";

// High-breadth configuration for critical design reviews
const result = await run({
  problem: "Design a resilient rate-limiter for distributed services",
  framesPerRun: 5,   // N parallel exploratory frames
  topK: 3,           // Deepen the best 3 ideas
});

console.log(renderText(result));
// Access result.shortlist (all ideas), result.nonObviousPick (most novel),
// result.traps (flagged pitfalls), and result.deepened (detailed sketches)

```

## Summary

- **ADHD uses parallel divergence** to multiply idea breadth by spawning N isolated cognitive frames, each reloading the full base context.
- **Cost scales linearly with N** due to the ~26K token substrate per branch, making the total expense approximately `N × base_context` plus critique and deepening costs.
- **Tune the tradeoff** by reducing `framesPerRun` or `topK`, or by using the standalone library to minimize base context overhead.
- **Deploy strategically** for complex architectural decisions where premature convergence is riskier than higher token bills.

## Frequently Asked Questions

### How much does ADHD increase token usage compared to a single LLM call?

ADHD typically consumes **5–10× the tokens** of a single call when using the standalone library, but can reach **20–30×** in Claude-Code integrated mode due to the 26,000-token base context substrate that each of the N branches must reload. The exact multiplier depends on your `framesPerRun` (N) and `topK` settings.

### Can I reduce costs while keeping the divergent thinking benefits?

Yes. Reduce `framesPerRun` to 3 instead of 5 for a 40% cost reduction while maintaining multi-perspective coverage. Alternatively, set `topK` to 1 to eliminate multiple deepening passes. Using the **standalone npm package** rather than the Claude-Code skill also minimizes base context overhead.

### Why does the base context matter more than the number of API calls?

Because ADHD is optimized for quality over call count. The dominant cost is not the generation tokens but the **context window** required to maintain each branch's cognitive state. Each parallel frame independently loads the full session history (prompts, file contents, and prior reasoning), making the substrate cost proportional to N regardless of how short the actual generation is.

### Is ADHD suitable for real-time applications?

No. The deliberate tradeoff of token cost for breadth makes ADHD inappropriate for per-keystroke autocomplete, low-latency chat, or real-time streaming. The framework is designed for offline or asynchronous decision-making where spending 30–60 seconds and thousands of tokens is acceptable to avoid architectural traps or missed alternatives.