How ADHD Balances Ideation Breadth and Token Cost in AI Agents

ADHD handles the breadth-cost tradeoff by spawning parallel divergent thought branches and filtering them through a critic phase, accepting linearly scaling token costs for exponentially richer idea generation.

The UditAkhourii/adhd repository implements a structured reasoning framework that deliberately trades computational expense for creative breadth. Rather than generating a single answer, the system runs multiple isolated cognitive frames in parallel, then scores, clusters, and deepens the most promising results.

The Two-Phase Architecture

ADHD operates on a diverge-then-focus loop implemented in [src/engine.ts](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts). This architecture separates wide exploration from critical evaluation, ensuring that cost is incurred only after breadth has been established.

Parallel Divergence Frames

During the divergence phase, the engine spawns N isolated frames (default N = 5), each operating as an independent Claude-Code session with full context. Each frame explores the problem from a distinct cognitive angle—biology, logistics, game theory, or systems design—without cross-contamination.

According to the evaluation data in the README, this parallelization increases breadth scores from 4.8 to 9.0 (approximately 1.9× improvement) and novelty by 2×–3× compared to single-shot generation.

The Critic and Deepening Phase

After divergence, a critic phase scores the generated ideas and performs clustering to identify the top-K candidates (default K = 3). These survivors undergo a deepening pass where the system expands rough concepts into detailed execution plans, including trap detection and non-obvious implications.

Understanding the Cost Model

The cost structure is not merely "number of calls multiplied by average token count." Instead, it follows a substrate-heavy formula documented in [documentation/when-to-use.md](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md#cost--speed):

cost ≈ N × (base_context + branch_output)   // divergence phase
     + critic_context                       // scoring & clustering
     + K × deepen_context                    // focus passes

The Hidden Cost of Base Context

The dominant expense is the base context that each branch must reload. In the Claude-Code integration, every divergence branch carries the full session substrate (approximately 26,000 tokens by default). With N = 5 branches, the substrate alone consumes ~130,000 tokens before any novel content is generated.

This explains why the intuitive "5–10× cost" estimate often understates the actual expense for integrated agent workflows versus the standalone library.

Configuring the Tradeoff

ADHD exposes parameters to tune the breadth-cost ratio based on problem stakes and budget constraints.

Adjusting N and K

  • Reduce framesPerRun (N) when exploring familiar problem spaces or when token budgets are tight. Setting N = 3 cuts substrate costs by 40%.
  • Lower topK (K) to minimize deepening passes. Setting K = 1 focuses resources on the single most promising idea rather than developing three alternatives.

Use the CLI flags to adjust these dynamically:


# Reduce cost by lowering breadth and depth

adhd "Design a caching strategy" --frames 3 --top 1

Standalone vs Claude-Code Integration

The standalone library (import { run } from "adhd-agent") maintains a minimal context substrate, keeping costs closer to the theoretical N× multiplier. Conversely, the Claude-Code skill version ([skills/adhd/SKILL.md](https://github.com/UditAkhourii/adhd/blob/main/skills/adhd/SKILL.md)) incurs higher base context costs but benefits from rich session history.

When to Accept Higher Costs

The framework is designed for high-stakes decisions where a cheap but wrong answer is more expensive than a thorough exploration. According to the usage guidelines, deploy ADHD for architecture decisions, API design, and complex debugging—scenarios where trap detection and non-obvious alternatives justify the token spend.

Avoid the breadth-cost tradeoff for low-stakes lookups, syntax questions, or real-time autocomplete where single-shot latency and cost are paramount.

Implementation Example

Configure the breadth-depth tradeoff programmatically in [src/engine.ts](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts):

import { run, renderText } from "adhd-agent";

// High-breadth configuration for critical design reviews
const result = await run({
  problem: "Design a resilient rate-limiter for distributed services",
  framesPerRun: 5,   // N parallel exploratory frames
  topK: 3,           // Deepen the best 3 ideas
});

console.log(renderText(result));
// Access result.shortlist (all ideas), result.nonObviousPick (most novel),
// result.traps (flagged pitfalls), and result.deepened (detailed sketches)

Summary

  • ADHD uses parallel divergence to multiply idea breadth by spawning N isolated cognitive frames, each reloading the full base context.
  • Cost scales linearly with N due to the ~26K token substrate per branch, making the total expense approximately N × base_context plus critique and deepening costs.
  • Tune the tradeoff by reducing framesPerRun or topK, or by using the standalone library to minimize base context overhead.
  • Deploy strategically for complex architectural decisions where premature convergence is riskier than higher token bills.

Frequently Asked Questions

How much does ADHD increase token usage compared to a single LLM call?

ADHD typically consumes 5–10× the tokens of a single call when using the standalone library, but can reach 20–30× in Claude-Code integrated mode due to the 26,000-token base context substrate that each of the N branches must reload. The exact multiplier depends on your framesPerRun (N) and topK settings.

Can I reduce costs while keeping the divergent thinking benefits?

Yes. Reduce framesPerRun to 3 instead of 5 for a 40% cost reduction while maintaining multi-perspective coverage. Alternatively, set topK to 1 to eliminate multiple deepening passes. Using the standalone npm package rather than the Claude-Code skill also minimizes base context overhead.

Why does the base context matter more than the number of API calls?

Because ADHD is optimized for quality over call count. The dominant cost is not the generation tokens but the context window required to maintain each branch's cognitive state. Each parallel frame independently loads the full session history (prompts, file contents, and prior reasoning), making the substrate cost proportional to N regardless of how short the actual generation is.

Is ADHD suitable for real-time applications?

No. The deliberate tradeoff of token cost for breadth makes ADHD inappropriate for per-keystroke autocomplete, low-latency chat, or real-time streaming. The framework is designed for offline or asynchronous decision-making where spending 30–60 seconds and thousands of tokens is acceptable to avoid architectural traps or missed alternatives.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →