# ADHD Performance Characteristics: Latency, Token Usage, and Concurrency Impact

> Discover ADHD performance characteristics: latency, token usage, and concurrency impact. Understand how LLM branches and prompt substrate affect costs.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: performance
- Published: 2026-08-01

---

**ADHD typically completes a run in 30–90 seconds using 4 concurrent LLM branches, but token costs scale linearly with frame count because each isolated branch reloads the full system prompt substrate.**

The [UditAkhourii/adhd](https://github.com/UditAkhourii/adhd) repository implements a divergent thinking engine for LLMs that trades computational overhead for reasoning quality. Understanding its performance characteristics—specifically how latency, token usage, and concurrency interact—is essential for cost estimation and production deployment. The architecture intentionally isolates reasoning branches to guarantee divergent output, which fundamentally shapes its resource consumption model.

## How ADHD Manages Parallel Execution

The ADHD engine orchestrates multiple independent LLM sessions to explore a problem space simultaneously. Each branch operates as a clean-slate conversation, ensuring that ideas develop without cross-contamination from sibling branches.

### The Branch Isolation Model

In [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts), the orchestration logic spawns isolated LLM calls through a concurrency semaphore. Because every branch is an independent session, the **first token** of any response cannot be generated until the entire *system prompt* (referred to as the "substrate") has been transmitted. This architectural decision means latency is dominated by the number of parallel branches rather than the length of the generated ideas themselves. A default run issues approximately **10 LLM calls**: 5 for divergence, 1 for scoring, 1 for clustering, and 3 for deepening.

### Default Concurrency Settings

The default **concurrency limit is 4**, defined in [`src/types.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/types.ts) within the `LLMOptions` interface and exposed via the CLI in [`src/cli.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/cli.ts). This setting controls the maximum number of simultaneous `callLLM` invocations allowed by the semaphore. The implementation in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) manages this queue, ensuring that exceeding the limit waits for in-flight requests to complete before spawning new ones.

## Latency Characteristics in ADHD

Wall-clock latency for a standard run ranges from **30 to 90 seconds** when using the default concurrency of 4. This duration encompasses the entire pipeline: divergent generation across multiple frames, scoring of candidates, clustering of similar ideas, and deepening of the top survivors.

Latency scales inversely with concurrency but linearly with frame count. With `concurrency = 1`, execution stretches to roughly **2–3× longer** because branches must execute sequentially. Conversely, increasing concurrency beyond 4 yields diminishing returns due to API rate limits and network bandwidth constraints, offering only marginal latency improvements while increasing the risk of throttling.

## Token Usage and Cost Model

Call count is a misleading metric for ADHD expenses. The critical cost driver is the **substrate multiplier**: each branch reloads the complete base context including system prompts, host-session state, and tool definitions.

If the base substrate measures approximately **26,000 tokens**, then `N` frames consume `N × 26,000` tokens before any novel content is generated. The total token formula is:

```typescript
total_tokens ≈ N × base_context_tokens + Σ(idea_tokens)

```

Here, `idea_tokens` represent the actual generated content and typically constitute a small fraction of the total cost. When ADHD operates as a skill within a larger agent, the host's context is appended to each branch, further increasing the multiplier.

### Calculating Total Token Costs

To estimate costs programmatically, measure your substrate size and apply the linear model:

```typescript
const baseTokens = 26_000;      // Size of system prompt substrate
const N = 5;                    // Number of frames (default)
const ideaTokens = 200;         // Average tokens per generated idea
const total = N * baseTokens + (N * ideaTokens);

console.log(`≈ ${total.toLocaleString()} tokens`); 
// Output: ≈ 131,000 tokens for default settings

```

## Concurrency Impact on Performance

Concurrency controls parallelism without altering the total token expenditure. Each branch still pays the full substrate cost regardless of execution order.

| Setting | Wall-Clock Latency | Token Cost |
|---------|-------------------|------------|
| `concurrency = 4` (default) | 30–90 seconds | N × base_context + idea_tokens |
| `concurrency = 1` | ~2–3× longer | Identical to default |
| `concurrency = 8` | Slightly faster (API limits apply) | Identical to default |

The trade-off is strictly between **speed and API load**. Higher concurrency reduces latency but increases instantaneous API request volume, while lower concurrency stretches execution time while keeping the identical token budget.

## Measuring and Configuring Performance

### CLI Configuration

Control concurrency directly from the command line using the flag defined in [`src/cli.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/cli.ts):

```bash

# Default execution: 5 frames, concurrency 4

adhd "design a retry/timeout strategy for an LLM-calling CLI"

# Reduced concurrency: slower execution, identical token cost

adhd --concurrency 2 "design a retry/timeout strategy for an LLM-calling CLI"

```

### Programmatic Control

For TypeScript implementations, pass performance parameters to the `run` function:

```typescript
import { run } from "adhd";

await run({
  problem: "design a retry/timeout strategy for an LLM-calling CLI",
  frames: 5,          // Number of parallel divergent branches (N)
  ideasPerFrame: 6,   // Ideas generated per frame (k)
  topK: 3,            // Survivors selected for deepening (K)
  concurrency: 8,     // Max parallel LLM calls
});

```

### Benchmarking Latency

Measure execution time using standard Node.js timing methods:

```typescript
console.time("ADHD");
await run({ 
  problem: "Optimize database query performance",
  concurrency: 4 
});
console.timeEnd("ADHD"); 
// Prints: ADHD: 45.234s (typical range 30-90s)

```

For detailed implementation of the parallel execution logic, refer to [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts). The token cost model and isolation rationale are documented in [`documentation/when-to-use.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md) and [`documentation/how-it-works.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/how-it-works.md), while generated API documentation appears in [`docs/index.html`](https://github.com/UditAkhourii/adhd/blob/main/docs/index.html).

## Summary

- **Latency**: Default runs complete in 30–90 seconds with concurrency set to 4, scaling linearly with frame count and inversely with concurrency limits.
- **Token Usage**: Costs follow the formula `N × base_context_tokens`, making the substrate multiplier the dominant expense factor rather than generated idea length.
- **Concurrency**: Adjusting the concurrency parameter in [`src/types.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/types.ts) or via CLI changes wall-clock latency but leaves total token consumption unchanged.
- **Architecture**: Branch isolation in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) guarantees divergent reasoning by reloading the full system prompt for every parallel branch.

## Frequently Asked Questions

### Does increasing concurrency reduce token costs in ADHD?

No. Increasing concurrency reduces wall-clock latency by allowing more branches to execute simultaneously, but the total token count remains identical. Each branch still transmits the full substrate context regardless of execution order, so the cost model `N × base_context_tokens` stays constant.

### Why does ADHD reload the full system prompt for every branch?

The architecture enforces **branch isolation** to guarantee truly divergent reasoning. By treating each frame as an independent LLM session with no shared context beyond the initial substrate, the system prevents early ideas from biasing subsequent generation. This isolation is implemented in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) and documented in [`documentation/how-it-works.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/how-it-works.md).

### How can I estimate token usage before running ADHD?

Calculate the substrate size of your system prompt (including tool definitions and host context), then multiply by the number of frames (`N`). Add a small buffer for generated idea tokens. For example, with a 26,000-token substrate and 5 frames, expect approximately 130,000 input tokens plus output tokens before execution.

### What is the optimal concurrency setting for ADHD?

The default value of **4** balances latency against API rate limits for most use cases. Settings above 4 offer minimal latency improvements due to network bandwidth and provider throttling, while settings below 4 linearly increase execution time. Adjust based on your API tier's rate limits and your tolerance for wall-clock duration.