# Cost and Latency Implications of Running 5 Frames × 6 Ideas with Critic Passes in ADHD

> Discover the cost and latency of 5 frames x 6 ideas with critic passes in ADHD. Learn how it impacts LLM calls, token cost, and wall-clock time for efficient performance.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: performance
- Published: 2026-08-19

---

**Running 5 frames with 6 ideas each and full critic passes in ADHD consumes roughly 3S + K LLM calls per run, driving linear token cost and wall-clock latency that typically stays within 2–3× a single LLM call under default concurrency.**

The ADHD repository by **UditAkhourii/adhd** orchestrates divergent reasoning through parallel cognitive frames, and every production deployment must budget for the token and time footprint of a full run. When you execute a configuration with 5 frames, 6 ideas per frame, and the default three-phase critic pipeline—scoring, clustering, and deepening—the framework follows a predictable resource model. Understanding the exact cost and latency implications of running 5 frames × 6 ideas with critic passes ensures you can tune parameters before scaling.

## How the ADHD Pipeline Consumes LLM Calls

The divergent phase in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) launches **S independent cognitive frames**. For each frame, the generator produces **6 ideas** in a single LLM call via the `diverge` routine. After generation, three distinct **critic passes** run per frame:

- **Score** — assigns a numeric quality score to each idea (1 call per frame via `engine.scoreIdeas`)
- **Cluster** — groups similar ideas and prunes traps (1 call per frame via `engine.clusterIdeas`)
- **Deepen** — re-queries the top **K** ideas, defaulting to 3, to elaborate them (K calls via `engine.deepenIdeas`)

That yields a total call count of:

```text
Total calls = S (diverge) + S (score) + S (cluster) + K (deepen)
            = 3S + K

```

With the default deepen pool of **K = 3**, a standard run issues **3S + 3 LLM calls**.

## Token Cost Breakdown

According to [`documentation/when-to-use.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md), ADHD models divergence cost as:

```text
cost ≈ N × (base_context + branch_output)

```

In this formula, **N equals S**, the number of frames. Because each frame reloads the full base context and appends its own output tokens, the 6-idea payload directly inflates the **branch_output** term. The critic passes add a roughly equivalent linear factor per frame; they reuse the same frame prompt as the base context and emit only a small scoring or clustering payload. The deepen stage adds another linear term scaled by **K**, since each selected idea is submitted as a separate follow-up call.

Consequences of this model include:

- **Token spend** grows linearly with frame count.
- **ideasPerFrame** increases output tokens per branch but does not multiply the number of network calls.
- The dominant cost driver is the number of parallel branches (frames), not the critic pass count alone.

## Latency and Parallel Execution

As documented in [`documentation/how-it-works.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/how-it-works.md), ADHD executes divergent and critic calls under a configurable semaphore. The default **concurrency** limit is **4**, meaning the runtime parallelizes up to four simultaneous LLM requests.

The expected latency follows this batching pattern:

```text
latency ≈ (⌈S / C⌉ + ⌈S / C⌉ + ⌈K / C⌉) × average_llm_latency

```

Here **C** is the concurrency limit. The first term covers the initial `diverge` wave, the second covers the combined critic passes, and the third covers the `deepen` stage. In practice, with the default settings—**S = 5**, **K = 3**, and **C = 4**—a full run completes in approximately **2–3× the latency of a single LLM call**, because the small critic payloads and pipeline overlap keep the wall-clock time well below the raw call count.

## Configuring a Run in Code

You can parameterize a run directly through the main entry point. The snippet below sets up **S = 5** frames with **6 ideas** each and the default critic pipeline:

```typescript
import { run } from "./src/index.js";

const result = await run({
  problem: "How can we make our API more resilient?",
  framesPerRun: 5,      // S frames
  ideasPerFrame: 6,     // 6 ideas per frame
  topK: 3,              // ideas to deepen (K)
  // Optional: override the model used for critic passes
  // criticModel: "claude-2.1",
});

```

Internally, this invokes `engine.diverge` for the initial generation wave, followed by `engine.scoreIdeas`, `engine.clusterIdeas`, and finally `engine.deepenIdeas`. Each stage is traced through [`src/llm.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/llm.ts), which handles token accounting and request dispatch.

## Key Source Files

For readers who want to verify the behavior in source, the following files define the pipeline:

- [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) — orchestrates `diverge`, `scoreIdeas`, `clusterIdeas`, and `deepenIdeas`
- [`src/frames.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/frames.ts) — defines the cognitive frame payloads that form the base context
- [`src/llm.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/llm.ts) — low-level wrapper around the LLM SDK, managing token usage and latency
- [`documentation/when-to-use.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md) — documents the `cost ≈ N × (base_context + branch_output)` model
- [`documentation/how-it-works.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/how-it-works.md) — explains the semaphore-based parallel execution strategy
- [`documentation/frames.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/frames.md) — lists built-in frames and selection logic

## Summary

- A full run with **S frames**, **6 ideas**, and critic passes issues **3S + K LLM calls**, where **K** defaults to 3.
- Token cost scales **linearly** with frame count according to the formula in [`documentation/when-to-use.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md).
- ADHD parallelizes frames and critics via a semaphore with default **concurrency = 4**.
- Default wall-clock latency for **S = 5** and **K = 3** stays near **2–3× a single LLM call**.
- You can tune `framesPerRun`, `ideasPerFrame`, `topK`, and `concurrency` to balance quality against budget.

## Frequently Asked Questions

### How many LLM calls does one full run generate?

A run generates **3S + K** calls, where **S** is `framesPerRun` and **K** is `topK`. With the default configuration of 5 frames and a deepen pool of 3, that equals **18 LLM calls** per run.

### Does increasing ideasPerFrame raise token cost?

Yes. Raising `ideasPerFrame` increases the **branch_output** tokens emitted during each `diverge` call, which directly increases the per-frame token spend. However, it does not increase the number of network calls.

### Can I reduce latency without lowering the frame count?

Yes. You can raise the **concurrency** limit in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) or via configuration to allow more simultaneous LLM calls. Alternatively, you can skip the deepen stage or reduce `topK` to cut the final batch of calls.

### Where is the cost model officially documented?

The linear cost model is documented in [`documentation/when-to-use.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/when-to-use.md), which states `cost ≈ N × (base_context + branch_output)`. The parallel execution model that governs latency is described in [`documentation/how-it-works.md`](https://github.com/UditAkhourii/adhd/blob/main/documentation/how-it-works.md).